Dynamic layout optimization of annotation labels in volume rendering
By optimizing the annotation label layout, considering the visibility of labels and regions of interest, and using iterative optimization and ray tracing simulation, the problem of low efficiency in annotation label layout in existing technologies is solved, achieving efficient, visible, and real-time optimization of labels in images.
Patent Information
- Application Number
- CN202310116467.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-02-10
- Filing Date
- 2023-02-09
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2043-02-09
AI Technical Summary
Existing technologies struggle to efficiently optimize the layout of annotation labels in images, especially in large images or image animation scenarios. Automatic label placement algorithms are computationally inefficient and cannot calculate the optimal solution in real time.
By considering the visibility of labels in the image and the visibility of regions of interest, an iterative optimization algorithm is used to optimize the label layout. Combined with ray tracing simulation and depth map generation, the position and opacity of the labels are determined to reduce occlusion between labels and between labels and other elements, thereby achieving a globally optimal and temporally coherent layout.
It improves the visibility and recognizability of annotation tags in images, reduces occlusion between tags and other elements, achieves real-time optimization in large image or image animation scenarios, and improves image review efficiency.
Smart Images

Figure CN116580398B_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims the benefit of EP 22156156.6, filed on February 10, 2022, which is incorporated herein by reference in its entirety. Technical Field
[0003] Various examples of this disclosure generally relate to annotation labels in rendered images. More specifically, various examples of this disclosure relate to optimizing the layout of annotation labels for displaying semantic contextual information corresponding to regions of interest in an image based on medical volume rendering. Background Technology
[0004] Annotation tags can be used to enhance images with textual semantic information, thereby explaining the visible features in the image to the viewer. For example, information about anatomical structures, organs, pathology, etc., can be used to enhance medical images, thereby facilitating radiologists' review of medical images. Annotation tags can be overlaid on images according to the annotation layout. Sometimes, a series of images needs to be annotated; the image sequence can form a movie. In such scenarios, the temporal dependence of the annotation layout is also determined.
[0005] Creating globally optimal and temporally coherent annotation layouts for labels in images without human intervention is a common task in imaging processes. Automatic label placement is an NP-hard problem and often cannot efficiently compute optimal solutions. Traditionally, genetic algorithms for automatically labeling mapping data are known, where layout optimization relies on heuristics inspired by the concept of natural selection in Darwinian evolution. In a simple example, the fitness function to be maximized is defined as tracking the number of non-overlapping labels, which drives a process of selection, crossover, and mutation of a pool of possible solutions.
[0006] However, for use cases with large images or image animations, these algorithms often cannot be computed in real time. Summary of the Invention
[0007] Therefore, advanced technologies are needed to optimize the layout of annotation labels in rendered images. In particular, technologies are needed to mitigate or alleviate at least some of the aforementioned limitations and drawbacks.
[0008] In the following description, the solution is explained with respect to the claimed method and the claimed computing device, computer program, and storage medium, wherein features, advantages, or alternative embodiments may be assigned to other claimed objects, and vice versa. In other words, features described in the context of the method may be used to improve the claims relating to the computing device, computer program, and storage medium.
[0009] Based on various examples, the layout of the labels is determined using measures of the region of interest in the image associated with the label and / or the visibility of the label—that is, defining the placement of the label relative to the image including features annotated by the label.
[0010] Labels may include and / or represent semantic information; for example, they may include a textual description of semantic information associated with the labeled feature / region of interest. Labels may include a graphical link between the textual description and the corresponding labeled feature; for example, a label may be anchored at an anchor point near or at the corresponding feature. A line may be drawn from the anchor point to the label text.
[0011] Labels may include one or more of 2D or 3D glyphs, 2D or 3D text boxes, surfaces, and connecting lines. It should be understood that any other graphic or text element can be used as a label.
[0012] This is based on the finding that, depending on the specific label layout, labels may occlude one or more other labels, regions of interest, and / or landmarks in the rendered image. Such visibility of at least a portion of the image should be used or considered when optimizing label layout.
[0013] A method is provided for optimizing label layout to annotate at least one region of interest and / or markers representing regions of interest in a rendered image based on an imaging dataset, the method comprising the following actions.
[0014] During the action, at least one image is acquired.
[0015] Generally, an image can be received from a data storage device or another computing device. The image can be loaded from a hospital's image repository (e.g., a Picture Archiving and Communication System (PACS)). In various examples, the image can be a rendered image, and / or obtaining at least one rendered image can include rendering an image from a dataset, such as an imaging dataset that includes imaging data, such as output data from an imaging process (e.g., a camera or MRI imaging method). The dataset can include volumetric or non-volumetric imaging data, particularly medical volumetric imaging data. Therefore, in various examples, an image based on an imaging dataset can be referred to as a rendered image.
[0016] In another action, based on an image, multiple regions of interest (ROIs) and / or landmarks corresponding to a region of interest (ROI) can be obtained in the image, and the locations of the ROIs and / or landmarks can be obtained. The ROIs in the image may correspond to parts or regions of particular interest to the user, i.e., image structures such as anatomical structures that display specific anatomical structures to the user. Landmarks may correspond to specific portions of the ROIs, which can represent the ROIs and can, for example, serve as anchor points for labels with semantic information. The locations of at least one ROI and / or ROIs, as well as the locations of the landmarks and / or landmarks, in at least one rendered image can be obtained. Obtaining these can include using the image, automatically or manually, or a combination thereof, determined or identified by an algorithm. Obtaining these can include receiving them through user input or from a database.
[0017] Various algorithms are known to facilitate the analysis of images, particularly medical images, to detect and locate regions of interest (ROIs). For example, certain anatomical regions or organs can be segmented. Bounding boxes for such ROIs can be determined. Such algorithms can employ machine learning techniques. The algorithm can be a neural network algorithm. Here, training is facilitated by obtaining ground truth values through supervised learning on multiple training images, and then training the neural network algorithm to detect and locate ROIs based on the ground truth values. Since the implementation of such algorithms for detecting ROIs and associated semantic information is generally known to those skilled in the art, further details are unnecessary to describe in this regard. Furthermore, the specific implementation of such algorithms is not closely related to the operation of the techniques described herein. This is because the techniques described herein involving determining label layouts can be flexibly coupled with different types and kinds of algorithms that determine the location of ROIs and / or obtain semantic information associated with the location of such ROIs.
[0018] In another action, semantic information associated with multiple regions of interest and / or landmarks is obtained, wherein obtaining may include receiving from a data storage device or other computing device, or determining the semantic information based on an image (e.g., based on the identified regions of interest and / or landmarks).
[0019] For example, classification algorithms can be used to categorize visible features within a region of interest. Contextual information can be obtained from algorithms used to determine the locations of multiple regions of interest. For instance, an algorithm can be trained to detect fractures; this algorithm can then provide the detected fracture type and / or severity as semantic information, serving as contextual information.
[0020] In some examples, such semantic information may also be input by the user via a human-computer interface.
[0021] In another action, based on semantic information and the positions of multiple regions of interest (ROIs) and / or landmarks, and considering the visibility of labels in the image and the further visibility of ROIs and / or landmarks, a layout of labels for annotating multiple landmarks in at least one rendered image is determined. Considering the visibility of labels in the image and the further visibility of ROIs and / or landmarks may include considering whether one or more of the ROIs and / or landmarks, and / or labels are visible, or whether they are occluded by another of one or more ROIs and / or landmarks, and / or labels. In other words, the criteria used to determine and optimize the label layout in the image are based on the visibility of these features to the user in the annotated image. In other words, the label layout can be determined and optimized based on the visibility of labels and / or ROIs and / or landmarks (i.e., representative portions of the ROIs). For example, considering visibility may include considering the opacity of the labels, and / or the contrast between one or more labels and / or the image background around the labels (i.e., the image portion directly adjacent to the labels), and / or the opacity determined for the labels.
[0022] In other words, label layout can be determined based on and / or using measures of visibility of one or more of the following: labels in the image, and / or landmarks, and / or regions of interest (e.g., anatomical structures) associated with the landmarks. When labels are displayed in the image, i.e., when the labels are displayed together with the image, the labels can represent at least a portion of semantic information in the image.
[0023] Labels can be overlaid on images for display. That is, the corresponding text description of semantic information can be overlaid on image pixels. Labels can be displayed at a certain opacity level.
[0024] Labels can be anchored at corresponding landmarks to annotate areas of interest.
[0025] The layout of a label may include size, position, style, graphic or text content and / or the type and number of label components, and / or one or more of the arrangement of the label in the image relative to other labels and / or landmarks.
[0026] For example, a landmark associated with a region of interest (ROI) can be defined by the center location of the ROI. A landmark associated with a ROI can also be defined by a feature specific to that ROI. In practical terms, the landmark to which a label is anchored might be the centroid of an anatomical structure. For instance, a landmark could correspond to a substructure of an organ, such as the top region of the liver.
[0027] In various examples, an initial layout can be determined, which can be optimized over multiple iterations.
[0028] Iterative numerical optimization can be used to seek to maximize or minimize the objective function. A global minimum or maximum value can be found by varying the parameters within a parameter space, which can be bounded by one or more constraints.
[0029] As a concrete example, between subsequent iterations, the label positions might be shifted / adjusted, starting from the label positions of the previous iteration. The direction and / or step size of such shifts can be defined by an optimization algorithm that typically considers the current value of the objective function or the impact of the current label position on the objective function. For example, gradient descent optimization can be used, which considers the change in the objective function as a function of the label layout. Genetic optimization can be used. Landweber optimization can be used.
[0030] Generally, label layout can be optimized using or based on semantic information and the location of landmarks and / or regions of interest (ROIs). Therefore, optimization can be based on the visibility of labels, ROIs, and / or landmarks in an image. In other words, the output of each iteration of the optimization process can be an improved label layout in the image, where the layout is optimized to improve the visibility of labels and their corresponding landmarks and ROIs in the image. For example, a metric for label layout quality can be determined using an objective function and the label layout as input. Based on the improved objective function value as a measure of the visibility of labels, ROIs, and landmarks when displayed with the rendered image, the output of the optimization iteration can be an improved label layout, such as by applying the objective function to the layout and what the image provides.
[0031] In optional actions, rendered images can be provided or displayed to the user along with labels, where the labels display semantic annotation information to the user in an improved way, which is more intuitive and easier for the user to identify.
[0032] Determining the label layout may include determining an initial layout and iteratively optimizing the initial layout based on an objective function that is based on the visibility of the labels and the further visibility of the landmarks.
[0033] For example, if a label is obscured or overlapped by another label, its visibility may be reduced. Alternatively or additionally, if the label is placed in an area of an underlying image with a contrast value similar to the label text, its visibility may be reduced.
[0034] For example, if the region of interest is obscured by a label placed on top of it, the visibility of the region of interest may be reduced.
[0035] For example, the visibility of a label can be reduced to various locations near the edges of an image.
[0036] The objective function can be based on at least one label and its corresponding landmark, or multiple labels relative to each other's proximity.
[0037] Therefore, proximity relationships can specify the relative arrangement between different labels or between labels and landmarks. For example, proximity relationships can use a metric that penalizes a larger distance between label text and label anchor points. Alternatively, proximity relationships can use a metric that penalizes a smaller distance between adjacent label text. Alternatively, proximity relationships can use a metric that penalizes overlapping label lines connecting anchor points and label text.
[0038] The objective function can be based on an occlusion rating associated with the label; the occlusion rating is based on the label's location in the image. An occlusion rating can represent the occlusion rate associated with a label or accumulated across all labels, such as occlusion by another label, landmark, and / or region of interest in the image. A larger occlusion rating may result in greater overlap. To give just a few examples, a smaller occlusion rating might be produced by a higher opacity value used for the label text.
[0039] The objective function can be based on the depth location of multiple landmarks in a volumetric medical imaging dataset.
[0040] The camera depth of certain features in an image can be determined through optical simulations of volumetric data (e.g., volumetric ray casting algorithms or ray tracing simulations), or in other words, the depth location can be represented by a corresponding depth value indicating the distance from the camera. The depth value of each of multiple segments in the image can be stored as a depth map of the image. In the optical simulation, a line of sight from the camera position and through pixels in the viewport is traced, and the optical opacity of the volumetric data is integrated along this line of sight—that is, accumulated. Points along the line of sight are identified where the accumulated opacity reaches certain thresholds. The camera depth (represented by depth values) at those locations is the distance from the point to the camera. Therefore, the depth location can be defined based on the accumulated opacity values in the optical simulation based on the volumetric dataset.
[0041] The opacity value of the overlay image can be defined based on the depth position determined for multiple locations along the rays used for ray tracing rendering of at least one rendered image. The depth position can be defined with respect to the distance between the camera position used for rendering and the corresponding region of interest. The depth position can be defined in a reference coordinate system in which the volumetric medical imaging dataset is defined.
[0042] At least one rendered image may comprise a movie sequence of multiple rendered images, where the objective function is based on the relative changes in layout between subsequent rendered images in the movie sequence. For example, to avoid abrupt changes in label layout, the objective function may penalize the larger optical flow associated with the label layout.
[0043] As will be understood from the above, various criteria for constructing the objective function have been disclosed. It should be understood that other criteria may be considered, for example, alone or in combination with the criteria disclosed above. The disclosed criteria may also be used in various combinations or individually.
[0044] The objective function can be based on multiple criteria, where the weighting factors associated with the multiple criteria are user-adjustable between subsequent iterations of the iterative optimization.
[0045] Therefore, closed-loop user interaction can be achieved when determining the label layout. That is, during the optimization process, the user can adjust the objective function based on the current layout of the labels that can be output to the user. Thus, the user can adjust the optimization criteria according to their needs during optimization, thereby obtaining low-latency feedback.
[0046] Iterative optimization is not necessary in all scenarios to determine the label layout. Other scenarios are also conceivable.
[0047] The label layout can be determined using an artificial neural network algorithm, which operates based on one or more inputs, namely semantic information, at least one rendered image, and the positions of multiple landmarks.
[0048] The layout can be a three-dimensional layout that defines the position of the labels in a reference coordinate system to which the volumetric medical imaging dataset is registered, wherein the method may include, based on the label layout, rendering a corresponding overlay image depicting the label for each of at least one rendered image.
[0049] Based on multiple opacity values determined for segments of the corresponding overlay image, each of at least one rendered image can be merged with the corresponding overlay image.
[0050] This means that opacity values can be determined not only for individual labels, but also for larger segments of overlaid images, each potentially including multiple labels.
[0051] Based on rendering an image from a dataset, a first depth map corresponding to the dataset image can be determined. The rendered image can be rendered using a ray tracing simulation, where multiple rays originating from the camera viewpoint are traced based on the imaging dataset. At each interaction / event affecting the rays based on the imaging dataset, the ray's opacity value can be updated, allowing the calculation of cumulative opacity values, including all previous interactions / events. A given segment in the rendered image can correspond to one or more rays in the ray tracing simulation. When the cumulative opacity value of the rays reaches a first opacity threshold at an event, a depth value corresponding to the event depth can be determined and stored in the first depth map. Therefore, each first depth value in the first depth map can be associated with a corresponding segment of the dataset image, where the first depth value corresponds to a depth at which the rays in the optical simulation (e.g., volumetric ray casting algorithm or ray tracing simulation) used to render the image associated with the segment reach the first cumulative opacity threshold originating from the camera viewpoint. For example, the depth of a segment in an annotated overlay image can be determined based on the position or depth of an intersecting image (i.e., an image plane in volumetric imaging data), and / or based on one or more landmarks, and / or based on the image or the rendering process of the image, and / or the layout of the labels, wherein the opacity of the segment in the annotated overlay image can be adjusted based on the depth of the segment in the annotated overlay image and a first depth map.
[0052] A second depth map corresponding to the dataset image can be correspondingly determined as a first depth map, comprising multiple depth values corresponding to each segment of the dataset image, wherein each second depth value of the second depth map is associated with a corresponding segment of the dataset image, wherein the second depth value corresponds to a depth at which the cumulative opacity of the rays in the volumetric ray casting algorithm used to render the image associated with the segment reaches a second cumulative opacity threshold at the camera viewpoint, wherein the second opacity threshold is higher than a first opacity threshold. The determined depth of the segment in the annotated overlay image is used and compared with the first and second depth values of the corresponding segment in the dataset image, and the opacity of the segment in the annotated overlay image is adjusted based on the depth of the segment in the annotated overlay image, the first depth value, and the second depth value. The first and second depth maps can be generated and used in the same manner, i.e., a description of the first or second depth map can be applied to the other of the first or second depth maps.
[0053] Adjusting the opacity of each annotation fragment in the annotation overlay image may include determining full opacity for the annotation image fragment if the depth value of the annotation image fragment is less than a first depth value of the dataset image fragment corresponding to the annotation fragment, and / or determining a scaled opacity value, which is a value between full opacity (typically 1) and full transparency (typically 0), by comparing the depth value of the annotation image fragment with the corresponding first and / or second depth values if the depth value of the annotation image fragment is between the first and second depth values of the rendered image fragment. For example, the scaled opacity value may be determined based on the ratio of the depth of the annotation image fragment to the corresponding first and / or second depth values. The adjustment may further include determining full transparency for the annotation fragment if the depth value of the annotation fragment is greater than a second depth value corresponding to the dataset image fragment.
[0054] The imaging dataset can be a medical volumetric imaging dataset, and / or the landmarks can be anatomical landmarks.
[0055] The layout of annotation labels may include one or more of the following: the position of the label relative to the dataset image, the depth of the label relative to the dataset image, the size of the label, the color of the label, the text / content displayed in the label, and the distance of the label to the corresponding landmark in the dataset image. It should be understood that any spatial characteristics of the labels, and / or any spatial relationships between labels and / or between labels and landmarks, can be defined in the layout.
[0056] A computing device is provided that includes at least one processor and a memory, the memory including instructions executable by the processor, wherein when the instructions are executed in the processor, the computing device is configured to perform actions according to any method or combination of methods of this disclosure.
[0057] A computer program or computer program product and a non-transitory computer-readable storage medium including program code are provided. The program code can be executed by at least one processor. When the program code is executed, the at least one processor performs any method or combination of methods according to this disclosure.
[0058] It should be understood that, without departing from the scope of this disclosure, the features mentioned above and those to be explained below can be used not only in the indicated corresponding combinations, but also in other combinations or individually. Attached Figure Description
[0059] Figure 1 The illustrations depict 3D anatomical visualizations of patient data using photorealistic volumetric path tracing (cinematic rendering) based on various examples.
[0060] Figure 2The illustration shows annotation overlays applied to photorealistic renderers for 3D medical images, based on various examples.
[0061] Figure 3-5 The illustration shows frames from translation animations based on various examples, where annotation labels gradually disappear as the corresponding 3D position moves out of the rendering viewport.
[0062] Figure 6-8 The illustration shows the distance from the landmark label to the visible anatomical structure based on various examples, as well as the effects of the landmark label fading in and out.
[0063] Figure 9 The illustrations schematically depict the actions of methods for optimizing the layout of labels in a rendered image, based on various examples.
[0064] Figure 10 The illustrations schematically depict computing devices configured to perform methods according to the present disclosure, based on various examples. Detailed Implementation
[0065] Some examples of this disclosure generally provide multiple circuits or other electrical devices. All references to circuits and other electrical devices, and to the functionality provided for each, are not intended to limit coverage to what is illustrated and described herein. While specific labels may be assigned to the various circuits or other electrical devices disclosed, such labels are not intended to limit the scope of operation of the circuits and other electrical devices. Such circuits and other electrical devices may be combined and / or separated from each other in any way based on a desired particular type of electrical implementation. It should be understood that any circuit or other electrical device disclosed herein may include any number of microcontrollers, general-purpose processor units (CPUs), graphics processing units (GPUs), integrated circuits, memory devices (e.g., FLASH, random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other suitable variations thereof), and software that cooperates with each other to perform one or more of the operations disclosed herein. Furthermore, any one or more of the electrical devices may be configured to execute program code embodied in a non-transitory computer-readable medium programmed to perform any number of the functions disclosed.
[0066] In the following, examples of this disclosure will be described in detail with reference to the accompanying drawings. It should be understood that the following description of the examples should not be construed as limiting. The scope of this disclosure is not intended to be limited by the examples or drawings described below, which are to be understood as illustrative only.
[0067] The accompanying drawings are to be considered schematic representations, and the elements illustrated in the drawings are not necessarily shown to scale. Rather, the various elements are shown such that their function and general purpose will be obvious to those skilled in the art. Any connection or coupling between functional blocks, devices, components, or other physical or functional units shown in the drawings or described herein may also be achieved through indirect connection or coupling. Coupling between components may also be established through wireless connections. Functional blocks may be implemented in hardware, firmware, software, or a combination thereof.
[0068] The following section describes a technique for the optimized automatic placement of labels associated with landmarks identified in an image.
[0069] Annotation tags can be used to enhance images with textual semantic information, that is, to display semantic information and other image features to the user, thereby explaining the visible features in the image to the viewer. For example, information about anatomical structures, organs, pathology, etc., can be used to enhance medical images, thereby facilitating radiologists' review of medical images. Annotation tags can be overlaid on images according to the annotation layout. Sometimes, a series of images needs to be annotated; the image sequence can form a movie. In such scenarios, the temporal dependence of the annotation layout is also determined.
[0070] Creating globally optimal and temporally coherent annotation layouts in images without human intervention is a common challenge. The disclosed technique addresses the display of text labels for body landmarks and organ contours derived from AI-based automated segmentation algorithms for both educational and clinical workflows. The layout utilizes iterative optimization based on heuristic rules and simulated annealing running in real-time for typical data sizes, while also taking into account the visibility of anatomical structures in 3D volumetric rendering.
[0071] Specifically, automatic layout optimization of annotation labels is performed based on the visibility of anatomical structures, a temporal coherence heuristic is used to support automatic animation and interactive applications, and an improved approximate depth embedding is provided by using opaque geometry on top of annotation overlays from dual-depth images.
[0072] Various examples of these methods are applicable to applications involving the rendering of landmarks and segmented structures, such as cinematic anatomy for educational workflows and the Digital OncologyCompanion (DOC) extension in AI-Rad Companion for clinical workflows. These techniques are also applicable to virtual reality and augmented reality applications involving 3D rendering and patient anatomical images. A subset of these techniques is further broadly applicable to non-medical and non-volumerative rendering applications, such as CAD / CAM.
[0073] Automatic label placement is an NP-hard problem, and in general, the optimal solution cannot be computed in a reasonable amount of time. Although several algorithms exist for computing label layouts with desirable properties, this paper focuses specifically on a fast local optimization algorithm that can be implemented in real time for typical problem sizes in single-case medical visualization.
[0074] Traditionally, genetic algorithms for automatically labeled mapping data are known, where layout optimization relies on heuristics inspired by the concept of natural selection in Darwinian evolution. In a simple example, the fitness function to be maximized is defined as tracking the number of non-overlapping labels, which drives the selection, crossover, and mutation process of a pool of possible solutions. Simulated annealing is a variant of local optimization, where optimization steps that lead to a worse outcome for the objective cost function can maintain a probability proportional to the change in the cost function. This technique approximates the global optimum by forcing large random changes at the start of the iterative optimization process with the goal of avoiding local maxima. In both traditional 2D and immersive 3D visual environments, the techniques disclosed herein contribute rules and heuristics related to layout optimization based on visible anatomical features in 3D medical volumetric rendering.
[0075] Traditional volumetric visualization methods based on ray projection, also known as ray tracing simulations, are still used in many current state-of-the-art medical visualization products. They simulate only the emission and absorption of radiant energy passing through volumetric data along the main line of sight. Using absorption coefficients derived from patient data, the radiant energy emitted at each point along the ray to the observer's position is absorbed according to the Beer-Lambert law. Based on local volumetric gradients (local illumination), renderers typically use only standard local shading models (e.g., the Blinn-Phong model) to calculate shading. While fast, these methods do not simulate the complex light scattering and extinction associated with photorealism (global illumination).
[0076] Cinematic rendering is another type of ray tracing simulation that implements physically based Monte Carlo light transport, which uses stochastic processes to simulate light paths through volumetric data, each path having multiple scattering events. As more and more paths are simulated, the solution converges to an accurate estimate of the irradiance of incident light from all directions at each point. Based on properties derived from anatomical data, the renderer employs a hybrid of volumetric scattering and surface-like scattering, modeled separately by phase functions and BRDFs, as described, for example, in Kroes, Thomas, Frits H. Post, and Charl P. Botha. "Exposure render: An interactive photo-realistic volume rendering framework." PLOS ONE 7.7 (2012): e38586. Medical data is illuminated using image-based lighting via a high dynamic range light probe, which can be captured by a 360-degree camera to approximate lighting conditions in a real-world setting. Such lighting results in a more natural appearance of the data compared to images created using synthetic light sources typically applied to direct volumetric rendering methods. When combined with accurate simulations of photon scattering and absorption, the renderer produces photorealistic images that include many shading effects observed in nature, such as soft shadows, ambient occlusion, volumetric scattering, and subsurface photon interactions.
[0077] Figure 1 The illustrations depict 3D anatomical visualizations of patient data using photorealistic volumetric path tracing (cinematic rendering) based on various examples.
[0078] As in Figure 1 As can be seen, the 3D anatomical visualization of patient data using photorealistic volumetric path tracing (cinematic rendering) is shown as an unlabeled 3D image. Here, a plane is chosen to display the cross-sectional image based on the medical volumetric dataset, where areas with low opacity are shown as transparent, i.e., the chest is shown as partially hollow. The lungs are shown as blank spaces, a result of the voxel optical opacity defined by voxel classification.
[0079] Figure 2 The illustration shows an annotation overlay of multiple labels 1 according to various examples, which is applied to a photorealistic renderer of 3D medical images 3 based on Monte Carlo volume path tracing.
[0080] As in Figure 2As can be seen, the segmented organ rendering image 3 directly visualizes the mask voxels associated with multiple defined regions of interest 4 (i.e., organs) and associated landmarks 2, and demonstrates the deep embedding of the overlay image relative to the pure volumetric rendering image 3. This overlay image includes labels 1, i.e., landmark glyphs, connecting lines, and text boxes. According to this disclosure, the labels 2 of the segmented organs are displayed based on automatic label generation, layout optimization, and overlay rendering.
[0081] Here, the visibility of organs and corresponding landmarks, as well as lines and text boxes, has been considered according to the criteria described above. Specifically, the proximity of spatial relationships (e.g., distance between two elements or overlap between two elements) between different labels and their corresponding organs and / or landmarks (i.e., specific representative parts of organs in the image) has been optimized, which may be part of the objective function. Labels are arranged such that they do not overlap with each other or with organs, and the distances between labels and regions of interest are regular and within a predetermined threshold of average distance. In this example, an occlusion rating has been minimized, which may be part of the objective function and may describe the overlap or occlusion of labels and / or regions of interest and / or landmarks.
[0082] The techniques disclosed herein improve the accuracy of depth embedding for highly opaque surfaces by generating a first and a second depth map based on first and second cumulative opacity thresholds and simulating volumetric extinction under the assumption of a homogeneous participating medium. While approximating, this method allows for decoupling of annotation-overlay rendering from volumetric rendering algorithms. Even in path-tracing systems, the generation of the second depth map can be performed by separate volumetric ray projection independent of the first depth map, because for large optically transparent areas of the volumetric data, the Monte Carlo solver can converge very slowly to the correct distance range along the main line of sight. The techniques described herein are broadly applicable to a range of rendering algorithms, including systems that generate only a single depth layer from volumetric data, or systems employing alternative methods for embedding surfaces within volumetric data (e.g., depth-based stripping).
[0083] The following description uses examples from medical volumetric rendering to illustrate steps for optimizing label layout to display semantic annotations associated with multiple landmarks in a rendered image. However, it should be understood that the techniques presented can generally be adapted to any image rendering process as needed.
[0084] In one action, an image is acquired. In various examples, frames are received from or rendered from a volumetric dataset as 3D volumetric medical images or from 4D medical images.
[0085] In another action, based on the image, automatic annotation and organ contour drawing can be performed. In various examples, landmarks can be detected in the image, such as the top or bottom or center of the kidney, the tip of the rib at T12, and similar landmarks.
[0086] Organ landmarks, such as the centroids of organ voxels and / or organ mesh outlines, can be generated from contour drawing data. Alternatively, annotation and organ contour drawing can be performed manually.
[0087] In various examples, images can be obtained independently of annotation labels and / or their layout. In other words, two separate images can be obtained and / or determined: one image is a rendered image based on an imaging dataset, and the other image can be an overlay image containing annotation labels according to the label layout, to be overlaid on the rendered image to display the labels in the image. Therefore, the rendered image can include multiple rendered segments at associated locations, and the overlay image can include multiple overlay segments at associated locations, where the pixel positions of the rendered and overlay segments are identical, allowing them to be combined by overlaying the overlay segments onto the rendered segments, optionally adjusting the opacity values of the overlay segments, for example, based on attributes of the rendered segments (e.g., color, opacity, depth). This result can be referred to as the result image, incorporating features of both the rendered and overlay images.
[0088] In another action, an annotation layout can be selected for labels to be displayed in the image. Labels can represent and display semantic annotation information associated with landmarks. Annotation labels can consist of 3D glyphs, 2D or 3D text boxes at the 3D landmark locations, and connecting elements. In various examples, the label layout can be in the primary 3D space, for example, for AR / VR applications, where labels can be associated with multiple different depth values. Labels can also be in a 2D layout on a focal plane with 3D landmarks and connecting elements visualized, where labels can be associated with a common depth value from the camera viewpoint. In various embodiments, one label, multiple labels, or all labels can lie in a single plane, i.e., have a common depth from the camera viewpoint. It is also possible that a group or several groups of labels can lie in different common planes. Alternatively, a convergence plane can be used for stereoscopic display applications.
[0089] In one action, during volumetric rendering, post-classified volumetric data statistics for each line of sight in the ray projection channel can be determined. In various examples, a predefined first opacity threshold (e.g., a cumulative opacity threshold of 0.35) can be used to generate a representative depth D1. If the threshold is not reached, the center of a depth range containing non-zero optical opacity can be used instead.
[0090] A second predefined opacity threshold (e.g., a cumulative opacity of 0.999) can be used to generate a fully opaque saturation depth D2. If the light rays leave the volumetric data boundary before the cumulative opacity reaches the threshold, the depth value can be inferred based on the cumulative opacity at the exit point and the distance to the first opacity sample, assuming a homogeneous participating medium. Photorealistic rendering based on Monte Carlo volumetric path tracing may require a separate ray casting pass to collect ray data.
[0091] In another action, the initial placement of the labels can be performed, where the initial positions of the labels in the rendered image are determined.
[0092] In various examples, an occlusion factor can be determined for each annotation based on the 3D landmark position. The occlusion factor could be, for example, an occlusion rating indicating how much of a label / region of interest is occluded by another label. Initial placement can be based on rules, such as landmarks always being visible. Alternatively, the occlusion factor can be determined by the depth (i.e., camera space depth) of the 3D landmarks D1 and D2. In various examples, the annotation can fade in 5mm in front of D1 and 5mm behind D2. As a result, the label and landmark fade in and out as the clipping plane moves through the data, or as the dynamic landmark moves away from visible structures. The occlusion factor can also be determined by the distance to the viewport; for example, the landmark fades away near the edge of the screen; alternatively, the landmark fades away within a short distance from the viewport frame, such that the connecting lines and labels remain partially visible as the landmark moves off-screen (e.g., during camera panning).
[0093] The position of 3D landmarks can be calculated in the camera's viewport.
[0094] An initial plane in 3D space can be selected for rendering annotation labels. A view-aligned annotation plane can be selected, along with a depth that minimizes the distance to the landmark. Alternatively, principal component analysis (PCA) can be performed on the 3D landmark location, where a plane defined by the principal and secondary axes can be used as the layout plane. Therefore, the initial layout can include the depth of the plane in which the label will be displayed; in this case, the initial layout defines a common depth value for both the plane and the label.
[0095] The initial label layout can be calculated based on various constraints, including the label positions. In various examples, labels can be grouped, for example, left / right, or divided into screen quadrants. Alternatively, constraints can be omitted, and the initial positions of the text boxes can be landmark positions in viewport space. A fast greedy algorithm or a simple heuristic can be used, such as placing each consecutive text box in the first available non-overlapping space.
[0096] In another action, iterative optimization of label placement in 2D or 3D layouts can be performed. Optimization can be tailored to different final layouts, such as clustered or circular layouts.
[0097] According to this disclosure, multiple steps can be used to optimize label layout based on the visibility of labels and landmarks in an image.
[0098] In other words, the input to the optimization algorithm can be labels, a label layout, and an image including landmarks. One or more heuristic rules can be used to modify the label layout. When the rendered and overlaid images are displayed together, the heuristic rules can be based on the visibility of labels and / or landmarks in the resulting image. An objective function can be used to evaluate the modified layout, applied to the labels in the provided layout in the image, where the objective function provides its output value based on the visibility of labels and / or landmarks, and where the objective function evaluates the layout based on multiple criteria. The objective function and / or criteria can be based on the visibility of images of interest and / or landmarks and / or image structures in the resulting image. The modified layout can be provided as the output of the optimization algorithm.
[0099] The layout can be evaluated and optimized using an objective function based on multiple criteria, which may be based on or take into account the visibility of labels and / or landmarks. It should be understood that the criteria presented are merely examples, and many other criteria, such as those based on the visibility of labels and / or landmarks, can be used and combined in the objective function.
[0100] In various examples, during optimization, for each label, one or more of the following heuristics (i.e., instructions for rearranging labels) can be applied to change the layout. It should be understood that the heuristics presented are merely examples, and many other rules, such as those based on the visibility of labels and / or landmarks, can be used.
[0101]
[0102]
[0103] When applying heuristic rules, the action selection for each iteration can be guided by the weights assigned by the user and the application for each criterion.
[0104] The objective function can define an objective metric for layout quality, and gradient descent or other numerical optimization methods can be applied to optimize the layout. The quality metric can be based on user ratings. Machine learning methods can also be used.
[0105] The optimization process can be implemented using genetic algorithms. Simulated annealing can be used to avoid the optimization from converging to a globally suboptimal but locally optimal solution.
[0106] Some embodiments may move the focal plane of the 2D label layout to minimize the overlap of connecting lines.
[0107] From the dataset, multiple images corresponding to different planes can be rendered, for example, which can be displayed in animations, or users can interactively change to different planes / images while viewing the dataset. Temporal constraints can be applied to the tag-based mobile application to ensure temporal coherence during such animations and / or during interactive rendering in AR / VR applications with head tracking.
[0108] Implementations of this system can instead perform independent optimizations for each animation keyframe and explicitly interpolate layout parameters.
[0109] Using layout quality metrics, the optimization phase can calculate the intermediate transition layout that maximizes the minimum layout quality during animation.
[0110] In another action, an annotation rendering pass can be executed.
[0111] In various examples, annotation rendering can be performed on individual overlay images. In some examples, depth images D1 and D2 are used, and the overlay images can be displayed and depth-merged independently of the volume rendering process.
[0112] When using D1 and D2, for each fragment in the overlay rendering, a normalized distance within the range [D1, D2] can be calculated. Then, assuming extinction from a homogeneous participating medium, the fragment opacity can be scaled based on the normalized distance, where the opacity scaling is 1.0 at distance D1 and 0.0 at distance D2.
[0113] The following examples will explain how the disclosed technique is implemented. Body landmark detection and organ contour rendering are performed in the preprocessing steps. Additionally, smooth surface meshes are generated for the detected organs, along with labels for each organ, anchored at the organ voxel closest to the centroid. In these examples, images are rendered from a volumetric dataset, where semantic information includes the names of specific parts of human organs.
[0114] Figure 3-5 The illustration shows frames from translational animations according to various examples, where label 1 (i.e., the marker glyph, its text box, and connecting line) is displayed in the corresponding image plane involving multiple regions of interest 4 and corresponding markers 2, and label 1 gradually disappears as the corresponding 3D position leaves the rendering viewport. Figure 3-5The image shows multiple images, or frames, from a translation animation (head to toe), illustrating visibility processing based on organ labels 1 contained within the viewport. As the camera moves, the camera viewport moves (the viewport is a virtual representation displayed in the 3D scene). As landmarks enter and leave the visible viewport, their opacity changes, and landmarks with zero opacity are not included in the layout optimization. This results in a less cluttered visualization when new landmarks 2 enter the view or organ mesh and their labels 1 fade in during animation. For example, a region of interest can be associated with multiple landmarks relating to different parts of the region of interest (e.g., the top / center / bottom of an organ).
[0115] Therefore, a movie sequence of multiple rendered images can be provided, where the objective function is based on the relative changes in layout between subsequent rendered images in the movie sequence. Here, the label layout can be dynamically adjusted based on user input, where the objective function is based on multiple criteria, and where the weighting factors associated with the multiple criteria are user-adjustable between subsequent iterations of the iterative optimization. When determining the label layout for each image in the movie sequence, a temporal coherence criterion can be determined and included in the objective function. This criterion describes, for example, the temporal coherence of the layout position, such that changes in layout exceeding a predetermined threshold within a predetermined time period and / or number of frames can be penalized in the objective function.
[0116] Figure 6-8 Further illustrations show the fade-in and fade-out effects of marker 2 based on the distances from marker 2 to visible anatomical structures 4 and their marker 2, according to various examples. Figure 6-8 In the video, image 3, frame 3, is shown from the clip animation (back to front), illustrating the visibility processing of landmark label 1 based on distance to visible anatomical structures. The video contains a single clip plane aligned with the anterior patient orientation and moving from the back of the body towards the front. Landmark 2 has a fading distance of 10 mm. The same set of rules applies. Figure 2-5 While layout optimization is possible, other implementations can introduce heuristics to indicate, for example, invisible landmarks 2, or apply only a subset of heuristics.
[0117] In the described example, the layout is a three-dimensional layout that defines the position of the labels in a reference coordinate system to which the volumetric medical imaging dataset is registered. The method further includes rendering a corresponding overlay image depicting the label based on the label layout and for each of at least one rendered image, wherein the rendered image is merged with the corresponding overlay image based on a plurality of opacity values determined for a segment of the corresponding overlay image.
[0118] The label layout in the described example can be determined using an artificial neural network algorithm that operates based on one or more inputs, which are based on semantic information, at least one rendered image, and the positions of multiple regions of interest.
[0119] Generally, considering the visibility of labels and regions of interest (ROIs) may include determining the overlap and / or proximity of a label or a portion of a label to a corresponding landmark or ROI. Considering visibility may include applying an objective function to the layout, wherein the objective function determines a measure of the visibility of one or more of the labels, any portions of the labels, ROIs, and landmarks, wherein the measure may include an occlusion rating (e.g., percentage overlap) for each of the labels and / or ROIs or landmarks, or one or more of the proximity relationships between each of the labels and / or ROIs or landmarks.
[0120] Figure 9 The illustrations schematically depict the actions of methods for optimizing the layout of labels in a rendered image, based on various examples.
[0121] The method begins with action S10. In action S20, at least one rendered image is obtained. In action S30, the locations of multiple regions of interest (ROIs) in the at least one rendered image are determined. In action S40, semantic information associated with the multiple ROIs is obtained. In action S50, based on the semantic information and the locations of the multiple ROIs, and considering the visibility of the labels and the further visibility of the ROIs, a layout of labels for annotating the multiple ROIs in the at least one rendered image is determined. The method ends in action S60.
[0122] Figure 10 A computing device configured to perform any method according to the present disclosure is schematically illustrated according to various examples. The computing device 100 includes at least one processor 110 and a memory 120, the memory 120 including instructions executable by the processor 110, wherein when the instructions are executed in the processor 110, the computing device 100 is configured to perform actions according to any method or combination of methods according to the present disclosure.
[0123] An improved method for optimizing label layout to display annotation information in an image is described, wherein the image is rendered independently of the labels, and in the merged result image, the labels are displayed together with the image using an overlay image containing the labels. The opacity of the overlay image can be adjusted. The label layout optimization process, particularly including the position of the labels in the overlay / result image, is implemented based on several heuristics describing how the layout can be modified and an objective function that evaluates the layout based on the visibility of labels and landmarks in the result image. In various examples, the opacity of the overlay image can be adjusted by using first and second depth maps used for rendering the image, which can be determined during rendering or in a separate ray tracing simulation based on a ray tracing simulation (e.g., a volumetric ray casting algorithm) used for rendering the image. The depth image can be calculated as part of the volumetric ray casting algorithm by recording the locations along the line of sight where the accumulated optical opacity reaches certain thresholds.
[0124] Although the invention has been shown and described with respect to certain preferred embodiments, equivalents and modifications will occur to those skilled in the art upon reading and understanding the specification. The invention includes all such equivalents and modifications and is limited only by the scope of the appended claims.
[0125] For illustration purposes, various scenarios have been presented above using rendered images from volumetric medical imaging datasets. Similar techniques can be easily applied to other types and kinds of rendered images, such as rendered images from volumetric monitoring images, etc.
Claims
1. A method for optimizing the layout of labels to annotate multiple regions of interest in at least one rendered image of a volumetric medical imaging dataset with semantic information, the method comprising: Obtain the at least one rendered image; Determine the positions of the plurality of regions of interest in the at least one rendered image; Obtain semantic information associated with the multiple regions of interest; and Based on semantic information and the locations of the multiple regions of interest (ROIs), and considering the visibility of the labels and the further visibility of the multiple ROIs, the layout of the labels used to annotate the multiple ROIs in the at least one rendered image is determined. The layout described therein is a three-dimensional layout that defines the position of the labels in a reference coordinate system, and the volumetric medical imaging dataset is registered to the reference coordinate system, wherein the method further includes: Based on the layout of the labels, and for each of the at least one rendered image, render the corresponding overlay image depicting the label.
2. The method of claim 1, further comprising determining the positions of a plurality of markers corresponding to regions of interest in the at least one rendered image, wherein the markers represent the positions of the corresponding regions of interest, and further considering the visibility of the markers for determining the layout of the labels.
3. The method of claim 1, wherein determining the layout of the label includes determining an initial layout when the label is displayed in an image and iteratively optimizing the initial layout based on an objective function, said objective function being based on the visibility of the label and the further visibility of the region of interest.
4. The method of claim 3, wherein the objective function is further based on the proximity relationship between at least one of the labels and the corresponding landmarks relative to each other.
5. The method of claim 4, wherein the objective function is further based on an occlusion rating associated with a label, the occlusion rating being based on the position of the label in the rendered image.
6. The method of claim 5, wherein the occlusion rating is based on the overlap between one of the tags and another of the tags, the region of interest, and the markers of the region of interest.
7. The method of claim 3, wherein the at least one rendered image comprises a movie sequence of a plurality of rendered images including the at least one rendered image, wherein the objective function is based on the relative change in layout between subsequent rendered images of the movie sequence of the plurality of rendered images.
8. The method of claim 3, wherein the objective function is based on a plurality of criteria, wherein the weighting factors associated with the plurality of criteria are user-adjustable between subsequent iterations of the iterative optimization.
9. The method according to claim 1, The layout of the labels is determined using an artificial neural network algorithm, which operates based on one or more inputs, namely semantic information, the at least one rendered image, and the positions of the plurality of regions of interest.
10. The method of claim 1, further comprising: Based on multiple depth positions determined for segments of the corresponding overlay image, each of the at least one rendered image is merged with the corresponding overlay image.
11. The method of claim 10, wherein the depth position is defined based on opacity values determined for multiple positions along a ray used for ray-traced rendering of the at least one rendered image based on a volumetric medical imaging dataset.
12. The method according to claim 11, wherein, Based on rendering images from the dataset: A first depth map is generated comprising a plurality of first depth values of the at least one rendered image, wherein each first depth value of the first depth map is associated with a corresponding segment of the at least one rendered image, wherein the first depth value corresponds to a depth position at which the cumulative opacity value of rays in a volumetric ray casting algorithm for rendering the at least one rendered image associated with the segment reaches a first opacity threshold at the camera viewpoint. The depth values of segments in the overlay image are determined based on the layout, and these depth values correspond to segments of the rendered image in the overlay. The opacity of the segments in the overlay image is adjusted based on the depth values of the segments in the overlay image and the first depth map.
13. The method of claim 12, wherein a second depth map of the at least one rendered image is generated, wherein each second depth value of the second depth map is associated with a corresponding segment of the at least one rendered image, wherein the second depth value corresponds to a depth location at which the cumulative opacity value of a ray in a ray tracing simulation associated with the segment in the at least one rendered image reaches a second opacity threshold, wherein the second opacity threshold is higher than a first opacity threshold, and The opacity value of the segments in the overlay image is adjusted based on the depth of the segments, the first depth map, and the second depth map.
14. The method of claim 13, wherein adjusting the opacity value of a segment in the overlay image comprises: When the depth value of a fragment is less than the first depth value of the corresponding fragment in the at least one rendered image, the full opacity value of the fragment in the overlay image is determined. When the depth value of a segment in the overlay image is between the first and second depth values of the corresponding segment in the at least one rendered image, a scaling opacity value between a full opacity value and a full transparency value is determined based on a comparison between the depth value of the segment in the overlay image and the corresponding first and second depth values of the corresponding segment in the at least one rendered image; or When the depth value of a fragment in the overlay image is greater than the second depth value of the corresponding fragment in the at least one rendered image, the full transparency value of the fragment in the overlay image is determined.
15. The method of claim 1, wherein the region of interest is an anatomical structure displayed in the at least one rendered image.
16. The method of claim 1, wherein the label comprises one or more of 2D or 3D glyphs, 2D or 3D text boxes, surfaces, and connecting lines, the connecting lines being anchored at corresponding landmarks.
17. The method of claim 1, wherein the layout of the label includes one or more of the following: the position of the label in the at least one rendered image, the depth position of the label, the size of the label, the color of the label, and the text or content displayed in the label.
18. A system comprising: A computer includes at least one processor and a memory, the memory including instructions executable by the processor, wherein when the instructions are executed in the processor, the computing device is configured to: Obtain the rendered image; Determine the location of the region of interest in the rendered image; Obtain semantic information associated with the region of interest; and Based on semantic information and the location of regions of interest (ROIs), and considering the visibility of labels and further visibility of ROIs, the layout of labels used to annotate ROIs in the rendered image is determined. The layout described therein is a three-dimensional layout that defines the position of the labels in a reference coordinate system, and the volumetric medical imaging dataset is registered to the reference coordinate system. Based on the layout of the labels, and for each of the at least one rendered image, render the corresponding overlay image depicting the label.