Automatic selection and display layout of medical images from clinical descriptions
By receiving clinical descriptions and viewing preferences input by users, using clinical knowledge ontology database and machine learning model, the display layout of medical images is automatically selected and generated, which solves the problem that clinicians have difficulty configuring display layout and achieves efficient medical image reading.
Patent Information
- Application Number
- CN202411634205.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-11-15
- Publication Date
- 2025-05-20
AI Technical Summary
It is difficult for clinicians to configure the display layout of medical images, which requires professional technical knowledge, and the prior art cannot automatically select and display medical images based on clinical descriptions.
By receiving descriptions and viewing preferences of desired medical images input by users, using clinical knowledge ontology databases and machine learning models, the corresponding display layout is automatically matched and generated.
It realizes the display layout of automatically selecting and generating medical images from clinical descriptions, simplifying the operation of clinicians and improving the efficient reading ability of medical images.
Smart Images

Figure CN120020971A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the automatic selection and display layout of medical images, and particularly to the automatic selection and display layout of medical images from clinical descriptions. Background Art
[0002] In radiology, medical images of patients are acquired by radiologists for clinical analysis. To facilitate efficient reading of such medical images, radiologists typically prefer a consistent and optimal display layout of such medical images. Currently, the display layout of medical images is defined by configuring data properties and view settings, which requires technical expertise in specific vendor software implementations. However, clinicians typically communicate in clinical terms and do not have the technical knowledge required to configure the display layout of medical images. Summary of the Invention
[0003] According to one or more embodiments, a system and method for automatic selection and display layout of medical images are provided. A user input is received, the user input including 1) a description of a desired medical image and 2) a viewing preference for the desired medical image. One or more nodes of a clinical knowledge ontology database that match the description of the desired medical image are determined. The one or more matching nodes are associated with one or more medical images in the clinical knowledge ontology database. A display layout for the one or more medical images is generated based on the viewing preference. The display layout is output.
[0004] In one embodiment, the user input further includes a temporal description of the desired medical image. One or more nodes of the clinical knowledge ontology database that match the description of the desired medical image are determined based on the temporal description of the desired medical image.
[0005] In one embodiment, the user input is parsed into a vector. A vector search is performed between the vector representing the description of the desired medical image and the vectors representing the nodes of the clinical knowledge ontology database to identify a list of ranked nodes. One or more of the highest ranked nodes that meet a ranking threshold are identified as the one or more nodes.
[0006] In one embodiment, a clinical knowledge ontology database is generated by the steps of: receiving one or more medical images, extracting features from the one or more medical images using one or more machine learning models, associating the extracted features with corresponding nodes of the clinical knowledge ontology database, and outputting a clinical knowledge ontology database having the extracted features associated with the corresponding nodes.
[0007] In one embodiment, the description of the desired medical image includes a description of the anatomical object of interest to navigate to within one or more medical images, and the one or more matching nodes include one or more matching nodes associated with the coordinates of the anatomical object of interest in the one or more medical images.
[0008] In one embodiment, the description of the desired medical image is defined based on at least one of imaging modality, acquisition parameters, acquisition orientation, image appearance, anatomical field of view, or classification and detection. In one embodiment, the viewing preference is defined based on at least one of anatomical display orientation, windowing, or rendering mode.
[0009] In one embodiment, a language model is used to translate the viewing preference into rendering parameters.
[0010] In one embodiment, the display layout is displayed on a display device.
[0011] These and other advantages of the present invention will be apparent to those of ordinary skill in the art by reference to the following detailed description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A method for automatically generating a display layout for one or more medical images is shown in accordance with one or more embodiments;
[0013] Figure 2 A workflow for automatically generating a display layout for one or more medical images is shown in accordance with one or more embodiments;
[0014] Figure 3 A method for generating a clinical knowledge ontology database is shown in accordance with one or more embodiments;
[0015] Figure 4 A workflow for automatic medical image navigation is shown in accordance with one or more embodiments;
[0016] Figure 5 An exemplary artificial neural network that can be used to implement one or more embodiments is shown;
[0017] Figure 6 A convolutional neural network that can be used to implement one or more embodiments is shown;
[0018] Figure 7 A schematic structure of a recurrent machine learning model that can be used to implement one or more embodiments is shown; and
[0019] Figure 8 A high-level block diagram of a computer that can be used to implement one or more embodiments is shown. Detailed implementation mode
[0020] The present invention generally relates to methods and systems for the automatic selection and display layout of medical images from clinical descriptions. Embodiments of the present invention are described herein to give a visual understanding of such methods and systems. Digital images often consist of digital representations of one or more objects (or shapes). Herein, the digital representation of an object is often described in terms of identifying and manipulating the object. Such manipulation is virtual manipulation performed in the memory of a computer system or other circuitry / hardware. Thus, it should be understood that embodiments of the present invention may be executed within a computer system using data stored within the computer system. Further, the pixels of an image referred to herein may equivalently refer to the voxels of the image, and vice versa. Embodiments of the present invention are described with reference to the figures, where like reference numerals refer to the same or similar elements.
[0021] The embodiments described herein provide a display layout system for the automatic selection and display layout of medical images from any natural language clinical statement received as user input from a user (e.g., a radiologist). The user input includes a description of the desired medical image and the viewing preferences for the desired medical image. The description of the desired medical image is matched with one or more nodes of a clinical knowledge ontology database of clinical terms, where the one or more nodes are associated a priori with one or more medical images. A display layout for one or more medical images is automatically generated based on the viewing preferences. Advantageously, the display layout system provides a robust transformation of user input of instructions in clinical terms into closely related medical images. The use of a language model also enables the transformation of partially described viewing preferences into internal viewing implementations.
[0022] Figure 1 A method 100 for automatically generating a display layout for one or more medical images according to one or more embodiments is shown. The steps and sub-steps of method 100 may be performed by one or more suitable computing devices, such as Figure 8 computer 802. Figure 2 A workflow 200 for automatically generating a display layout for one or more medical images according to one or more embodiments is shown. Figure 1 And Figure 2 will be described together.
[0023] In Figure 1At step 102, receive user input, where the user input includes 1) a description of the desired medical image and 2) viewing preferences for the desired medical image. The user input can be received from a radiologist, a clinician, or any other user. The user input can be defined in clinical terms in any, unstructured format. The user input can be received as text, speech, or in any other suitable form. An example of user input is: "Display of T2 with the most recent previous date, in the anterior orientation, in a lung window". In one example, as shown in Figure 2 workflow 200 of
[0024] the user input is a text / speech command 202. For example, the description of the desired medical image can be defined based on at least one of the imaging modality, acquisition parameters, acquisition orientation, image appearance, anatomical field of view, classification, and detection (e.g., the presence of an anatomical object of interest, such as an organ, blood vessel, bone, tumor, pathology, etc.), and / or any other suitable description of the desired medical image. For example, the viewing preferences for the desired medical image can be defined based on at least one of the anatomical display orientation, windowing, rendering mode, and / or any other suitable viewing preference.
[0025] In one embodiment, the user input received at Figure 2 step 102 further includes a temporal description of the desired medical image, such as a relative description of a date or time order.
[0026] For example, by loading the user input from a storage device or memory of a computer system (e.g., Figure 8 memory 810 or storage device 812 of computer 802 of Figure 8 ), or by receiving the user input from a remote computer system (e.g., Figure 8 computer 802 of
[0027] ), the user input can be received from a user interacting with an I / O (input / output) device of the computer system (e.g., Figure 1 I / O device 808 of computer 802 of Figure 2 workflow 200 of
[0028] To identify one or more nodes that match the description of a desired medical image, as shown in the workflow 200 of Figure 2 the text / speech command 202 input by the user is first interpreted by a language model parser 204 to extract 1) a description of the desired medical image, 2) viewing preferences, and 3) (optionally) a temporal description of the desired medical image as corresponding vector representations. The language model parser 204 can be implemented according to any suitable (e.g., well-known) method. The vector representation of the description of the desired medical image extracted from the text / speech command 202 is then input into a large clinical knowledge AI embedding database 206 to search for one or more nodes having vector representations that match the vector representation of the description of the desired medical image. Any suitable vector similarity algorithm can be used to perform the search, such as, for example, cosine similarity. The result of the vector search is a list of ranked nodes within the large clinical knowledge AI embedding database 206 that are similar to the description of the desired medical image. The highest matching node (or top N nodes, where N is any positive integer) that also meets a ranking threshold is selected as the one or more matching nodes, and one or more medical images associated with the one or more matching nodes are identified.
[0029] Optionally, in one embodiment, in the case where a temporal description of the desired medical image is received at step 102 of Figure 1 the search can be further constrained to nodes within the large clinical knowledge AI embedding database 206 that satisfy the temporal description (e.g., date range). In one embodiment, the search can be constrained according to the temporal description by prompting the language model 210 with a list of available image acquisition dates and asking the language model 210 to identify the nodes that satisfy the temporal description. Then, a vector search can be performed on the identified nodes that satisfy the temporal description.
[0030] The language model 210 can be any suitable language model. In one embodiment, the language model 210 can be a machine learning-based language model. For example, in one embodiment, the language model 210 can be a customized, relatively small language model for natural language processing, such as, for example, BERT (Bidirectional Encoder Representations from Transformers). In another embodiment, the language model 210 can be pre-trained deep learning based on an LLM (Large Language Model). For example, the LLM can be based on the transformer architecture, which uses self-attention mechanisms to capture long-term dependencies in the text. An example of a transformer-based architecture is GPT (Generative Pretrained Transformer), which has a multi-layer transformer decoder architecture that can be pre-trained to optimize the next token prediction task and then fine-tuned using labeled data for various downstream tasks. The GPT-based LLM can be trained using reinforcement learning with human feedback for performing various natural language processing tasks.
[0031] The large clinical knowledge AI embedding database 206 is generated a priori to model clinical terms of medical images as a graph including multiple nodes connected by edges. Each node is represented as a vector representation associated with a clinical term of a medical image. It can be generated according to Figure 3 the large clinical knowledge AI embedding database 206, which is described in detail below.
[0032] At Figure 1 step 106 of, a display layout of one or more medical images is generated based on viewing preferences. In one example, as shown in the workflow 200 of Figure 2 , the viewing preferences extracted from the text / speech command 202 by the language model parser 204 are sent to the language prompt 220 to convert the viewing preferences into rendering parameters, and the renderer 222 generates the display layout 224 according to the rendering parameters. The language prompt 220 is a prompt or input to the language model 210 that defines a specific rendering input mode. For example, the language prompt 220 can instruct the language model 210 that the renderer 222 can support viewing the patient in the front, left, and right orientations and instruct the language model 210 to convert the user-input viewing preferences into one of these orientations. Other viewing preferences (e.g., windowing presets, available rendering modes, etc.) can be converted similarly.
[0033] At Figure 1 step 108 of, the display layout is output. For example, it can be output by displaying the display layout on a display device of a computer system (e.g., the I / O device 808 of the computer 802 of Figure 8 , or by storing the display layout in a memory or storage device of the computer system (e.g.,Figure 8 store the display layout on the memory 810 or storage device 812 of the computer 802, or output the display layout by transmitting the display layout to a remote computer system (e.g., Figure 8 the computer 802).
[0034] The embodiments described herein can be implemented, for example, in a PACS (Picture Archiving and Communication System) viewer application. For example, when a user loads a patient in a PACS viewer application, the user can provide user input in a chat window, and the system can automatically select medical images and generate a display layout for displaying the medical images to the user.
[0035] Figure 3 FIG. 300 shows a method 300 for generating a clinical knowledge ontology database according to one or more embodiments. The steps or sub-steps of method 300 can be performed by one or more suitable computing devices, such as, for example, Figure 8 the computer 802. Reference will continue to be made to Figure 2 workflow 200 to describe Figure 3 method 300. In one example, method 300 is performed before step 104 of Figure 1 to generate a clinical knowledge ontology database utilized at step 104 of Figure 3 In step 302, one or more medical images are received. In one example, as shown in workflow 200 of Figure 1 the one or more medical images are scanner images 218. The one or more medical images can depict any anatomical object of interest, such as, for example, organs, bones, blood vessels, abnormalities, etc. The one or more medical images can be of any suitable modality, such as, for example, CT (Computed Tomography), MRI (Magnetic Resonance Imaging), US (Ultrasound), x-ray, or any other medical imaging modality or combination of medical imaging modalities. The one or more medical images can be 2D (two-dimensional) images, 3D (three-dimensional) volumes, or 4D (four-dimensional) volumes (3D plus time). The one or more medical images can be received, for example, by directly receiving from an image acquisition device (e.g.,
[0036] At Figure 3 the image acquisition device 814) when the medical image is acquired, by loading previously acquired medical images from the storage device or memory of a computer system (e.g., Figure 2 the memory 810 or storage device 812 of the computer 802), or by receiving medical images from a remote computer system (e.g., Figure 8 the computer 802). Figure 8 the memory 810 or storage device 812 of the computer 802), or by receiving medical images from a remote computer system (e.g., Figure 8 the computer 802).
[0037] At Figure 3 step 304, one or more machine learning models are used to extract features from one or more medical images. In one example, as shown in workflow 200 of Figure 2 , features 214 are extracted from scanner image 218 using AI detector 216. The features may include, for example, metadata of the image (e.g., acquisition type, location, label), patient records (e.g., landmark coordinates, measurements, models, ROI (region of interest) centerlines, images, anatomical fields of view, presence of pathologies such as injuries or foreign bodies, etc.), or any other suitable features extracted from the one or more medical images. The one or more machine learning models convert the imaging data of the one or more medical images into text-based labels and metadata. Any suitable (e.g., well-known) techniques for performing medical imaging analysis tasks may be used to implement the one or more machine learning models.
[0038] At Figure 3 step 306, the extracted features are associated with corresponding nodes in a clinical knowledge ontology database. In one example, as shown in workflow 200 of Figure 2 , the clinical knowledge ontology database is large clinical knowledge AI embedding database 206, and graph model 208 and language model 210 are used to map features 214 to nodes of clinical knowledge ontology 212 to generate large clinical knowledge AI embedding database 206. A clinical knowledge ontology database is created to include all nodes within an ontology database that models clinical terms. The nodes are represented as language embedding vectors of descriptions and aliases. The clinical knowledge ontology database may be based on a standard ontology graph, such as, for example, RadLex, UMLS (Unified Medical Language System), or any other suitable imaging ontology. Optionally, a graph model 208 may be created to store the relationships between nodes and edges to allow for more indirect description associations. To generate an embedding or vector representation of features 214, an embedding model (e.g., text-ada-embedding-002 or even a fine-tuned BERT) is used. In language model 210, the features of the image and DICOM (Digital Imaging and Communications in Medicine) tags (e.g., modality, study date, view position, pixel array, study description) and their links to ontology concepts (e.g., RadLex identifiers) are converted into paragraph descriptions and stored as embeddings. In graph model 208, the links between ontology concepts are stored as embeddings (e.g., the MR concept is a subclass of diagnostic imaging). Then these two types of embeddings are combined to find the one that best matches the user request. Using graph model 208, because if the request is for abstract diagnostic imaging, then it will still find MR images.
[0039] In one embodiment, the extracted features are further associated with corresponding nodes based on image metadata, which can be non - standardized and potentially vendor - specific. For example, the extracted features can include image acquisitions using contrast with an MR (magnetic resonance) protocol. The image acquisition tags and associated MR acquisition parameters are then associated with the corresponding nodes for subsequent retrieval.
[0040] At Figure 3 step 308 of, a clinical knowledge ontology database with the extracted features associated with the corresponding nodes is output. For example, the clinical knowledge ontology database can be output by storing it on the memory or storage device (e.g., Figure 8 the memory 810 or storage device 812 of computer 802) of a computer system, or by transmitting the clinical knowledge ontology database to a remote computer system (e.g., Figure 8 computer 802).
[0041] In one embodiment, Figure 1 method 100 of can be applied to automatically navigate to a specific anatomical location within a medical image based on user input. Figure 4 Workflow 400 for automatic medical image navigation according to one or more embodiments is shown. Figure 4 Workflow 400 of is Figure 2 an extension of workflow 200 of to enable automatic medical image navigation. Workflow 400 of will be described with reference to Figure 1 method 100 of Figure 4 Workflow 400 of.
[0042] To generate a large clinical knowledge AI embedding database 206, for example, during a previous preprocessing stage, an AI detector 216 is applied to detect features 214 from scanner images 218. The features 214 include spatial features representing anatomical text labels associated with, for example, detected spatial coordinates of an anatomical object of interest (e.g., an organ, bone, blood vessel, tumor, etc.), a region of interest mask, a model, etc. The anatomical text labels are concepts expressed in high detail within the large clinical knowledge AI embedding database 206 as nodes of an interconnected graph. Each of the detected anatomical labels is associated with one or more corresponding nodes. The detected spatial features 214 are then stored in a patient record database associated with the nodes. Similarly, text features 214 can be extracted from text-based reports 402 detected using the AI detector 216. Such text features 214 can include expressions in the report 402, such as, for example, an expression describing the anatomical region where an abnormality is described, an expression describing the pathology occurring at a specific anatomical structure, etc. The extracted text features 214 are associated with corresponding nodes in the large clinical knowledge AI embedding database 206. Since the pathology and the location where such pathology occurs are connected within the ontology graph, an association between the pathology and the location is established. Based on the previously generated large clinical knowledge AI embedding database 206 of spatial feature relationships, method 100 can be executed for automated medical image navigation.
[0043] At Figure 1 step 102, a user input is received, the user input including 1) a description of a desired medical image and 2) viewing preferences for the desired medical image. The description of the desired medical image includes a description of an anatomical object of interest to navigate to within one or more medical images. Examples of user input are: "Go to the lungs when a nodule is detected, otherwise start at the head of the pancreas using an abdominal window" or "Go to the low density of the liver at segment IV."
[0044] At Figure 1 step 104, one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image are determined. Within the patient image coordinate nodes, the highest ranked matching node associated with a medical image that has detected an anatomical object of interest that matches the description of the desired medical image is determined. If a match is found, navigation is made to that location within the associated medical image using the associated coordinates. In one example, as Figure 4As shown in workflow 400, spatial computing 404 computes spatial information (e.g., coordinates, paths, models, etc.) for navigation. In one embodiment, post-processing steps may be performed to convert 3D coordinates into a specific view. For example, post-processing steps may be performed to generate a reformatted, clipped planar, or surface view from multiple coordinates. In another example, the coordinates may be used to further sample nearby pixels to optimize windowing or transfer functions for viewing the area.
[0045] At Figure 1 step 106, a display layout of one or more medical images is generated based on viewing preferences, and at Figure 1 step 108, the display layout is output.
[0046] The embodiments described herein are described with respect to the claimed systems and with respect to the claimed methods. The features, advantages, or alternative embodiments herein may be assigned to other claimed subject matter, and vice versa. In other words, the claims and embodiments of the system may be improved using the features described or claimed in the context of the corresponding method. In this case, the functional features of the method are implemented by the physical units of the system.
[0047] Furthermore, the specific embodiments described herein are described with respect to methods and systems that utilize trained machine learning models, and with respect to methods and systems for providing trained machine learning models. The features, advantages, or alternative embodiments herein may be assigned to other claimed subject matter, and vice versa. In other words, the claims and embodiments for providing trained machine learning models may be improved using the features described or claimed in the context of utilizing trained machine learning models, and vice versa. In particular, the datasets used in the methods and systems for utilizing trained machine learning models may have the same properties and characteristics as the corresponding datasets used in the methods and systems for providing trained machine learning models, and the trained machine learning models provided by the corresponding methods and systems may be used in the methods and systems for utilizing trained machine learning models.
[0048] Generally, a trained machine learning model mimics the cognitive functions associated with human thought by other humans. In particular, through training based on training data, a machine learning model is able to adapt to new environments and detect and infer patterns. Another term for "trained machine learning model" is "trained function".
[0049] Generally, the parameters of a machine learning model can be adapted through training. In particular, supervised training, semi-supervised training, unsupervised training, reinforcement learning, and / or active learning can be used. Further, representation learning (an alternative term is "feature learning") can be used. In particular, the parameters of a machine learning model can be iteratively adapted through a number of training steps. In particular, within training, a specific cost function can be minimized. In particular, within the training of a neural network, the backpropagation algorithm can be used.
[0050] In particular, a machine learning model, such as, for example Figure 2 the AI detector 216 of Figure 3 one or more machine learning models utilized at step 304 of Figure 4 the AI detector 216 of, can include, for example, neural networks, support vector machines, decision trees, and / or Bayesian networks, and / or the machine learning model can be based on, for example, k-means clustering, Q-learning, genetic algorithms, and / or association rules. In particular, the neural network can be, for example, a deep neural network, a convolutional neural network, or a convolutional deep neural network. Further, the neural network can be, for example, an adversarial network, a deep adversarial network, and / or a generative adversarial network.
[0051] Figure 5 An embodiment of an artificial neural network 500 is shown, which can be used to implement one or more machine learning models described herein. Alternative terms for "artificial neural network" are "neural network", "artificial neural net", or "neural net".
[0052] The artificial neural network 500 includes nodes 520,..., 532 and edges 540,..., 542, where each edge 540,..., 542 is a directed connection from a first node 520,..., 532 to a second node 520,..., 532. Generally, the first node 520,..., 532 and the second node 520,..., 532 are different nodes 520,..., 532, and it is also possible that the first node 520,..., 532 and the second node 520,..., 532 are the same. For example, in Figure 5 the edge 540 is a directed connection from node 520 to node 523, and the edge 542 is a directed connection from node 530 to node 532. The edge 540,..., 542 from the first node 520,..., 532 to the second node 520,..., 532 is also labeled as the "incoming edge" of the second node 520,..., 532 and the "outgoing edge" of the first node 520,..., 532.
[0053] In this embodiment, the nodes 520, …, 532 of the artificial neural network 500 may be arranged in layers 510, …, 513, where the layers may include an inherent order introduced by the edges 540, …, 542 between the nodes 520, …, 532. In particular, the edges 540, …, 542 may only exist between adjacent node layers. In the illustrated embodiment, the input layer 510 includes only the nodes 520, …, 522 without incoming edges, the output layer 513 includes only the nodes 531, 532 without outgoing edges, and the hidden layers 511, 512 are located between the input layer 510 and the output layer 513. Generally, the number of hidden layers 511, 512 can be arbitrarily selected. The number of nodes 520, …, 522 in the input layer 510 is generally related to the number of input values of the neural network, and the number of nodes 531, 532 in the output layer 513 is generally related to the number of output values of the neural network.
[0054] In particular, (real) numbers may be assigned as values to each of the nodes 520, …, 532 of the neural network 500. Here, x (n) i denotes the value of the i-th node 520, …, 532 in the n-th layer 510, …, 513. The values of the nodes 520, …, 522 in the input layer 510 are equal to the input values of the neural network 500, and the values of the nodes 531, 532 in the output layer 513 are equal to the output values of the neural network 500. Further, each edge 540, …, 542 may include a weight as a real number, in particular, the weight is a real number within the interval [-1, 1] or within the interval [0, 1]. Here, w (m,n) i,j denotes the weight of the edge between the i-th node 520, …, 532 in the m-th layer 510, …, 513 and the j-th node 520, …, 532 in the n-th layer 510, …, 513. Further, the abbreviation w (n) i,j is defined for the weight w (n,n+1) i,j defines.
[0055] In particular, in order to calculate the output values of the neural network 500, the input values are propagated through the neural network. In particular, the value of the node 520, …, 532 in the (n + 1)-th layer 510, …, 513 can be calculated based on the value of the node 520, …, 532 in the n-th layer 510, …, 513 by the following formula
[0056]
[0057] In this text, the function f is a transfer function (another term is "activation function"). Known transfer functions are step functions, sigmoid functions (e.g., logistic function, generalized logistic function, hyperbolic tangent, arctangent function, error function, smoothed step function), or rectifier functions. Transfer functions are mainly used for normalization purposes.
[0058] In particular, the values are propagated layer by layer through the neural network, where the value of the input layer 510 is given by the input of the neural network 500, where the value of the first hidden layer 511 can be calculated based on the value of the input layer 510 of the neural network, where the value of the second hidden layer 512 can be calculated based on the value of the first hidden layer 511, and so on.
[0059] To set the values of the edges the neural network 500 must be trained using training data. In particular, the training data includes training input data and training output data (labeled as t i ). For the training step, the neural network 500 is applied to the training input data to generate the calculated output data. In particular, the training data and the calculated output data include a certain number of values, the number of which is equal to the number of nodes in the output layer.
[0060] In particular, the comparison between the calculated output data and the training data is used to recursively adapt the weights within the neural network 500 (backpropagation algorithm). In particular, the weights are changed according to the following formula:
[0061]
[0062] where γ is the learning rate, and the number δ (n) j can be recursively calculated as
[0063]
[0064] based on δ (n+1) j , if the (n + 1)-th layer is not the output layer, and
[0065]
[0066] if the (n + 1)-th layer is the output layer 513, where f' is the first derivative of the activation function, and t (n+1) j is the comparison training value of the j-th node of the output layer 513.
[0067] A convolutional neural network is a neural network that uses convolutional operations instead of general matrix multiplication in at least one of its layers (so-called "convolutional layers"). In particular, a convolutional layer performs a dot product of one or more convolutional kernels with the input data / image of the convolutional layer, where the entries of the one or more convolutional kernels are parameters or weights adapted through training. In particular, the Frobenius inner product and the ReLU activation function can be used. A convolutional neural network may include additional layers, such as pooling layers, fully connected layers, and normalization layers.
[0068] By using a convolutional neural network, input images can be processed in a very efficient manner because convolutional operations based on different kernels can extract various image features, so that by adapting the weights of the convolutional kernels, relevant image features can be found during training. Furthermore, based on weight sharing in the convolutional kernels, fewer parameters need to be trained, which prevents overfitting during the training phase and allows for faster training or more layers in the network, thus improving the performance of the network.
[0069] Figure 6 An embodiment of a convolutional neural network 600 that can be used to implement one or more machine learning models described herein is shown. In the illustrated embodiment, the convolutional neural network 600 includes an input node layer 610, a convolutional layer 611, a pooling layer 613, a fully connected layer 614, and an output node layer 616, as well as hidden node layers 612, 614. Alternatively, the convolutional neural network 600 may include several convolutional layers 611, several pooling layers 613, and several fully connected layers 615, as well as other types of layers. The order of the layers can be arbitrarily selected, and generally the fully connected layer 615 is used as the last layer before the output layer 616.
[0070] In particular, within the convolutional neural network 600, the nodes 620, 622, 624 of the node layers 610, 612, 614 can be considered to be arranged as a d-dimensional matrix or a d-dimensional image. In particular, in the two-dimensional case, the value of the node 620, 622, 624 indexed by i and j in the nth node layer 610, 612, 614 can be denoted as x(n)[i, j]. However, the arrangement of the nodes 620, 622, 624 of a node layer 610, 612, 614 has no impact on the calculations performed within such a convolutional neural network 600 because these are given only by the structure of the edges and the weights.
[0071] The convolutional layer 611 is a connection layer between the previous node layer 610 (with node values x(n-1)) and the subsequent node layer 612 (with node values x(n)). In particular, the convolutional layer 611 is characterized by the structure and weights of the incoming edges that form a convolutional operation based on a specific number of kernels. In particular, the structure and weights of the edges of the convolutional layer 611 are selected such that the value x(n) of the node 622 in the subsequent node layer 612 is calculated as a convolution x(n) = K * x(n-1) based on the value x(n-1) of the node 620 in the previous node layer 610, where the convolution * is defined in the two-dimensional case as
[0072]
[0073] Here, the kernel K is a d-dimensional matrix (in this embodiment, a two-dimensional matrix), which is typically small compared to the number of nodes 620, 622 (e.g., a 3×3 matrix or a 5×5 matrix). In particular, this implies that the weights of the edges in the convolutional layer 611 are not independent, but are selected such that they produce the said convolutional equation. In particular, for a kernel that is a 3×3 matrix, there are only 9 independent weights (each entry of the kernel matrix corresponds to an independent weight), regardless of the number of nodes 620, 622 in the previous node layer 610 and the subsequent node layer 612.
[0074] Generally, the convolutional neural network 600 uses node layers 610, 612, 614 with multiple channels, especially due to the use of multiple kernels in the convolutional layer 611. In those cases, the node layer can be considered a (d+1)-dimensional matrix (the first dimension indexes the channels). Then, the operation of the convolutional layer 611 is defined by the following two-dimensional example:
[0075]
[0076] where x (n-1)a corresponds to the a-th channel of the previous node layer 610, x (n)b corresponds to the b-th channel of the subsequent node layer 612, and K a,b corresponds to one of the kernels. If the convolutional layer 611 acts on a previous node layer 610 with A channels and outputs a subsequent node layer 612 with B channels, there are A·B independent d-dimensional kernels K a,b .
[0077] Generally, an activation function is used in the convolutional neural network 600. In this embodiment, ReLU (the acronym for "Rectified Linear Units") is used, where R(z) = max(0, z), so the operation of the convolutional layer 611 in the two-dimensional example is
[0078]
[0079] It is also possible to use other activation functions, for example, ELU (acronym for "Exponential Linear Unit"), LeakyReLU, Sigmoid, Tanh, or Softmax.
[0080] In the illustrated embodiment, the input layer 610 includes 36 nodes 620, arranged in a two-dimensional 6×6 matrix. The first hidden node layer 612 includes 72 nodes 622, arranged in two two-dimensional 6×6 matrices, each of which is the result of the convolution of the values of the input layer with a 3×3 kernel within the convolutional layer 611. Equivalently, the nodes 622 of the first hidden node layer 612 can be interpreted as arranged in a three-dimensional 2×6×6 matrix, where the first dimension corresponds to the channel dimension.
[0081] The advantage of using the convolutional layer 611 is that by implementing a local connection pattern between the nodes of adjacent layers, particularly by each node only connecting to a small area of the nodes of the previous layer, the spatial local correlation of the input data can be utilized.
[0082] The pooling layer 613 is a connection layer between the previous node layer 612 (with node values x(n - 1)) and the subsequent node layer 614 (with node values x(n)). In particular, the characteristics of the pooling layer 613 can lie in the structure and weights of the edges forming the pooling operation and the activation function based on the non-linear pooling function f. For example, in the two-dimensional case, the value x(n) of the nodes 624 in the subsequent node layer 614 can be calculated based on the value x(n - 1) of the nodes 622 in the previous node layer 612 as follows
[0083]
[0084] In other words, by using the pooling layer 613, the number of nodes 622, 624 can be reduced by repositioning a number d1·d2 of adjacent nodes 622 in the previous node layer 612 and calculating a single node 622 in the subsequent node layer 614 as a function of the values of the said number of adjacent nodes. In particular, the pooling function f can be the maximum function, the average value, or the L2 norm. In particular, for the pooling layer 613, the weights of the incoming edges are fixed and not modified through training.
[0085] The advantage of using the pooling layer 613 is to reduce the number of nodes 622, 624 and the number of parameters. This results in a reduction in the amount of computation in the network and control of overfitting.
[0086] In the illustrated embodiment, the pooling layer 613 is a max pooling layer that replaces four adjacent nodes with only one node, the value being the maximum value among the values of the four adjacent nodes. Max pooling is applied to each d-dimensional matrix of the previous layer; in this embodiment, max pooling is applied to each of the two two-dimensional matrices, thereby reducing the number of nodes from 72 to 18.
[0087] Generally, the last layer of the convolutional neural network 600 is the fully connected layer 615. The fully connected layer 615 is a connection layer between the previous node layer 614 and the subsequent node layer 616. The features of the fully connected layer 613 can lie in the fact that there are most, especially all, edges between the nodes 614 of the previous node layer 614 and the nodes 616 of the subsequent node layer, and the weights of each of these edges can be adjusted individually.
[0088] In this embodiment, the nodes 624 in the previous node layer 614 of the fully connected layer 615 are shown both as a two-dimensional matrix and additionally as unconnected nodes (indicated as node rows, where the number of nodes is reduced for better presentation). Such an operation is also labeled as "flattening". In this embodiment, the number of nodes 626 in the subsequent node layer 616 of the fully connected layer 615 is less than the number of nodes 624 in the previous node layer 614. Alternatively, the number of nodes 626 can be equal to or greater.
[0089] Furthermore, in this embodiment, the Softmax activation function is used within the fully connected layer 615. By applying the Softmax function, the sum of the values of all the nodes 626 of the output layer 616 is 1, and all the values of all the nodes 626 of the output layer 616 are real numbers between 0 and 1. In particular, if the convolutional neural network 600 is used for classifying input data, the values of the output layer 616 can be interpreted as the probabilities that the input data falls into one of the different classes.
[0090] In particular, the convolutional neural network 600 can be trained based on the backpropagation algorithm. To prevent overfitting, regularization methods can be used, such as dropout of the nodes 620,..., 624, stochastic pooling, the use of artificial data, weight decay based on the L1 or L2 norm or the max norm constraint.
[0091] According to one aspect, a machine learning model may include one or more Residual Networks (ResNets). In particular, a ResNet is an artificial neural network that includes at least one skip or bypass connection for skipping at least one layer of the artificial neural network. In particular, a ResNet may be a convolutional neural network that includes one or more bypass connections for bypassing one or more convolutional layers accordingly. According to some examples, a ResNet may be represented as an m-layer ResNet, where m is the number of layers in the corresponding architecture, and according to some examples, may take on the values of 34, 50, 101, or 152. According to some examples, such an m-layer ResNet may include (m - 2) / 2 bypass connections accordingly.
[0092] A bypass connection may be regarded as a bypass that leaks through one or more bypassed layers and directly feeds the output of the previous layer to the layer succeeding the one or more bypassed layers. Instead of having to directly fit the desired mapping, the bypassed layer will have to fit the residual mapping that "balances" the directly fed output.
[0093] Fitting the residual mapping is computationally easier to optimize than the directed mapping. More importantly, this alleviates the problem of vanishing / exploding gradients during optimization when training the machine learning model: if the bypassed layer encounters such a problem, its contribution can be skipped by regularizing the directly fed output. Thus, the benefit of using ResNets is that much deeper networks can be trained.
[0094] In particular, a recurrent machine learning model is a machine learning model as described below: its output depends not only on the input values and the parameters of the machine learning model adapted by the training process, but also on a hidden state vector, where the hidden state vector is based on previous inputs to the recurrent machine learning model. In particular, a recurrent machine learning model may include additional storage state or additional structure that incorporates a time delay or includes a feedback loop.
[0095] In particular, the underlying structure of a recurrent machine learning model may be a neural network, which may be labeled as a recurrent neural network. Such a recurrent neural network may be described as an artificial neural network where the connections between nodes form a directed graph along a time series. In particular, a recurrent neural network may be interpreted as a directed acyclic graph. In particular, a recurrent neural network may be a finite impulse recurrent neural network or an infinite impulse recurrent neural network (where the finite impulse network can be unfolded and replaced with a strictly feedforward neural network, and the infinite impulse network cannot be unfolded and cannot be replaced with a strictly feedforward neural network).
[0096] In particular, training a recurrent neural network can be based on the BPTT algorithm (acronym for "backpropagation through time"), the RTRL algorithm (acronym for "real-time recurrent learning"), and / or a genetic algorithm.
[0097] By using a recurrent machine learning model, input data including variable-length sequences can be used. In particular, this implies that the method cannot be used only for a fixed number of input data sets (and needs to be trained differently for each other number of input data sets used as input), but can be used for any number of input data sets. This implies that, independent of the number of input data sets included in different sequences, the entire training data set can be used within the training, and the training data is not reduced to the training data corresponding to a specific number of successive input data sets.
[0098] Figure 7 Both the recurrent representation 702 and the unfolded representation 704 show the schematic structure of the recurrent machine learning model F, which can be used to implement one or more of the machine learning models described herein. The recurrent machine learning model takes a number of input data sets x, x 1 , …, x N 706 as input and creates a corresponding set of output data sets y, y 1 , …, y N 708. Further, the output depends on the so-called hidden vectors h, h 1 , …, h N 710, which implicitly includes information about the input data sets previously used as input to the recurrent machine learning model F 712. By using these hidden vectors h, h 1 , …, h N 710, the sequential nature of the input data sets can be exploited.
[0099] In a single step of processing, the recurrent machine learning model F 712 takes the hidden vector h n-1 created in the previous step and the input data set x n as input. Within this step, the recurrent machine learning model F generates an updated hidden vector hn and an output data set y n as output. In other words, one step of processing computes (y n , h n ) = F(x n , h n-1), or by splitting the recurrent machine learning model F712 into a part F(y) for calculating the output data and a part F(h) for calculating the hidden vector, one step of the processing calculates y n = F (y) (x n , h n-1 ) and h n = F (h) (x n , h n-1 ). For the first processing step, h 0 can be randomly selected or filled with all-zero entries. The parameters of the recurrent machine learning model F712 trained previously based on the training dataset do not change between different processing steps.
[0100] In particular, the output data and the hidden vector of the processing step depend on all the previous input datasets used in the previous steps. y n = F (y) (x n , F (h) (x n-1 , h n-2 ) and h n = (x n , F (h) (x n-1 , h n-2 ))
[0101] The systems, apparatuses, and methods described herein can be implemented using digital circuits or using one or more computers that utilize well-known computer processors, memory units, storage devices, computer software, and other components. Typically, a computer includes a processor for executing instructions and one or more memories for storing instructions and data. The computer may also include or be coupled to one or more mass storage devices, such as one or more disks, internal hard drives and removable disks, magneto-optical disks, optical disks, etc.
[0102] The systems, apparatuses, and methods described herein can be implemented using computers operating in a client-server relationship. Typically, in such a system, the client computer is located remotely from the server computer and interacts via a network. The client-server relationship can be defined and controlled by computer programs running on the respective client and server computers.
[0103] The systems, apparatuses, and methods described herein can be implemented within a network-based cloud computing system. In such a network-based cloud computing system, servers or other processors connected to the network communicate with one or more client computers via the network. For example, a client computer can communicate with a server via a web browser application that resides and operates on the client computer. The client computer can store data on the server and access the data via the network. The client computer can transmit requests for data or requests for online services to the server via the network. The server can perform the requested services and provide data to the client computer(s). The server can also transmit data that is adapted to cause the client computer to perform a specific function, such as to perform a calculation, display specific data on a screen, etc. For example, the server can transmit a request that is adapted to cause the client computer to perform one or more of the steps or functions of the methods and workflows described herein, including Figures 1-4 one or more of the steps or functions of. Specific steps or functions of the methods and workflows described herein, including Figures 1-4 one or more of the steps or functions of, can be performed by a server in a network-based cloud computing system or by another processor. Specific steps or functions of the methods and workflows described herein, including Figures 1-4 one or more of the steps, can be performed by a client computer in a network-based cloud computing system. Steps or functions of the methods and workflows described herein, including Figures 1-4 one or more of the steps, can be performed by a server and / or client computer in a network-based cloud computing system in any combination.
[0104] The systems, apparatuses, and methods described herein can be implemented using a computer program product that is tangibly embodied in an information carrier, such as in a non-transitory machine-readable storage device, for execution by a programmable processor; and the method and workflow steps described herein, including Figures 1-4 one or more of the steps or functions of, can be implemented using one or more computer programs executable by such a processor. A computer program is a set of computer program instructions that can be used directly or indirectly in a computer to perform a specific activity or produce a specific result. A computer program can be written in any form of programming language, including a compiled or interpreted language, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0105] Figure 8The high-level block diagram of an example computer 802 that can be used to implement the systems, apparatuses, and methods described herein is depicted. Computer 802 includes a processor 804 operatively coupled to a data storage device 812 and a memory 810. The processor 804 controls the overall operation of computer 802 by executing computer program instructions that define such operations. The computer program instructions can be stored in the data storage device 812 or other computer-readable media and are loaded into the memory 810 when it is desired to execute the computer program instructions. Thus, Figures 1-4 the methods and workflow steps or functions of Figures 1-4 can be defined by computer program instructions stored in the memory 810 and / or data storage device 812 and are controlled by the processor 804 that executes the computer program instructions. For example, the computer program instructions can be implemented as computer-executable code programmed by those skilled in the art to perform Figures 1-4 the methods and workflow steps or functions of
[0106] Thus, by executing the computer program instructions, the processor 804 performs
[0107] the methods and workflow steps or functions of. Computer 802 may also include one or more network interfaces 806 for communicating with other devices via a network. Computer 802 may also include one or more input / output devices 808 (e.g., a display, keyboard, mouse, microphone, speaker, buttons, etc.) that enable a user to interact with computer 802.
[0106] The processor 804 can include both general and special-purpose microprocessors and can be the sole processor of computer 802 or one of multiple processors. For example, the processor 804 can include one or more central processing units (CPUs). The processor 804, data storage device 812, and / or memory 810 can include one or more application-specific integrated circuits (ASICs) and / or one or more field-programmable gate arrays (FPGAs), supplemented or incorporated by one or more application-specific integrated circuits (ASICs) and / or one or more field-programmable gate arrays (FPGAs).
[0107] Each of the data storage device 812 and the memory 810 includes a tangible non-transitory computer-readable storage medium. The data storage device 812 and the memory 810 may each include high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate synchronous dynamic random access memory (DDR RAM), or other random access solid-state storage devices, and may include non-volatile memory, such as one or more disk storage devices (such as internal hard disks and removable disks), magneto-optical storage devices, optical disk storage devices, flash memory devices, semiconductor storage devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), digital versatile disc read-only memory (DVD-ROM) discs, or other non-volatile solid-state storage devices.
[0108] The input / output device 808 may include peripheral devices, such as printers, scanners, displays, etc. For example, the input / output device 808 may include a display device, such as a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor, for displaying information to the user, a keyboard, and a pointing device, such as a mouse or a trackball, through which the user may provide input to the computer 802.
[0109] The image acquisition device 814 may be connected to the computer 802 to input image data (e.g., medical images) into the computer 802. It is possible to implement the image acquisition device 814 and the computer 802 as one device. It is also possible for the image acquisition device 814 and the computer 802 to communicate wirelessly through a network. In a possible embodiment, the computer 802 may be remotely located with respect to the image acquisition device 814.
[0110] Any and all systems, apparatuses, and methods discussed herein may be implemented using one or more computers, such as the computer 802.
[0111] Those skilled in the art will recognize that the implementation of an actual computer or computer system may have other structures and may also include other components, and Figure 8 is a high-level representation of some of the components of such a computer for illustrative purposes.
[0112] Independently of the use of grammatical terms, individuals with male or female gender identities are included within the term.
[0113] The foregoing detailed description is to be understood as illustrative and exemplary in every aspect and not restrictive, and the scope of the invention disclosed herein is determined not from the detailed description but from the claims as interpreted in accordance with the full width permitted by patent law. It should be understood that the embodiments shown and described herein are merely illustrative of the principles of the invention and that various modifications can be effected by those skilled in the art without departing from the scope and spirit of the invention. Without departing from the scope and spirit of the invention, those skilled in the art can effect various other combinations of features.
[0114] The following is a list of non-limiting illustrative embodiments disclosed herein:
[0115] Illustrative Embodiment 1. A computer-implemented method, comprising: receiving user input, the user input including 1) a description of a desired medical image and 2) viewing preferences for the desired medical image; determining one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image, the one or more matching nodes being associated with one or more medical images in the clinical knowledge ontology database; generating a display layout for the one or more medical images based on the viewing preferences; and outputting the display layout.
[0116] Illustrative Embodiment 2. The computer-implemented method according to Illustrative Embodiment 1, wherein: receiving user input including 1) a description of a desired medical image and 2) viewing preferences for the desired medical image includes receiving user input further including a temporal description of the desired medical image; and determining one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image includes determining one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image based on the temporal description of the desired medical image.
[0117] Illustrative Embodiment 3. The computer-implemented method according to any one of Illustrative Embodiments 1-2, wherein determining one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image includes: parsing the user input into a vector; performing a vector search between the vector representing the description of the desired medical image and the vectors representing the nodes of the clinical knowledge ontology database to identify a list of ranked nodes; and identifying the one or more highest-ranked nodes that meet a ranking threshold as the one or more nodes.
[0118] Illustrative Example 4. The computer-implemented method according to any one of Illustrative Examples 1-3 further includes generating a clinical knowledge ontology database by the following steps: receiving one or more medical images; extracting features from the one or more medical images using one or more machine learning models; associating the extracted features with corresponding nodes of the clinical knowledge ontology database; and outputting the clinical knowledge ontology database having the extracted features associated with the corresponding nodes.
[0119] Illustrative Example 5. The computer-implemented method according to any one of Illustrative Examples 1-4, wherein the description of the desired medical image includes a description of the anatomical object of interest to navigate to within the one or more medical images, and wherein the one or more matching nodes associated with the one or more medical images in the clinical knowledge ontology database include one or more matching nodes associated with the coordinates of the anatomical object of interest in the one or more medical images.
[0120] Illustrative Example 6. The computer-implemented method according to any one of Illustrative Examples 1-5, wherein the description of the desired medical image is defined based on at least one of imaging modality, acquisition parameters, acquisition orientation, image appearance, anatomical field of view, or classification and detection.
[0121] Illustrative Example 7. The computer-implemented method according to any one of Illustrative Examples 1-6, wherein the viewing preference is defined based on at least one of anatomical display orientation, windowing, or rendering mode.
[0122] Illustrative Example 8. The computer-implemented method according to any one of Illustrative Examples 1-7, wherein generating a display layout of the one or more medical images based on the viewing preference includes: using a language model to convert the viewing preference into rendering parameters.
[0123] Illustrative Example 9. The computer-implemented method according to any one of Illustrative Examples 1-8, wherein outputting the display layout includes: displaying the display layout on a display device.
[0124] Illustrative Example 10. An apparatus includes: a component for receiving user input, the user input including 1) a description of a desired medical image and 2) a viewing preference for the desired medical image; a component for determining one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image, the one or more matching nodes being associated with one or more medical images in the clinical knowledge ontology database; a component for generating a display layout of the one or more medical images based on the viewing preference; and a component for outputting the display layout.
[0125] Exemplary Embodiment 11. The apparatus according to Exemplary Embodiment 10, wherein: the component for receiving user input including 1) a description of a desired medical image and 2) viewing preferences for the desired medical image includes a component for receiving user input further including a temporal description of the desired medical image; and the component for determining one or more nodes of the clinical knowledge ontology database that match the description of the desired medical image includes a component for determining, based on the temporal description of the desired medical image, one or more nodes of the clinical knowledge ontology database that match the description of the desired medical image.
[0126] Exemplary Embodiment 12. The apparatus according to any one of Exemplary Embodiments 10-11, wherein the component for determining one or more nodes of the clinical knowledge ontology database that match the description of the desired medical image includes: a component for parsing the user input into a vector; a component for performing a vector search between the vector representing the description of the desired medical image and the vectors representing the nodes of the clinical knowledge ontology database to identify a ranked list of nodes; and a component for identifying one or more of the highest-ranked nodes that meet a ranking threshold as the one or more nodes.
[0127] Exemplary Embodiment 13. The apparatus according to any one of Exemplary Embodiments 10-12, further including a component for generating the clinical knowledge ontology database by: a component for receiving one or more medical images; a component for extracting features from the one or more medical images using one or more machine learning models; a component for associating the extracted features with corresponding nodes of the clinical knowledge ontology database; and a component for outputting the clinical knowledge ontology database having the extracted features associated with the corresponding nodes.
[0128] Exemplary Embodiment 14. The apparatus according to any one of Exemplary Embodiments 10-13, wherein the description of the desired medical image includes a description of an anatomical object of interest to navigate to within one or more medical images, and wherein one or more matching nodes associated with the one or more medical images in the clinical knowledge ontology database include one or more matching nodes associated with the coordinates of the anatomical object of interest in the one or more medical images.
[0129] Exemplary Embodiment 15. A non-transitory computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform operations, the operations including: receiving user input, the user input including 1) a description of a desired medical image and 2) viewing preferences for the desired medical image; determining one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image, the one or more matching nodes being associated with one or more medical images in the clinical knowledge ontology database; generating a display layout for the one or more medical images based on the viewing preferences; and outputting the display layout.
[0130] Exemplary Embodiment 16. The non-transitory computer-readable medium according to Exemplary Embodiment 15, wherein: receiving user input including 1) a description of a desired medical image and 2) viewing preferences for the desired medical image includes receiving user input further including a temporal description of the desired medical image; and determining one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image includes determining one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image based on the temporal description of the desired medical image.
[0131] Exemplary Embodiment 17. The non-transitory computer-readable medium according to any one of Exemplary Embodiments 15-16, wherein the description of the desired medical image is defined based on at least one of an imaging modality, acquisition parameters, acquisition orientation, image appearance, anatomical field of view, or classification and detection.
[0132] Exemplary Embodiment 18. The non-transitory computer-readable medium according to any one of Exemplary Embodiments 15-17, wherein the viewing preferences are defined based on at least one of an anatomical display orientation, windowing, or rendering mode.
[0133] Exemplary Embodiment 19. The non-transitory computer-readable medium according to any one of Exemplary Embodiments 15-18, wherein generating a display layout for the one or more medical images based on the viewing preferences includes: using a language model to translate the viewing preferences into rendering parameters.
[0134] Exemplary Embodiment 20. The non-transitory computer-readable medium according to any one of Exemplary Embodiments 15-19, wherein outputting the display layout includes: displaying the display layout on a display device.
Claims
1. A computer-implemented method comprising: receiving user input, the user input comprising 1) a description of a desired medical image and 2) a viewing preference of the desired medical image; determining one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image, the one or more matching nodes being associated with one or more medical images in the clinical knowledge ontology database; generating a display layout of the one or more medical images based on the viewing preferences; and Output display layout.
2. The computer-implemented method of claim 1 , wherein: Receiving user input including 1) a description of the desired medical image and 2) a viewing preference of the desired medical image includes receiving user input further including a temporal description of the desired medical image; and Determining one or more nodes of the clinical knowledge ontology database that match the description of the expected medical image includes determining one or more nodes of the clinical knowledge ontology database that match the description of the expected medical image based on the temporal description of the expected medical image.
3. The computer-implemented method of claim 1 , wherein determining one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image comprises: Parse user input into vectors; performing a vector search between a vector representing a description of a desired medical image and a vector representing a node of a clinical knowledge ontology database to identify a list of ranked nodes; and One or more highest-ranked nodes that satisfy a ranking threshold are identified as the one or more nodes.
4. The computer-implemented method of claim 1 , further comprising generating the clinical knowledge ontology database by: receiving one or more medical images; extracting features from the one or more medical images using one or more machine learning models; Associating the extracted features with corresponding nodes of a clinical knowledge ontology database; and A clinical knowledge ontology database having the extracted features associated with the corresponding nodes is output.
5. A computer-implemented method according to claim 1, wherein the description of the expected medical image includes a description of an anatomical object of interest to which one or more medical images are navigated within, and wherein the one or more matching nodes associated with the one or more medical images in the clinical knowledge ontology database include one or more matching nodes associated with the coordinates of the anatomical object of interest in the one or more medical images.
6. The computer-implemented method of claim 1, wherein the description of the desired medical image is defined based on at least one of imaging modality, acquisition parameters, acquisition orientation, image appearance, anatomical field of view, or classification and detection.
7. A computer-implemented method according to claim 1, wherein the viewing preference is defined based on at least one of anatomical display orientation, windowing, or rendering mode.
8. The computer-implemented method of claim 1 , wherein generating a display layout of the one or more medical images based on the viewing preferences comprises: Use language models to convert viewing preferences into rendering parameters.
9. The computer-implemented method of claim 1 , wherein outputting the display layout comprises: Display the display layout on the display device.
10. An apparatus comprising: means for receiving user input, the user input comprising 1) a description of a desired medical image and 2) viewing preferences for the desired medical image; means for determining one or more nodes of a clinical knowledge ontology database that match the description of a desired medical image, the one or more matching nodes being associated with one or more medical images in the clinical knowledge ontology database; means for generating a display layout of one or more medical images based on the viewing preferences; and A widget used to output a display layout.
11. The device according to claim 10, wherein: means for receiving user input including 1) a description of a desired medical image and 2) a viewing preference of the desired medical image includes means for receiving user input further including a temporal description of the desired medical image; and The means for determining one or more nodes of the clinical knowledge ontology database that match the description of the expected medical image includes means for determining one or more nodes of the clinical knowledge ontology database that match the description of the expected medical image based on the temporal description of the expected medical image.
12. The apparatus of claim 10, wherein the component for determining one or more nodes of a clinical knowledge ontology database that matches the description of the expected medical image comprises: Components for parsing user input into vectors; means for performing a vector search between a vector representing a description of a desired medical image and a vector representing a node of a clinical knowledge ontology database to identify a ranked list of nodes; and Means for identifying one or more highest ranked nodes that satisfy a ranking threshold as the one or more nodes.
13. The apparatus according to claim 10, further comprising a component for generating the clinical knowledge ontology database in the following manner: means for receiving one or more medical images; means for extracting features from one or more medical images using one or more machine learning models; A component for associating the extracted features with corresponding nodes of a clinical knowledge ontology database; and Means for outputting a clinical knowledge ontology database having the extracted features associated with corresponding nodes.
14. An apparatus according to claim 10, wherein the description of the expected medical image includes a description of an anatomical object of interest to which one or more medical images are navigated within, and wherein the one or more matching nodes associated with the one or more medical images in the clinical knowledge ontology database include one or more matching nodes associated with the coordinates of the anatomical object of interest in the one or more medical images.
15. A non-transitory computer-readable medium comprising instructions which, when executed by a computer, cause the computer to perform operations comprising: receiving user input, the user input comprising 1) a description of a desired medical image and 2) a viewing preference of the desired medical image; determining one or more nodes of a clinical knowledge ontology database that match the description of the desired medical image, the one or more matching nodes being associated with one or more medical images in the clinical knowledge ontology database; generating a display layout of the one or more medical images based on the viewing preferences; and Output display layout.
16. The non-transitory computer readable medium of claim 15, wherein: Receiving user input including 1) a description of the desired medical image and 2) a viewing preference of the desired medical image includes receiving user input further including a temporal description of the desired medical image; and Determining one or more nodes of the clinical knowledge ontology database that match the description of the expected medical image includes determining one or more nodes of the clinical knowledge ontology database that match the description of the expected medical image based on the temporal description of the expected medical image.
17. The non-transitory computer readable medium of claim 15, wherein the description of the desired medical image is defined based on at least one of imaging modality, acquisition parameters, acquisition orientation, image appearance, anatomical field of view, or classification and detection.
18. The non-transitory computer-readable medium of claim 15, wherein the viewing preference is defined based on at least one of anatomical display orientation, windowing, or rendering mode.
19. The non-transitory computer-readable medium of claim 15, wherein generating a display layout of the one or more medical images based on the viewing preferences comprises: Use language models to convert viewing preferences into rendering parameters.
20. The non-transitory computer readable medium of claim 15, wherein outputting the display layout comprises: Display the display layout on the display device.