Training machine learning algorithms using digitally reconstructed radiological images

By combining DRR with annotated data to generate machine learning models, the problem of time-consuming and low quality of data collection in the prior art is solved, and the position of anatomical structures in medical images is automatically marked, training efficiency and accuracy is improved, and it is suitable for medical image guidance programs and surgical navigation systems.

CN120259835APending Publication Date: 2025-07-04BOYI LAI EUROPE AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510178589.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2019-09-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the process of training using machine learning algorithms, data collection takes time and is of low quality, making it difficult to automatically label anatomical structure locations in medical images, resulting in inefficient training.

Method used

Digital reconstruction radiographic images (DRR) are used to input machine learning algorithms together with annotation data. By generating the adaptability of the machine learning model, the likelihood relationship between anatomy and annotation is established, and the anatomy position is automatically marked.

Benefits of technology

Improves the efficiency and accuracy of medical image data training, enables automatic labeling of anatomical locations, and is suitable for image guidance programs, surgical navigation and cloud-based surgical planning systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259835A_ABST
    Figure CN120259835A_ABST
Patent Text Reader

Abstract

A computer-implemented method for training a likelihood-based computational model to determine a position of an image representation of an annotated anatomical structure in a two-dimensional X-ray image includes inputting a medical DRR together with an annotation to a machine learning algorithm to train the algorithm, the adaptive learnable parameters of the machine learning model are generated. An annotation may be derived from metadata associated with the DRR, or may be included in atlas data that matches the DRR to establish a relationship between the annotation included in the atlas data and the DRR. The machine learning algorithm thus generated may be re-used to analyze clinical or synthetic DRRs in order to appropriately add annotations to those DRRs and / or identify the location of anatomical structures in those DRRs.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a patent application for invention with the application date of September 20, 2019, application number 201980061756.8, and title "Training Machine Learning Algorithms Using Digitally Reconstructed Radiographs". Technical Field

[0002] The present invention relates to a computer-implemented method for training a likelihood-based computational model to determine the position of an image representation of an annotated anatomical structure in a two-dimensional X-ray image, a corresponding computer program, a computer-implemented method for determining the relationship between an anatomical structure represented in a two-dimensional medical image and an annotation of the anatomical structure, a program storage medium storing the program, a computer for executing the program, and a medical system including an electronic data storage device and the above computer. Background Art

[0003] Previously, synthetic models have been used to train machine learning algorithms. For example, Microsoft Kinect is trained with 3D models that provide labels for each pose.

[0004] There is no mention in the literature of using digitally reconstructed radiographs (DRRs) to train an algorithm and then using it with other data sets. The difficulty with training only on clinical data is that collecting the data can be time-consuming, the quality of the data can be low, and classification and labeling must be performed for machine learning. With DRRs, the labeling process can be automated.

[0005] The object of the present invention is to provide an improved method for training and using artificial intelligence (AI) algorithms for applying annotations to medical image data or for detecting the image position of a predetermined anatomical structure in medical image data.

[0006] The present invention can be used in image-guided procedures, for example, related to all products of Brainlab AG for radiotherapy (such as and ), surgical navigation (such as or ) or cloud-based surgical planning (such as ).

[0007] Aspects, examples, and exemplary steps of the present invention and their embodiments are disclosed below. Different exemplary features of the present invention can be combined according to the present invention as long as it is technically appropriate and feasible. Summary of the Invention

[0008] In the following, a brief description of specific features of the present invention is given, and it should not be understood that the present invention is limited to the features or combinations of features described in this part.

[0009] The method of the present disclosure includes inputting a medical DRR together with annotations into a machine learning algorithm to train the algorithm, i.e., to generate adaptable learnable parameters of a machine learning model. The annotations can be derived from metadata associated with the DRR or can be included in atlas data that matches the DRR to establish a relationship between the annotations included in the atlas data and the DRR. The machine learning algorithm thus generated can be used again to analyze clinical or synthetic DRRs in order to appropriately add annotations to those DRRs and / or identify the locations of anatomical structures in those DRRs.

[0010] In this section, a description of the general features of the present invention is given, for example, by reference to workable embodiments of the present invention.

[0011] Generally, in a first aspect, a solution of the present invention to achieve the above object is a computer-implemented medical method for training a likelihood-based computational model to determine the location of an image representation of an annotated anatomical structure in a two-dimensional X-ray image. The method according to the first aspect includes performing, on at least one processor of at least one computer (e.g., at least one computer that is part of a navigation system), the following exemplary steps performed by the at least one processor.

[0012] In an (e.g., first) exemplary step of the method according to the first aspect, image training data is obtained, the image training data describing a synthetic two-dimensional X-ray image (e.g., a digital reconstructed radiograph - DRR), also referred to as a training image, that includes an image representation of an anatomical structure. This step corresponds to inputting a set of training DRRs for training the likelihood-based computational model. The term "anatomical structure" includes abnormal tissues, such as pathological tissues, such as tumors or bone fractures or skeletal displacements or medical implants (such as screws or artificial intervertebral discs or prostheses).

[0013] In an exemplary step (e.g., the second) of the method according to the first aspect, annotation data is obtained, which describes the annotation of the anatomical structure. The annotation is at least one of, for example, the following information: information describing the anatomical perspective of the image representation defining the anatomical structure (e.g., information describing whether the image representation is generated from the left side or the right side of the anatomical structure), information describing a subset or segmentation of the image representation (e.g., the bounding box defining the subset), or information describing the classification of the attributes defining the anatomical structure (e.g., the pathological degree of the anatomical structure or its identity, such as its anatomical designation and / or name). For example, the annotation data is determined from the metadata included in the image training data. In one example, atlas data is obtained, which describes the image-based model of the anatomical structure, and then the annotation data is determined, for example, based on the image training data and the atlas data. For example, this can be done by matching the training image with the image-based model, for example, by performing an image fusion algorithm on the two data sets to find the corresponding image structures. The image-based model includes, for example, data objects, such as representations of anatomical landmarks, whose geometry can be matched with the image components of the training image to find the structures corresponding to certain structures visible in the training image in the image-based model. Annotations can be defined with respect to the corresponding structures in the image-based model, and the annotations can be transferred to the training image based on the matching.

[0014] In an exemplary step (e.g., the third) of the method according to the first aspect, model parameter data is determined, which describes the model parameters (e.g., learnable parameters such as biases, batch normalization, or weights) of a likelihood-based computational model for establishing a likelihood-based relationship (e.g., likelihood-based association) between the anatomical structure and the annotation in a two-dimensional X-ray. For example, the computational model includes or consists of an artificial intelligence (AI) algorithm (e.g., a machine learning (ML) algorithm); in one example, a convolutional neural network is part of the computational model. For example, the model parameter data is determined by inputting the image training data and the annotation data into a function that establishes a likelihood-based relationship (and then executing the function based on the input). For example, the function establishes a likelihood-based relationship between the position of the anatomical structure in the two-dimensional X-ray image and the position for displaying the annotation in the two-dimensional X-ray image. Thus, a computational model such as a machine learning algorithm can be trained to establish a relationship between the position of the landmark and the position for labeling it in association with the training image (e.g., in the image).

[0015] In an example of the method according to the first aspect, medical image data is obtained, the medical image data depicting a three-dimensional medical image comprising an image representation of an anatomical structure, wherein determining the training image data is by determining an image value threshold (such as an intensity threshold) associated with the image representation of the anatomical structure in the three-dimensional medical image, defining a corresponding intensity mapping function, and generating, based on the intensity mapping function, an image representation of the anatomical structure in each two-dimensional synthetic X-ray image from the image representation of at least one (such as exactly one, a proper subset, or all) of the anatomical structures in the three-dimensional medical image.

[0016] In an example of the method according to the first aspect, atlas data is obtained, the atlas data depicting an image-based model of an anatomical structure and at least one projection parameter for generating a two-dimensional medical image. Then, a two-dimensional medical image is generated based on the at least one projection parameter. In particular, projection parameters (such as the perspective angle for generating a training image (DRR)) are obtained from the atlas data.

[0017] In a second aspect, the invention pertains to a computer-implemented method for determining a relationship between an anatomical structure represented in a two-dimensional medical image and an annotation of the anatomical structure. The method according to the second aspect includes performing, on at least one processor of at least one computer (such as at least one computer forming part of a navigation system), the following exemplary steps performed by the at least one processor.

[0018] In an (e.g., first) exemplary step of the method according to the second aspect, patient image data is obtained, the patient image data depicting a (synthetic or clinical, i.e., real) two-dimensional X-ray image that includes an image representation of an anatomical structure of a patient. For example, the patient image data has been generated by synthesizing a two-dimensional X-ray image from a three-dimensional image of the anatomical structure, or has been generated by applying an X-ray-based imaging modality (such as a fluoroscopic imaging modality or a tomographic imaging modality, such as computed X-ray tomography or magnetic resonance imaging of the anatomical structure) (and, in the latter case, generating the two-dimensional X-ray image by computed X-ray tomography or magnetic resonance tomography, respectively).

[0019] In an (e.g., second) exemplary step of the method according to the second aspect, structure annotation prediction data is determined, the structure annotation prediction data describing, according to a certain likelihood determined by a computational model, the location of the image representation of the anatomical structure in the two-dimensional X-ray image (the image being described by patient annotation data and an annotation of the anatomical structure), wherein the structure annotation data is determined by inputting the patient image data into a function that establishes a likelihood-based relationship between the image representation of the anatomical structure in the two-dimensional X-ray image and the annotation of the anatomical structure, the function being part of a computational model that has been trained by performing the method according to the first aspect (and then performing the function on this input).

[0020] In an example of the method according to the second aspect, a likelihood-based relationship is established between the position of an anatomical structure in a two-dimensional X-ray image described by patient image data and the position for displaying an annotation in the two-dimensional X-ray image described by patient image data, and the structural annotation data describes a likelihood-based relationship (such as a likelihood-based association) between the position of an image representation of an anatomical structure in a two-dimensional X-ray image described by patient image data and the position for displaying an annotation in the two-dimensional X-ray image described by patient image data.

[0021] In a third aspect, the present invention is directed to a computer program which, when running on at least one processor (e.g., one processor) of at least one computer (e.g., one computer) or when loaded into at least one memory (e.g., one memory) of at least one computer (e.g., one computer), causes the at least one computer to perform the above-described method according to the first or second aspect. Alternatively or additionally, the present invention may relate to a (e.g., physically, e.g., electrically generated by technical means) signal wave carrying information representing a program, such as the above program, e.g., a digital signal wave, such as an electromagnetic carrier wave, the program including, for example, code means adapted to perform any or all steps of the method according to the first aspect. In one example, the signal wave is a data carrier signal carrying the above computer program. A computer program stored on a disk is a data file which, when read and transmitted, becomes a data stream in the form of, for example, a (e.g., physically, e.g., electrically generated by technical means) signal. The signal may be implemented as a signal wave, such as the electromagnetic carrier wave described herein. For example, the signal (e.g., signal wave) is configured to be transmitted via a computer network, such as a LAN, WLAN, WAN, mobile network (e.g., the Internet). For example, the signal (e.g., signal wave) is configured to be transmitted by optical or acoustic data transmission. Thus, the present invention according to the third aspect may alternatively or additionally relate to a data stream representing the above program.

[0022] In a fourth aspect, the present invention is directed to a computer-readable program storage medium storing a program according to the third aspect. The program storage medium is, for example, non-transitory.

[0023] In a fifth aspect, the present invention is directed to a program storage medium storing data defining model parameters and an architecture of a likelihood-based computational model that has been trained by performing the method according to the first aspect.

[0024] In a sixth aspect, the present invention is directed to a data carrier signal that carries data defining model parameters and an architecture of a likelihood-based computational model that has been trained by performing the method according to the first aspect, and / or a data stream that carries data defining model parameters and an architecture of a likelihood-based computational model (the computational model having been trained by performing the method according to the first aspect).

[0025] In a seventh aspect, the present invention is directed to at least one computer (e.g., a computer) that includes at least one processor (e.g., a processor) and at least one memory (e.g., a memory), where the program according to the third aspect runs on the processor or is loaded into the memory, or where at least one computer includes a computer-readable program storage medium according to the fourth aspect.

[0026] In an eighth aspect, the present invention is directed to a system for determining a relationship between an anatomical structure represented in a two-dimensional medical image and an annotation of the anatomical structure, comprising:

[0027] a) at least one computer according to the preceding claims;

[0028] b) at least one electronic data storage device storing patient image data;

[0029] c) a program storage medium according to the preceding claims; and

[0030] wherein at least one computer is operatively coupled to:

[0031] - at least one electronic data storage device for obtaining the patient image data from the at least one electronic data storage device and for storing at least structure annotation prediction data in the at least one electronic data storage device; and

[0032] - the program storage medium for obtaining data defining model parameters and an architecture of a likelihood-based computational model from the program storage medium.

[0033] Alternatively or additionally, the present invention according to the fifth aspect is directed to a non-transitory computer-readable program storage medium, for example, that stores a program for causing a computer according to the fourth aspect to perform data processing steps of the method according to the first or second aspect.

[0034] For example, the present invention does not relate to or particularly does not include or contain invasive steps that represent a substantial physical interference with the body and require professional medical measures to be taken on the body, and where even with the necessary professional care and measures, the body may still be subject to significant health risks.

[0035] For example, the present invention does not include a step of applying ionizing radiation to a patient's body to generate, for example, patient image data. Instead, the patient image data has already been generated before performing the inventive method according to the second aspect. For this reason alone, no surgical or therapeutic activities, in particular no surgical or therapeutic steps, are required by implementing the present invention. More specifically, the present invention does not relate to or in particular does not include or comprise any surgical or therapeutic activities. Instead, the present invention may relate to being applicable to processing medical image data.

[0036] The present invention also relates to using a system according to the eighth aspect or a computer according to the seventh aspect for training a likelihood-based computational model to determine an image representation of an annotated anatomical structure in a two-dimensional X-ray image or to determine a relationship between an anatomical structure represented in a two-dimensional medical image and an annotation for that anatomical structure by respectively performing the method according to the first aspect or the second aspect.

[0037] Definition

[0038] Definitions of specific terms used in this disclosure are provided in this section and they also form part of this disclosure.

[0039] A method according to the present invention is, for example, a computer-implemented method. For example, all steps or only some steps (i.e., less than the total number of steps) of a method according to the present invention can be performed by a computer (e.g., at least one computer). An example of a computer-implemented method is the use of a computer to perform a data processing method. An example of a computer-implemented method is a method involving computer operations such that the computer is operated to perform one, several or all steps of the method.

[0040] A computer includes, for example, at least one processor and at least one memory, for example, to process data technically, for example, electronically and / or optically. The processor is made of, for example, a semiconductor material or composition, for example, at least partially n-type and / or p-type doped semiconductor, for example, at least one of type II, III, IV, V, VI semiconductor materials, for example, (doped) silicon arsenide and / or gallium arsenide. The described computing steps or determination steps are performed by a computer, for example. The determination steps or computing steps are, for example, steps of determining data within the framework of a technical method (for example, within the framework of a program). A computer is, for example, any type of data processing device, for example, an electronic data processing device. A computer can be a device generally regarded as a computer, for example, a desktop personal computer, a laptop, a netbook, etc., but can also be any programmable device, for example, a mobile phone or an embedded processor. A computer can, for example, include a "sub-computer" system (network), where each sub-computer represents a computer in its own right. The term "computer" includes cloud computers, for example, cloud servers. The term "computer" includes server resources. The term "cloud computer" includes cloud computer systems, which, for example, include a system of at least one cloud computer and, for example, include a plurality of operably interconnected cloud computers, such as server farms. Such cloud computers are preferably connected to a wide area network such as the World Wide Web (WWW) and are located in the so-called cloud of all computers connected to the World Wide Web. Such an infrastructure is used for "cloud computing", which describes those computing, software, data access, and storage services that do not require the end user to know the physical location and / or configuration of the computer providing a particular service. For example, the term "cloud" is used metaphorically in this context for the Internet (World Wide Web). For example, the cloud provides computing infrastructure as a service (IaaS). A cloud computer can be used as a virtual host for an operating system and / or data processing application for performing the method of the present invention. A cloud computer is, for example, provided by Amazon Web Services TM)Elastic Compute Cloud (EC2) is provided. The computer includes, for example, an interface for receiving or outputting data and / or performing analog-to-digital conversion. The data represents, for example, physical properties and / or data generated from technical signals. Technical signals are generated, for example, by (technical) detection devices (such as devices for detecting tagging devices) and / or (technical) analysis devices (such as devices for performing (medical) imaging methods), where the technical signals are, for example, electrical signals or optical signals. The technical signals represent, for example, the data received or output by the computer. The computer is preferably operatively coupled to a display device that allows the information output by the computer to be displayed to, for example, a user. An example of the display device is a virtual reality device or an augmented reality device (also known as virtual reality glasses or augmented reality glasses), which can be used as "goggles" for navigation. A specific example of such augmented reality glasses is Google Glass (a trademark brand under Google, Inc.). The augmented reality device or virtual reality device can be used both for inputting information into the computer through user interaction and for displaying the information output by the computer. Another example of the display device is, for example, a standard computer monitor including a liquid crystal display, which is operatively connected to a computer for receiving display control data from the computer for generating signals for displaying image information content on the display device. A specific embodiment of such a computer monitor is a digital light box. An example of such a digital light box is a product of Brainlab AG The monitor can also be, for example, a handheld portable device such as a smart phone or a personal digital assistant or a digital media player.

[0041] The invention also relates to a program that, when run on a computer, causes the computer to perform one, more, or all of the method steps described herein; and / or relates to a program storage medium storing the program (especially in a non-transitory form); and / or relates to a computer including the program storage medium; and / or relates to a (physical, for example electrical, for example technically generated) signal wave carrying information representing the program (such as the above-mentioned program), for example a digital signal wave, such as an electromagnetic carrier wave, the program including, for example, code means adapted to perform any or all of the method steps described herein.

[0042] Within the framework of the present invention, a computer program unit can be embodied by hardware and / or software (which includes firmware, resident software, microcode, etc.). Within the framework of the present invention, a computer program unit can take the form of a computer program product, which can be implemented by a computer-usable, e.g., computer-readable data storage medium, which includes computer-usable, e.g., computer-readable program instructions, and the "code" or "computer program" embodied in said data storage medium is for use on or in conjunction with an instruction execution system. Such a system can be a computer; a computer can be a data processing device including means for executing a computer program unit and / or program according to the present invention, e.g., a data processing device including a digital processor (central processing unit or CPU) for executing the computer program unit, and optionally including a volatile memory (e.g., random access memory or RAM) for storing data for executing the computer program unit and / or generated by executing the computer program unit. Within the framework of the present invention, a computer-usable, e.g., computer-readable data storage medium can be any data storage medium that can contain, store, communicate, propagate, or transport a program for use on or in conjunction with those instruction execution systems, devices, or apparatuses. A computer-usable, e.g., computer-readable data storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or a propagation medium such as the Internet. A computer-usable or computer-readable data storage medium can even be, for example, paper or other suitable medium on which the program can be printed, since the program can be captured electronically, e.g., by optically scanning the paper or other suitable medium and then compiled, decoded, or otherwise processed appropriately. The data storage medium is preferably a non-volatile data storage medium. The computer program products and any software and / or hardware described herein form various means for performing the functions of the present invention in example embodiments. A computer and / or data processing device can, for example, include a guidance information device, which includes means for outputting guidance information. The guidance information can be output to a user, for example, visually through a visual indication means (e.g., a monitor and / or a lamp) and / or auditorily through an auditory indication means (e.g., a speaker and / or a digital voice output device) and / or tactilely through a tactile indication means (e.g., a vibration element or a vibration element incorporated in an instrument). For the purposes of this document, a computer is a technical computer, which includes, for example, technical components such as tangible components, e.g., mechanical components and / or electronic components. Any device mentioned in this document is a technical device and is, for example, a tangible device.

[0043] The expression "acquire data" includes, for example, (within the framework of the computer-implemented method) scenarios where data is determined by a computer-implemented method or program. Determining data includes, for example, measuring a physical quantity and transforming the measured value into data, such as digital data, and / or calculating (e.g., outputting) the data by means of a computer and, for example, within the framework of the method according to the invention. The "determining" step as described herein includes, for example, issuing or consisting of an order to perform the determination described herein. For example, this step includes issuing or consisting of an order that causes a computer (such as a remote computer, such as a remote server, such as in the cloud) to perform the determination. Alternatively or additionally, the "determining" step as described herein includes, for example, the following steps or consists of them: receiving the result data of the determination described herein, for example, receiving the result data from a remote computer (such as from the remote computer that caused it to perform the determination). The meaning of "acquire data" also includes, for example, the following scenarios: receiving or retrieving data by a computer-implemented method or program (such as by input) from, for example, another program, a previous method step, or a data storage medium, for example, for further processing by the computer-implemented method or program. The generation of the data to be acquired may or may not be part of the method according to the invention. Thus, the expression "acquire data" may also, for example, mean waiting to receive data and / or receiving data. The received data may be input via an interface, for example. The expression "acquire data" may also mean that a computer-implemented method or program performs some steps to (actively) receive or retrieve data from a data source such as a data storage medium (such as ROM, RAM, database, hard disk drive, etc.) or via an interface (such as from another computer or network). The data acquired respectively by the method or device of the present disclosure can be obtained from a database located in a data storage device, which is operably connected to a computer for data transmission between the database and the computer, for example, data transmission from the database to the computer. The computer acquires data for use as an input to the "determine data" step. The determined data may be output again to the same or another database for storage for subsequent use. The database or the database used to implement the method of the present disclosure may be located in a network data storage device or a network server (such as a cloud data storage device or a cloud server) or a local data storage device (such as a mass storage device operably connected to at least one computer that performs the method of the present disclosure). The data can be brought into a "ready" state by performing additional steps before the acquisition step. According to this additional step, data is generated for acquisition. For example, detecting or capturing data (such as by an analysis device). Alternatively or additionally, according to the additional step, inputting data via an interface, for example. The generated data may be input (such as input into a computer), for example.According to an additional step (which is carried out before the acquisition step), data can also be provided by carrying out an additional step of storing data in a data storage medium (such as a ROM, RAM, CD, and / or hard disk drive), so that within the framework of the method or program according to the present invention, the data is made ready. Therefore, the step of "acquiring data" can also involve the command device acquiring and / or providing the data to be acquired. In particular, the acquisition step does not involve invasive steps, which represent substantial physical interference with the body and require professional medical measures, and even when carried out with the required professional care and measures, the body may be at significant health risk. In particular, the step of acquiring data, such as determining data, does not involve surgical steps, especially steps of treating the human or animal body using surgery or therapy. To distinguish different data used in this method, the data is represented as (i.e., called) "XY data", etc., and defined according to the information they describe, and then preferably called "XY information", etc.

[0044] Preferably, atlas data describing (such as defining, more particularly representing and / or as) the overall three-dimensional shape of a body anatomical part is acquired. Thus, the atlas data represents an atlas of the body anatomical part. An atlas typically consists of a plurality of object general models, and these object general models together form a composite structure. For example, the atlas constitutes a statistical model of a patient's body (such as a part of the body), which has been generated based on anatomical information collected from multiple human bodies, such as based on medical image data containing images of these human bodies. Therefore, in principle, the atlas data represents the statistical analysis result of such medical image data of multiple human bodies. This result can be output as an image - thus the atlas data contains or is equivalent to medical image data. Such a comparison can be carried out, for example, by applying an image fusion algorithm, where the image fusion algorithm performs image fusion between the atlas data and the medical image data. The comparison result can be a similarity measure between the atlas data and the medical image data. The atlas data includes image information (such as position image information), which can be matched (such as by applying an elastic or rigid image fusion algorithm) with, for example, the image information (such as position image information) contained in the medical image data, so that, for example, the atlas data is compared with the medical image data to determine the position of the anatomical structure in the medical image data corresponding to the anatomical structure defined by the atlas data.

[0045] Multiple human bodies (whose anatomical structures are used as inputs for generating atlas data) advantageously share at least one common characteristic such as gender, age, race, body measurements (e.g., height and / or weight), and pathological condition. Anatomical information describes, for example, the human body's anatomical structure and is extracted, for example, from medical image information about the human body. For example, an atlas of the femur may include the femoral head, femoral neck, body, greater trochanter, lesser trochanter, and lower limb as objects that together constitute a complete structure. For example, an atlas of the brain may include the telencephalon, cerebellum, diencephalon, pons, midbrain, and medulla oblongata as objects that together constitute a complex structure. One application of such an atlas is in medical image segmentation, where the atlas is matched to medical image data and the image data is compared to the matched atlas in order to assign points (pixels or voxels) of the image data to the objects of the matched atlas, thereby segmenting the image data into objects.

[0046] For example, atlas data includes information about body anatomical parts. This information is, for example, at least one of patient-specific, non-patient-specific, indication-specific, or non-indication-specific. Thus, the atlas data describes at least one of patient-specific, non-patient-specific, indication-specific, or non-indication-specific atlases. For example, the atlas data includes movement information indicating the degrees of freedom of movement of a body anatomical part relative to a given reference (e.g., another body anatomical part). For example, the atlas is a multi-modal atlas that defines atlas information for multiple (i.e., at least two) imaging modalities and contains mappings between the atlas information in different imaging modalities (e.g., mappings between all modalities) such that these atlases can be used to transform medical image information from its image depiction in a first imaging modality to its image depiction in a second imaging modality different from the first imaging modality, or to compare different imaging modalities to each other (e.g., match or register).

[0047] Movements for treating a body part are caused, for example, by movements referred to hereinafter as "vital activities". Reference is also made in this regard to EP2189943A1 and EP2189940A1, which have also been published as US2010 / 0125195A1 and US2010 / 0160836A1 respectively, in which these vital activities are discussed in detail. In order to determine the position of the body part to be treated, an analysis device such as an X-ray device, a CT device or an MRT device is used to generate an analysis image of the body (such as an X-ray image or an MRT image). For example, the analysis device is configured to perform a medical imaging method. The analysis device uses, for example, a medical imaging method and is, for example, a device for analyzing a patient's body (for example by using waves and / or radiation and / or energy beams, such as electromagnetic waves and / or radiation, ultrasonic waves and / or particle beams). The analysis device is, for example, a device that generates an image (for example, a two-dimensional or three-dimensional image) of the patient's body (and, for example, the internal structure and / or anatomical parts of the patient's body) by analyzing the body. The analysis device is used, for example, for medical diagnosis in radiology. However, it may be difficult to identify the body part to be treated in the analysis image. For example, it is easier to identify body parts that are indicative of changes in the position of the body part to be treated and, for example, of movements associated with the body part to be treated. The indicative body part is tracked such that the movement of the body part to be treated can be tracked based on a known correlation between changes in the position (for example, movement) of the indicative body part and changes in the position (for example, movement) of the body part to be treated. As an alternative or supplement to tracking the indicative body part, a marker detection device can be used to track a marker device (which can be used as an indicator and is therefore referred to as a "marker indicator"). The position of the marker indicator has a known (predetermined) correlation (for example, a fixed relative position) with the position of an indicator structure (such as the chest wall, for example true ribs or false ribs, or the diaphragm or intestinal wall, etc.) whose position changes due to vital activities.

[0048] In the medical field, imaging methods (also referred to as imaging modalities and / or medical imaging modalities) are used to generate image data (e.g., two-dimensional or three-dimensional image data) of the human anatomical structure (such as soft tissues, bones, organs, etc.). The term "medical imaging method" should be understood to mean an imaging method (advantageously device-based), such as so-called medical imaging modalities and / or radiological imaging methods, such as computed tomography (CT) and cone beam computed tomography (CBCT, such as volumetric CBCT), X-ray tomography, magnetic resonance tomography (MRT or MRI), conventional X-ray, ultrasound scanning and / or ultrasound examination, and positron emission tomography. For example, a medical imaging method is performed by an analysis device. Examples of medical imaging modalities applied by medical imaging methods are: radiography, magnetic resonance imaging, medical ultrasonography or ultrasound, endoscopy, elastography, tactile imaging, thermography, medical photography, and nuclear medicine functional imaging techniques, such as positron emission tomography (PET) and single photon emission computed tomography (SPECT) as documented in Wikipedia.

[0049] The image data thus generated is also referred to as "medical imaging data". The analysis device is used, for example, to generate image data in a device-based imaging method. The imaging method is used, for example, for the medical diagnosis of the body's anatomical structure to generate an image described by the image data. The imaging method is also used, for example, to detect pathological changes in the human body. However, some changes in the anatomical structure, such as pathological changes in the structure (tissue), may not be detectable and may, for example, be invisible in the images generated by the imaging method. Tumors represent an example of a change in the anatomical structure. If a tumor grows, it can be considered to represent an expanded anatomical structure. Such an expanded anatomical structure may not be detectable. For example, only a part of the expanded anatomical structure may be detectable. For example, early / late stage brain tumors are usually visible in an MRI scan when a contrast agent is infiltrated into the tumor. The MRI scan represents an example of an imaging method. In the case of an MRI scan of such a brain tumor, the signal enhancement (caused by the infiltration of the contrast agent into the tumor) in the MRI image is considered to represent a solid tumor mass. Thus, the tumor can be detected and, for example, distinguishable in the images generated by the imaging method. In addition to these tumors referred to as "enhanced" tumors, it is considered that approximately 10% of brain tumors are indistinguishable in the scan and, for example, invisible to the user observing the images generated by the imaging method.

[0050] Image fusion can be elastic image fusion or rigid image fusion. In the case of rigid image fusion, the relative positions between the pixels of a 2D image and / or the voxels of a 3D image are fixed, while in the case of elastic image fusion, the relative positions are allowed to change.

[0051] In the present application, the term "image deformation" is also used as an alternative to the term "elastic image fusion", but has the same meaning.

[0052] The elastic fusion transformation (e.g., elastic image fusion transformation) is designed, for example, to be able to achieve a seamless transition from one data set (e.g., the first data set, e.g., the first image) to another data set (e.g., the second data set, e.g., the second image). For example, the transformation is designed such that one of the first and second data sets (images) is deformed, e.g., in such a way that corresponding structures (e.g., corresponding image elements) are set at the same positions as the other of the first and second images. The deformed (transformed) image transformed from one of the first and second images is, for example, as similar as possible to the other of the first and second images. Preferably, a (numerical) optimization algorithm is applied to find the transformation that results in the best similarity. Preferably, the similarity is measured by a measure of similarity (also referred to hereinafter as "similarity measure"). The parameters of the optimization algorithm are, for example, the vectors of the deformation field. These vectors are determined by the optimization algorithm in such a way that the best similarity is obtained. Thus, the degree of best similarity represents a condition, e.g., a constraint, for the optimization algorithm. The basis of the vectors lies, for example, at the voxel positions of one of the first and second images to be transformed, and the tips of the vectors lie at the corresponding voxel positions in the image to be transformed. Preferably, a plurality of these vectors are provided, e.g., more than twenty or one hundred or one thousand or ten thousand, etc. Preferably, there are (other) constraints on the transformation (deformation), e.g., in order to avoid pathological deformations (e.g., by the transformation all voxels are moved to the same position). These constraints include, for example, the constraint that the transformation is regular, which, for example, means that the Jacobian determinant calculated from the matrix of the deformation field (e.g., vector field) is greater than zero, and also includes the constraints that the transformed (deformed) image is not self-intersecting and, for example, the transformed (deformed) image does not include defects and / or ruptures. The constraints include, for example, the following constraint: if a regular grid is transformed simultaneously with the image and in a corresponding manner, the grid is not allowed to interlace and fold at any of its positions. The optimization problem is solved, for example, iteratively by an optimization algorithm, which is, for example, a first-order optimization algorithm, e.g., a gradient descent algorithm. Other examples of optimization algorithms include optimization algorithms that do not use derivatives, such as the downhill simplex algorithm, or algorithms that use higher-order derivatives, such as quasi-Newton algorithms. The optimization algorithm preferably performs local optimization. If there are multiple local optima, global algorithms such as simulated annealing algorithms or genetic algorithms can be used. In the case of a linear optimization problem, for example, the simplex method can be used.

[0053] In the steps of the optimization algorithm, a voxel is moved by an amplitude in one direction, for example, such that the similarity increases. The amplitude is preferably less than a predetermined limit, such as less than one tenth or one hundredth or one thousandth of the image diameter, and is, for example, approximately equal to or less than the distance between adjacent voxels. For example, due to a large number of (iterative) steps, large deformations can be implemented.

[0054] The determined elastic fusion transformation can be used, for example, to determine the similarity (or similarity measure, see above) between a first data set and a second data set (a first image and a second image). To this end, the deviation between the elastic fusion transformation and the identity transformation is determined. The degree of deviation can be calculated, for example, by determining the difference between the determinant of the elastic fusion transformation and the identity transformation. The greater the deviation, the lower the similarity, and thus the degree of deviation can be used to determine the measure of similarity.

[0055] The measure of similarity can be determined, for example, based on the determined correlation between the first data set and the second data set. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In the following, the present invention is described with reference to the accompanying drawings, which give a background explanation of the present invention and show specific embodiments of the present invention. However, the scope of the present invention is not limited to the specific features disclosed in the context of the drawings. In the figures:

[0057] Figure 1 Illustrates the basic flow of the method according to the first aspect;

[0058] Figure 2 Illustrates the basic flow of the method according to the second aspect;

[0059] Figure 3 Shows an example of the method according to the first aspect;

[0060] Figure 4 Shows the principle of using Figure 3 the example of;

[0061] Figure 5 is a schematic diagram of a system according to the fifth aspect; and

[0062] Figure 6 Shows the structure of a single neuron of a convolutional neural network. DETAILED DESCRIPTION

[0063] Figure 1 Illustrates the basic steps of the method according to the first aspect, wherein step S101 includes obtaining image training data, step S102 includes obtaining annotation data, and the subsequent step S103 includes determining model parameter data.

[0064] Figure 2Describe the basic steps of the method according to the second aspect, where step S104 includes obtaining patient image data, and step S105 includes determining structural annotation data.

[0065] Figure 3 Illustrate an example of the method according to the first aspect. In step S21, patient image data embodied by a medical data set is read, and in a subsequent step S22, an intensity threshold is found, which is used to identify a threshold for rendering a grayscale representation of bone tissue (e.g., a predetermined Hounsfield unit value). Based on this threshold, step S23 proceeds to define an intensity mapping function, including rendering parameters read in step S24. Then, in step S25, density / intensity mapping is used to generate a two-dimensional DRR from the medical data set. Then, in step S26, the DRR is used together with annotations (i.e., at least one annotation) to form the basis of a training data item that can be used to train a machine learning algorithm in step S28. Additionally, in step S27, annotated clinical data (such as a real fluoroscopy) can optionally be read and used as the basis for training the machine learning algorithm. The annotation of the DRR can be read from metadata associated with the medical data set or can be generated by using atlas data. To this end, step S29 can read an identifier such as the name of a relevant anatomical structure and use it to segment the medical data set based on the atlas data. Additionally, the atlas data can store projection parameters for generating the DRR, which can be extracted in step S211 and input into step S25. The segmented anatomical structure is projected from three dimensions to two dimensions in step S212, and in step S213, the two-dimensional projection is used to generate an annotation that can be associated with the DRR generated in step S25 by using two-dimensional coordinates associated with the annotation.

[0066] Figure 4 Illustrate how in step S30, three-dimensional data representing patient image data is read and used as an input for generating synthetic data in step S31, which can then be combined with atlas data read in step S33 and clinical data read in step S32 as an input to a trained AI model, and then the trained AI model is run on this input to generate an information output (such as a bounding box, the probability of an image component representing a certain anatomical structure, or key points, or landmark localization) in step S35, or to determine the image position of a predetermined anatomical structure (such as a single vertebra) in step S36.

[0067] Figure 5 Is a schematic diagram of a medical system 1 according to the eighth aspect. The system is overall labeled with reference numeral 1 and includes a computer 2 and an electronic data storage device (such as a hard disk) 3 for storing at least patient image data. The components of the medical system 1 have the functions and characteristics explained above with respect to the eighth aspect of the present disclosure.

[0068] The focus of the method disclosed according to the first aspect is to train a machine learning algorithm to detect objects and / or features in patient image data (such as 3D data (such as CT or MRI) or 2D data (such as X-ray images or fluoroscopy)). In addition to using real patient data, digitally reconstructed radiographs are also used to optimize the training, as this allows for better fine-tuning of the input data.

[0069] The benefit of using DRRs is that a large number of images can be generated, and the image quality and their content of these images may be affected. The ability to generate DRRs from CT datasets means that the algorithm can be used for 3D datasets and 2D images. By adjusting the bone threshold (which is automatically detected by applying known methods), the content of the output image can be adjusted so that it shows, for example, only the bone structure and not the soft tissue. The projection parameters can be freely defined, allowing images to be generated from various shooting directions that would be difficult to achieve in a clinical environment (and may not be achievable at all due to ethical issues / radiation dose / availability of clinical samples).

[0070] Hereinafter, reference Figure 6 is made to explain the convolutional neural network as an example of a machine learning algorithm used in combination with the invention of the present disclosure.

[0071] A convolutional network (also known as a convolutional neural network, or CNN) is an example of a neural network used to process data with a known grid-like topology. Examples include time series data (which can be regarded as a one-dimensional grid sampled at regular time intervals) and image data (which can be regarded as a two-dimensional grid of pixels). The name "convolutional neural network" indicates that the network employs the mathematical operation of convolution. Convolution is a linear operation. A convolutional network is a simple neural network that uses convolution instead of general matrix multiplication in at least one of its layers. There are various variants of the convolution function, which are widely used in neural networks in practice. Generally, the operations used in convolutional neural networks do not exactly correspond to the definition of convolution used in other fields (such as engineering or pure mathematics).

[0072] The main component of a convolutional neural network is the artificial neuron. Figure 6 is an example of a depicted single neuron. The middle node represents the neuron, which receives all the inputs (x1,…,x n ), and multiplies them by their specific weights (w1,…,w n ). The importance of an input depends on the value of its weight. The addition of these calculated values is called the weighted sum, which will be inserted into the activation function. The weighted sum z is defined as:

[0073]

[0074] The bias b is a value independent of the input, which modifies the boundaries of the threshold. The resulting value is processed by an activation function that determines whether to pass the input to the next neuron.

[0075] CNNs typically take a 1D or 3D tensor as their input, for example, an image with H rows, W columns, and 3 channels (R, G, B color channels). However, a CNN can process higher-order tensor inputs in a similar manner. The input then continues through a series of processes. One processing step is typically referred to as a layer, which can be a convolutional layer, a pooling layer, a normalization layer, a fully connected layer, a loss layer, etc. The details of these layers are described in the following sections.

[0076]

[0077] Equation 5 above illustrates the way a CNN operates layer by layer during the forward pass. The input is x 1 , which is typically an image (a 1D or 3D tensor). The parameters involved in the first layer processing are collectively referred to as the tensor w 1 . The output of the first layer is x 2 , which also serves as the input for the second layer processing. This processing continues until all layers in the CNN have been processed, and its output is x L . However, an additional layer for backpropagation of errors is added, which is a method for learning good parameter values in a CNN. Assume the current problem is an image classification problem of C classes. A common strategy is to output x L as a C-dimensional vector, whose i-th entry encodes the prediction (the posterior probability of x 1 comes from the i-th class). To make x L a probability mass function, the processing in the (L-1)-th layer can be set to the softmax transformation of x L-1 (comparing the distance measure with the data transformation node). In other applications, the output x L can have other forms and interpretations. The last layer is the loss layer. Assume t is the corresponding target value (ground truth) of the input x 1 , then a cost or loss function can be used to measure the difference between the CNN prediction x L and the target t. It should be noted that some layers may not have any parameters, that is, for some i, w i may be empty.

[0078] In the example of a CNN, ReLu is used as the activation function for the convolutional layer, while the softmax activation function provides information to give a classification output. The following sections will illustrate the purpose of the most important layers.

[0079] The input image is fed into a feature learning section that includes layers of convolution and ReLu, followed by a layer that includes pooling, which is then followed by further pairwise repetition of layers of convolution and ReLu and layers of pooling. The output of the feature learning section is fed into a classification section that includes layers for flattening, fully connecting, and max softening.

[0080] In a convolutional layer, multiple convolutional kernels are typically used. Assuming that D kernels are used and the spatial extent of each kernel is H×W, all the kernels are represented as f. f is a 4th-order tensor in l <. Similarly, index variables 0 ≤ i < H, 0 ≤ j < W, 0 ≤ d l < D L and 0 ≤ d < D are used to determine a particular element in the kernel. It should also be noted that the kernel set f refers to the same object as the symbol w in Equation 5 (see the architecture section). The notation is slightly changed to simplify the derivation process. It is also clear that the kernels remain unchanged even when the mini-batch strategy is used.

[0081] As long as the convolutional kernel is larger than 1×1, the spatial extent of the output is smaller than that of the input. Sometimes it is necessary for the input and output images to have the same height and width, and a simple padding trick can be used. If the input is H l ×W l ×D l , and the kernel size is H×W×D l ×D, the convolutional result has the size (H l - H + 1) × (W l - W + 1) × D.

[0082] For each input channel, if rows are padded (i.e., inserted) above the first row, rows are padded (i.e., inserted) below the last row, and columns are padded to the left of the first column, columns are padded to the right of the last column, the size of the convolutional output will be H ×W l ×D, i.e., having the same spatial extent as the input. b·c is the floor function. The elements of the padded rows and columns are usually set to 0, but other values are possible. l The stride is another important concept in convolution. The kernel is convolved with the input at every possible spatial extent, which corresponds to a stride s = 1. However, if s > 1, each movement of the kernel skips s - 1 pixel positions (i.e., convolution is performed every s pixels in the horizontal and vertical directions).

[0083] In this section, the simple case of stride 1 and no padding is considered. Thus, in

[0084] ​ There is y (or x l+1 ), where H l+1 =H l -H+1,W l+1 =W l -W+1, and D l+1 = D. In precise mathematics, the convolution process can be expressed as an equation:

[0085]

[0086] For all 0≤d≤D=D l+1 , and satisfy 0≤i l+1 <H l –H+1=H l+1 , 0≤j l+1 <W l –W+1=W l+1 Any spatial position (i l+1 ,j l+1 ) Repeat Equation 15. In this equation, Refers to the triple (i l+1 +i,j l+1 +j,d l ) indexed by x l elements. Usually the bias term b d Add to For greater clarity, this term is omitted.

[0087] Pooling functions replace the network output at a location with summary statistics of nearby outputs. For example, the max pooling operation reports the maximum output within a rectangular neighborhood of a table. Other popular pooling functions include the average of a rectangular neighborhood, the L2 norm of a rectangular neighborhood, or a weighted average based on the distance to the center pixel. In all cases, pooling helps make the representation approximately invariant to small translations of the input. Translation invariance means that if the input is translated by a small amount, the value of the pooled output will not change.

[0088] Since pooling aggregates responses over an entire neighborhood, it is possible to use fewer pooling units than detector units by reporting summary statistics for regions that are k pixels apart instead of one pixel. This improves the computational efficiency of the network because the next layer has approximately k times fewer inputs to process.

[0089] Assume that the CNN model w has been learned 1 ,…,w L-1 If we know all the parameters of , we can use the model to make predictions. Prediction simply involves running the CNN model in the forward direction, i.e., in the direction of the arrow in Equation 5 (see the Architecture section). Take the image classification problem as an example. From the input x 1Start by passing it through the processing of the first layer (the box with parameter w 1 ), and obtain x 2 . Then pass x 2 to the second layer in sequence, and so on. Finally, receive its estimated x 1 whose posterior probability belongs to class C. The CNN can output the prediction as:

[0090]

[0091] The problem at this time is: how to learn the model parameters?

[0092] Just like in many other learning systems, the parameters of the CNN model are optimized to minimize the loss z, that is, it is desired that the predictions of the CNN model match the ground truth labels. Assume a training example x 1 is given to train such parameters. The training process involves running the CNN network in two directions. First, run the network in the forward pass to obtain x L , to make a prediction using the current CNN parameters. Instead of outputting the prediction, the prediction needs to be compared with the target t corresponding to x 1 , that is, continue running the forward pass until the last loss layer. Finally, obtain the loss z. The loss z is a supervision signal that guides how to correct (update) the model parameters.

[0093] There are several algorithms for optimizing the loss function, and the CNN is not limited to a specific algorithm. An example algorithm is called Stochastic Gradient Descent (SGD). This means that the parameters are updated by using the gradients estimated from a (usually) small subset of the training examples.

[0094]

[0095] In Equation 3, the ← symbol implicitly indicates that the parameter w i of (layer i) is updated from time t to t + 1. If the time index t is explicitly used, the equation will be written as:

[0096]

[0097] In Equation 4, the partial derivative measures the growth rate of z with respect to the different dimensional changes of w i . This partial derivative vector is called the gradient in mathematical optimization. Therefore, in a small local area near the current value of w i , moving w i in the direction determined by the gradient will increase the objective value z. To minimize the loss function, w i should be updated in the opposite direction of the gradient. This update rule is called gradient descent.

[0098] However, if we move too far in the negative gradient direction, the loss function may increase. Therefore, in each update, we change the parameters only by a small fraction of the negative gradient (controlled by η (learning rate)). Usually, η > 0 is set to a small number (e.g., η = 0.001). If the learning rate is not too high, one update based on x 1 will make the loss for this particular training example smaller. However, it is very likely to make the loss for some other training examples larger. Therefore, we need to use all training examples to update the parameters. When all training examples have been used to update the parameters, it is said that one learning cycle has been processed. Usually, one learning cycle will reduce the average loss of the training set until the learning system fits the training data. Therefore, we can repeat the gradient descent updates for learning cycles and terminate at a certain point to obtain the CNN parameters (e.g., when the average loss of the validation set increases, we can terminate).

[0099] The partial derivatives of the last layer are easy to calculate. x L is directly connected to z under the control of the parameter w L , so it is easy to calculate This step needs to be performed only when w L is not empty. Similarly, it is also easy to calculate For example, if we use the squared L2 loss, then is empty, and

[0100] In fact, for each layer, we calculate two sets of gradients: the partial derivative of z with respect to the layer parameter w i , and the input x i of this layer. As shown in Equation 3, the term can be used to update the parameters of the current (i-th layer). The term can be used to update the parameters backward, for example, to the (i - 1)-th layer. The intuitive explanation is that: x i is the output of the (i - 1)-th layer, and is the way how x i should be changed to reduce the loss function. Therefore, we can regard as part of the "error" regulatory information propagated backward from z to the current layer layer by layer. Therefore, we can continue the backpropagation process and use to backpropagate the error to the (i - 1)-th layer. This layer-by-layer backward update procedure can greatly simplify learning the CNN.

[0101] Take the i-th layer as an example. When updating the i-th layer, the backpropagation process of the (i + 1)-th layer must have been completed. That is, the terms and both have been calculated and stored in the memory and are ready for use. The task at this time is to calculate and Using the chain rule, we get:

[0102]

[0103] Since has already been calculated and stored in the memory, only matrix reshaping operation (vec) and an additional transpose operation are needed to obtain This is the first term in the right - hand side (RHS) of the two equations. As long as we can calculate and we can easily obtain the expected values (the left - hand side of the two equations).

[0104] and are much easier to calculate than directly calculating and because x i is directly related to x i through a function with parameter w i+1 directly.

[0105] In the context of neural networks, activations act as transfer functions between the inputs and outputs of neurons. They define under which conditions a node is activated, i.e., map the input values to an output that is then used as one of the inputs to subsequent neurons in the hidden layer. There are a large number of different activation functions with different characteristics.

[0106] The loss function quantifies how well the algorithm models the given data. To learn from the data and change the weights of the network, the loss function must be minimized (see 2.6). Generally, a distinction can be made between regression loss and classification loss. In classification, one tries to predict an output from a finite set of categorical values (class labels), while in regression, continuous values are predicted.

[0107] In the following mathematical formulas, the following parameters are defined as:

[0108] n is the number of training examples;

[0109] i is the i - th training example in the dataset;

[0110] y i is the ground - truth label of the i - th training example;

[0111] is the prediction for the i - th training example.

[0112] The most common setting for classification problems is the cross-entropy loss. It increases as the predicted probability deviates from the actual label. The logarithm of the actual predicted probability is multiplied by the ground truth class. An important aspect is that the cross-entropy loss severely penalizes confident but incorrect predictions. The mathematical formula can be described as:

[0113]

[0114] A typical example of regression loss is the mean squared error or L2 loss. As the name implies, the mean squared error is the average of the squared differences between the predicted values and the actual observed values. It only involves the average error magnitude and is independent of their direction. However, due to squaring, predictions that are far from the actual values suffer more severely compared to predictions with smaller deviations. Additionally, the MSE has good mathematical properties that make it easier to calculate the gradient. Its formula is as follows:

[0115]

[0116] For information on the functionality of convolutional neural networks, please refer to the following literature:

[0117] "Deep learning, chapter convolutional networks" by I. Goodfellow, Y. Bengio, and A. Courville, 2016, see http: / / www.deeplearningbook.org ;

[0118] "Introduction to convolutional neural networks" by J. Wu, see https: / / pdfs.semanticscholar.org / 450c / a19932fcef1ca6d0442cbf52fec38fb9d1e5.p df ;

[0119] "Common loss functions in machine learning", see

[0120] https: / / towardsdatascience.com / common-loss-functions-in-machine- learning-46af0ffc4d23 , last accessed: 2019-08-22;

[0121] "Imagenet classification with deep convolutional neural networks" by Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton, see

[0122] http: / / papers.nips.cc / paper / 4824-imagenet-classification-with-deep- convolutional-neural-networks.pdf ;

[0123] "Faster r-cnn: Towards real-time object detection with region proposal networks" by S. Ren, K. He, R. Girshick, and J. Sun, see

[0124] https: / / arxiv.org / pdf / 1506.01497.pdf ;

[0125] "Convolutional pose machines" by S.-E. Wei, V. Ramakrishna, T. Kanade, and Y. Sheikh, see https: / / arxiv.org / pdf / 1602.00134.pdf ;

[0126] "Fully convolutional networks for semantic segmentation" by Jonathan Long, Evan Shelhamer, and Trevor Darrell, see https: / / www.cv-foundation.org / openaccess / content_cvpr_2015 / papers / Long_Fully_Convolutional_Net works_2015_ CVPR_paper.pdf 。

Claims

1. A computer-implemented method for training a machine learning algorithm to determine the position of an image representation of an annotated anatomical structure in a two-dimensional X-ray image, the method comprising the steps of: a) obtaining image training data (S101), the image training data describing a synthesized two-dimensional X-ray image containing an image representation of the anatomical structure; b) obtaining annotation data (S102), the annotation data describing an annotation of the anatomical structure; and c) determining model parameter data (S103), the model parameter data describing model parameters of a machine learning algorithm for establishing a relationship between the anatomical structure in the two-dimensional X-ray image and the annotation, wherein the model parameter data is determined by inputting the image training data and the annotation data into a function for establishing the relationship; characterized in that d) obtaining atlas data, the atlas data describing the anatomical structure, and determining the annotation data based on the atlas data and the image training data.

2. The method according to claim 1, wherein, The annotation data is determined from metadata included in the image training data.

3. The method according to claim 1 or 2, wherein, The function establishes a relationship between the position of the anatomical structure in the two-dimensional X-ray image and the position for displaying the annotation in the two-dimensional X-ray image.

4. The method according to any one of the preceding claims, comprising the steps of: obtaining medical image data describing a three-dimensional medical image containing an image representation of the anatomical structure, wherein the image training data is determined by determining an image value threshold associated with the image representation of the anatomical structure in the three-dimensional medical image, defining a corresponding intensity mapping function, and generating the image representation of the anatomical structure in each synthesized two-dimensional X-ray image based on the intensity mapping function from the image representation of the anatomical structure in one of the three-dimensional medical images.

5. The method according to claim 4, wherein the atlas data describes at least one projection parameter for generating the two-dimensional X-ray image; generating the two-dimensional X-ray image based on the at least one projection parameter; the annotation data is determined by segmenting the anatomical structure from the medical image data based on the atlas data; the annotation data is determined by projecting the segmented anatomical structure from three dimensions to two dimensions and using the two-dimensional projection to generate an annotation.

6. The method according to any one of the preceding claims, wherein, A convolutional neural network is part of the machine learning algorithm.

7. The method according to any one of the preceding claims, wherein, The model parameters define the learnable parameters of the machine learning algorithm.

8. A computer-implemented method for determining a relationship between an anatomical structure represented in a two-dimensional medical image and an annotation of the anatomical structure, the method comprising the steps of: a) obtaining patient image data (S104), the patient image data describing a two-dimensional X-ray image containing an image representation of the patient's anatomical structure; and b) Determine structural annotation prediction data (S105), the structural annotation prediction data describing the position of the image representation of the anatomical structure in the two-dimensional X-ray image described by the patient image data and the annotation of the anatomical structure according to a certain likelihood determined by the machine learning algorithm, wherein determining the structural annotation data is by inputting the patient image data into a function that establishes a relationship between the image representation of the anatomical structure in the two-dimensional X-ray image and the annotation of the anatomical structure, and the function is part of a machine learning algorithm that has been trained by performing the method according to any one of claims 1 to 7.

9. The method according to claim 8, wherein, The patient image data has been generated by synthesizing the two-dimensional X-ray image from a three-dimensional image of the anatomical structure, or wherein the patient image data has been generated by applying a perspective imaging mode of the anatomical structure.

10. The method according to claim 8 or 9, wherein The function establishes a relationship between the position of the anatomical structure in the two-dimensional X-ray image described by the patient image data and the position for displaying the annotation in the two-dimensional X-ray image described by the patient image data, wherein the structural annotation data describes the relationship between the position of the image representation of the anatomical structure in the two-dimensional X-ray image described by the patient image data and the position for displaying the annotation of the anatomical structure in the two-dimensional X-ray image described by the patient image data.

11. A program storage medium, the program stored on the program storage medium, when running on at least one processor of a computer (2) or when loaded into the memory of the computer (2), causes the at least one processor of the computer (2) to execute the method steps of the method according to any one of claims 1 to 10.

12. A system (1) for determining the relationship between an anatomical structure represented in a two-dimensional medical image and an annotation of the anatomical structure, comprising: a) at least one computer (2); b) at least one electronic data storage device (3) storing patient image data; c) the program storage medium according to claim 11; wherein the at least one computer (2) is operably coupled to the at least one electronic data storage device (3) for obtaining the patient image data from the at least one electronic data storage device (3) and for storing at least structural annotation prediction data in the at least one electronic data storage device (3); The at least one computer (2) is operably coupled to the program storage medium for obtaining data defining the model parameters and the architecture of the machine learning algorithm from the program storage medium.

Citation Information

Patent Citations

  • Calculation of indicator body elements and pre-indicator trajectories

    EP2189940A1

  • Detection of vitally moved regions in an analysis image

    EP2189943A1

  • Determination of regions of an analytical image that are subject to vital movement

    US20100125195A1

  • Determination of indicator body parts and pre-indicator trajectories

    US20100160836A1