Imaging to depict a clothed person

A trained MLM predicts body shape information from image data to automate imaging parameter determination, addressing the challenge of manual estimation in clothing-covered individuals and enhancing imaging system accuracy.

DE102024204448B3Active Publication Date: 2025-10-30SIEMENS HEALTHINEERS AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102024204448
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-05-14
Publication Date
2025-10-30
Estimated Expiration
2044-05-14

AI Technical Summary

Technical Problem

Current imaging systems struggle to accurately estimate body shape information of a person wearing clothing, necessitating manual operator input and leading to suboptimal imaging configurations.

Method used

A computer-implemented method using a trained machine learning model (MLM) to predict body shape information from image data of a person wearing clothing, allowing for automated determination of imaging parameters.

Benefits of technology

Enhances the accuracy of imaging system configuration, reducing manual estimation errors and improving imaging quality by directly predicting body shape information from clothing-covered individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000016_0000
    Figure 00000016_0000
  • Figure 00000017_0000
    Figure 00000017_0000
  • Figure 00000017_0001
    Figure 00000017_0001
Patent Text Reader

Abstract

In a computer-implemented method for parameterizing an imaging system for imaging a clothed person (6), image data (11) of the clothed person (6) are obtained and body shape information of the person is determined by applying a trained machine learning model (MLM) (12) to the image data (11). At least one imaging parameter (13) of the imaging system is determined depending on the body shape information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a computer-implemented method for parameterizing an imaging system for imaging a clothed person and an associated computer-implemented training method for a machine learning model (MLM) for predicting body shape information of a person based on image data of the clothed person. The invention further relates to a data processing system for carrying out such computer-implemented methods or training methods, an imaging device with such a data processing system, a corresponding imaging method for imaging a clothed person, and a corresponding computer program product.

[0002] When medical images, such as X-rays, MRI scans, or PET scans, need to be acquired from clothed individuals, for example in emergencies or when undressing is not possible or desired for other reasons, information about body shape that is necessary or advantageous for configuring the imaging system is missing.

[0003] In current clinical practice, operators manually configure the imaging system, estimating body position and other body shape information by palpating the patient. While it's conceivable to use camera images of the person for a rough pre-initialization of the imaging system, the patient's clothing, which cannot always be removed, particularly in trauma areas, prevents an accurate automated estimation of body shape information.

[0004] In the publication X. Zou et al.: “CLOTH4D: A Dataset for Clothed Human Reconstruction.”, Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition 2023, CLOTH4D is presented, a dataset of clothed humans containing 1,000 subjects with diverse appearances, 1,000 3D outfits, and over 100,000 meshes of clothed and unclothed individuals. By evaluating and retraining clothed human reconstruction methods, new insights could be gained and performance improved.

[0005] In the publication R. Vidaurre et al.: “Fully Convolutional Graph Neural Networks for Parametric Virtual Try-On”, Computer Graphics Forum, Proc. of ACM SIGGRAPH Symposium on Computer Animation 2020, a learning-based approach for virtual clothing try-on is presented, based on a convolutional neural graph network. This network can handle a large family of garments represented as parametric, predefined 2D panels with arbitrary mesh topology, including long dresses, shirts, and fitted tops.

[0006] The publication by H. Zhang et al., “CloSET: Modeling Clothed Humans on Continuous Surface with Explicit Template Decomposition,” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition 2023, describes the creation of animatable avatars from static scans. This requires modeling clothing deformations in different poses. Point-based solutions are used, and it is proposed to decompose explicit clothing-related templates and then add position-dependent folds to them.

[0007] In the publication J. Zhu et al.: “Unpaired image-to-image translation using cycle-consistent adversarial networks.”, Proceedings of the IEEE international conference on computer vision 2017, an approach to image-to-image translation is presented. This is a class of image processing and graphics problems where the goal is to learn the mapping between an input image and an output image using a training set of matched image pairs. According to the proposed approach, it is possible to learn to translate an image from a source range X to a target range Y when no paired samples are available. This involves learning a mapping G : X → Y using adverse loss, such that the distribution of images from G(X) is indistinguishable from the distribution Y.

[0008] The Unity Engine (https: / / github.com / nielsdos / UnityClothSimulation.git, accessed April 22, 2024) is simulation software for simulating clothing fabrics. The Unreal Engine uses the Chaos Cloth Solver to simulate clothing (https: / / dev.epicgames.com / documentation / en-us / unreal-engine / clothing-tool-in-unreal-engine?application_version=5.2, accessed April 22, 2024).

[0009] Document US 2019 / 0057521A1 describes a method for predicting topograms from surface data in a medical imaging system. The topogram represents an internal organ of the patient as a projection by the patient. A sensor detects a patient's external surface. An image processor generates the topogram using a machine-trained generative adversarial network based on the surface data. The surface data is derived from an output of the external surface sensor. A display device shows the topogram. The external surface is the patient's skin or includes clothing. An image processor can configure a medical scanner based on the topogram, or the medical scanner can configure itself.

[0010] It is an object of the present invention to automatically estimate body shape information of a clothed person with higher accuracy.

[0011] This problem is solved by the respective subject matter of the independent claims. Advantageous further developments and preferred embodiments are the subject matter of the dependent claims.

[0012] The invention is based on the idea of ​​determining body shape information of a person, particularly in an unclothed state, based on image data of the person in a clothed state using a correspondingly trained machine learning model.

[0013] According to one aspect of the invention, a computer-implemented method for parameterizing an imaging system for imaging a clothed person is provided. Image data of the clothed person are obtained, and body shape information of the person is determined by applying a trained machine learning model (MLM) to the image data. At least one imaging parameter of the imaging system is determined based on the body shape information.

[0014] Unless otherwise specified, all steps of the computer-implemented method can be performed by a data processing system comprising at least one data processing device. In particular, the at least one data processing device is configured or adapted to perform the steps of the computer-implemented method. For this purpose, the at least one data processing device may, for example, store a computer program containing instructions which, when executed by the at least one data processing device, cause the at least one data processing device to perform the computer-implemented method. The computer-implemented method may also be implemented wholly or partly in hardware. The terms "data processing system" and "at least one data processing device" may be used interchangeably here and in the following. This also applies to corresponding derivatives.

[0015] In the event that the at least one data processing device comprises two or more data processing devices, certain steps performed by the at least one data processing device can also be understood as different data processing devices performing different steps or different parts of a step. In particular, it is not necessary for each data processing device to perform the steps. In other words, the execution of the steps can be distributed among the two or more data processing devices.

[0016] Each embodiment of the computer-implemented method results in a corresponding embodiment of a method for parameterizing an imaging system that is not purely computer-implemented, by including corresponding steps for generating the image data.

[0017] In general terms, a trained MLM can replicate cognitive functions that connect people with other human minds. Specifically, through training based on training data, the MLM can adapt to new circumstances and detect and extrapolate patterns. Another term for a trained MLM is "trained function."

[0018] In general, the parameters of a multi-level marketing (MLM) system can be adjusted or updated through training. This can involve supervised training, semi-supervised training, unsupervised training, reinforcement learning, and / or active learning. Furthermore, representational learning, also known as feature learning, can be employed. Specifically, the parameters of MLMs can be adjusted iteratively through multiple training steps. In particular, a specific loss function, also referred to as the cost function, can be minimized during training. When training an artificial neural network (ANN), the backpropagation algorithm can be used.

[0019] An MLM can, in particular, include an ANN, a support vector machine, a decision tree, and / or a Bayesian network, and / or the MLM can be based on k-means clustering, Q-learning, genetic algorithms, and / or association rules. Specifically, an ANN can be or include a deep neural network, a convolutional neural network (CNN), or a convolutional deep neural network. Furthermore, an ANN can be an adversarial network, a deep adverse network, and / or a generative adverse network (GAN).

[0020] In this case, the MLM is trained such that it can predict body shape information based on image-dependent input data, or that it can predict an output from which body shape information can be directly derived, such as a corresponding body model. The input data can either include the image data or be calculated based on it. For example, the input data can be calculated by encoding the image data, perhaps using another trained MLM. The input data is then provided by the encoded image data.

[0021] The image data depicts the person, who is to be imaged by the imaging system, while clothed, so that the person's body shape is not visible or not fully visible. The body shape information refers to the body shape of the same person while unclothed.

[0022] The term parameterization can be understood, in particular, as the determination of at least one imaging parameter, i.e., the determination of a corresponding value for each imaging parameter of the at least one imaging parameter. Which parameters constitute the at least one imaging parameter is predetermined and depends on the specific application, i.e., in particular on the type of imaging system and, if applicable, the type of imaging procedure to be performed.

[0023] For example, the imaging system is an X-ray-based imaging system, such as an X-ray angiography system, a C-arm X-ray system, an X-ray tomosynthesis system, or a computed tomography system. The at least one imaging parameter can then include, for example, at least one exposure setting of an X-ray source of the imaging system and / or a detector gain of an X-ray detector of the imaging system and / or a collimator position of a collimator of the imaging system and / or a size of a collimator aperture of the collimator and / or a shape of a collimator aperture of the collimator.

[0024] At least one exposure setting can include, for example, a peak kilovoltage (kVp) and / or a tube current of the X-ray source and / or a pulse duration of the X-ray pulses emitted by the X-ray source.

[0025] In other exemplary applications, the imaging system is a magnetic resonance imaging (MRI) system, and at least one imaging parameter includes a target imaging area within a patient tube of the MRI system. The same applies, for example, to other imaging systems, such as positron emission tomography (PET) systems.

[0026] If at least one imaging parameter has been determined as a result of the computer-implemented method according to the invention, the imaging system can be configured according to the determined at least one imaging parameter. The clothed person can then be imaged using the imaging system configured in this way.

[0027] According to the invention, the appropriately trained MLM is used to predict body shape information of the person from the image data that is not directly discernible to a human observer, and to derive the corresponding imaging parameters of the imaging system based on this prediction. It is therefore no longer necessary to manually scan the person's body to approximate their body shape. Furthermore, the MLM delivers more accurate results, leading to a better determination of the imaging parameters and ultimately to improved image quality. It has been shown that ANNs, such as GANs, CNNs, and transformer networks, are particularly well-suited as MLMs in this context.

[0028] Using the at least one imaging parameter determined according to the invention, the imaging system can then be configured, in particular automatically or partially automatically, according to the determined at least one imaging parameter. This allows the operation of the imaging system to be further automated.

[0029] Depending on the body shape information, secondary body information of the person is estimated and at least one imaging parameter is determined depending on the secondary body information.

[0030] Secondary body information relates in particular to the internal structure of the body. While this is not body shape information, it can be determined at least approximately based on it, which is why it is referred to as secondary information here and in the following. Such secondary body information can have a significant influence on the choice of at least one imaging parameter, making such embodiments particularly advantageous.

[0031] Secondary body information includes a person's body weight and / or the material composition of their body. Material composition can be, for example, the mass or volume ratio of one material in the person's body to another, such as bone tissue to fat tissue, water to fat tissue, muscle tissue to fat tissue, and so on. Material composition can also include body fat percentage, muscle mass, or similar data.

[0032] According to at least one embodiment, the image data includes a two-dimensional or two-and-a-half-dimensional image, in particular a camera image, of the clothed person.

[0033] The image of the clothed person can therefore be captured using a suitable camera if the person is to be imaged by the imaging system. Accordingly, the generation of the image data can be easily integrated into the clinical workflow. Furthermore, it is known, for example from the publications discussed at the beginning, that camera images are very well suited for the corresponding projection using MLMs.

[0034] A two-dimensional image depicts a scene in two dimensions. Each pixel in a two-dimensional arrangement of pixels has a corresponding intensity value. However, a two-dimensional image can also contain multiple channels, particularly color channels. In contrast, a two-and-a-half-dimensional image includes, for each pixel, not only an intensity value but also a depth value, which indicates the distance of the corresponding pixel in the scene from the camera. Such images can be generated, for example, with ToF cameras (Time-of-Flight) or flash lidar systems. In this context, two-and-a-half-dimensional images have the advantage that the additional depth information allows for a more accurate prediction or calculation of object shape information.

[0035] According to at least one embodiment, the image data includes a video of the clothed person.

[0036] In other words, the image data comprises a sequence of successive two-dimensional or two-and-a-half-dimensional camera images. These camera images show the clothed person from different perspectives. Effectively, the image data thus contains three-dimensional image information about the clothed person. This enables a more accurate prediction or calculation of body shape information.

[0037] According to at least one embodiment, the image data includes a two-and-a-half-dimensional or three-dimensional point cloud representing the clothed person.

[0038] Such point clouds can be generated, for example, with laser scanners. In a two-and-a-half-dimensional point cloud, each point is assigned a two-dimensional position and a distance or depth, analogous to the two-and-a-half-dimensional images described above. With a two-and-a-half-dimensional point cloud, the viewing direction is fixed. A three-dimensional point cloud, similar to a video with images from different viewing angles, contains this information for different viewing directions. This allows for a more accurate prediction or calculation of body shape information.

[0039] According to at least one embodiment, the body shape information describes a body contour of the person, particularly in an unclothed state.

[0040] In other words, body shape information indicates where the person's body begins and ends, which is not, or not reliably, discernible from the image data to operators of the imaging system due to clothing or similar factors. However, body contour can significantly influence the selection of at least one imaging parameter, making such embodiments particularly advantageous.

[0041] According to at least one embodiment, the body shape information specifies the respective positions of characteristic points of the person's body.

[0042] The characteristic points, also known as keypoints, are predefined points on the person's body, such as joints, for example, shoulder joints, elbow joints, knee joints, hip joints, wrist joints, ankle joints, and so on. Other examples of characteristic points are the solar plexus, defined points on the person's head or face, and so on.

[0043] The position of the characteristic points can be used to determine at least one imaging parameter, but may not be recognizable from the image data, or not reliably so, due to clothing worn by operators or similar factors. Therefore, such embodiments are particularly advantageous.

[0044] According to at least one embodiment, by applying the MLM to the image data a body model of the person is generated and the body shape information is determined depending on the body model.

[0045] As mentioned at the outset, there are known models that can predict how a person would look clothed from a body model, even for different clothing options. In the present embodiments of the computer-implemented method according to the invention, this approach is, in a sense, reversed, so that the body model is predicted from the clothed person. From the body model, the body shape information can then be extracted directly and / or the secondary body information determined. Such embodiments are particularly advantageous because, with a known body model, different body shape information and / or secondary body information can be determined for different imaging methods as needed, and the MLM does not have to be retrained each time.

[0046] According to a further aspect of the invention, a computer-implemented training method for an MLM, in particular an ANN, for predicting body shape information of a person based on image data of the clothed person, is specified, particularly for use in a computer-implemented method according to the invention. Training data is obtained, and the untrained or partially trained MLM is either supervised or unsupervised, depending on the training data, to predict the body shape information or to predict a body model of the person from which the body models can be derived by applying the MLM to the image data.

[0047] According to at least one embodiment, the MLM includes a convolutional neural network, CNN, or a transformer network, for example a vision transformer network.

[0048] The training is conducted under supervision. In this case, the training data comprises a large number of training datasets. Each training dataset contains training image data of a clothed person and associated baseline truth data.

[0049] The same principles apply to the training image data as described above for the image data. Specifically, it can include a two-dimensional or two-and-a-half-dimensional image of the clothed person and / or a video of the clothed person and / or a two-and-a-half-dimensional or three-dimensional point cloud. However, it can also consist of corresponding simulated images, videos, or point clouds, or images, videos, or point clouds of clothed phantom objects or the like.

[0050] The associated basic truth data of a training image dataset includes the corresponding body shape information or body model for the person, simulated person, or phantom object represented by the training image data of the same training image dataset. This is a one-to-one mapping between the training image data and the basic truth data of each training dataset. In other words, each training dataset contains a pair of corresponding datasets: the respective training image data and its associated basic truth data.

[0051] As mentioned earlier, such pairs of related data sets can be generated through simulation. For example, graphics engines like those used for computer games can be used for this purpose, such as the Unreal Engine mentioned earlier or the Unity Engine.

[0052] Furthermore, well-known supervised training methods can be used to train the CNN or transformer network.

[0053] According to at least one embodiment, the MLM includes a generative adverse network, GAN.

[0054] The GAN can include at least one discriminator and at least one generator. After training, one generator of the GAN is able to predict the body model or body shape information based on the image data. After training, to carry out a computer-implemented method according to the invention, only this generator, and not the entire GAN, may then be required.

[0055] For example, the training is unsupervised. The training data includes a large number of initial training datasets. Each of the initial training datasets contains training image data of a clothed person. The training data includes a large number of secondary training datasets, each of which contains training body shape information or a training body model of a person.

[0056] In this approach, unlike supervised training, the training image data of the first training datasets do not need to be mapped to the training body shape information or training body models of the second training datasets. Therefore, it is not necessary to provide mapped pairs as described above for supervised training. However, the training data can be generated analogously.

[0057] Furthermore, well-known supervised training methods can be used to train the GAN. For example, the GAN can be designed according to the cycle-consistent adversarial networks mentioned earlier.

[0058] According to a further aspect of the invention, an imaging method for imaging a clothed person is provided. In this method, at least one imaging parameter of an imaging system is determined by carrying out a computer-implemented method according to the invention. The clothed person is then imaged using the imaging system configured according to the at least one imaging parameter.

[0059] Further embodiments of the imaging method according to the invention follow directly from the various configurations of the computer-implemented method and the computer-implemented training method according to the invention, and vice versa.

[0060] According to another aspect of the invention, a data processing system is specified which is adapted to carry out a computer-implemented method according to the invention for parameterizing an imaging system.

[0061] According to another aspect of the invention, a further data processing system is specified which is adapted to carry out a computer-implemented training method according to the invention.

[0062] In the present disclosure, the terms "data processing system" and "at least one data processing device" can be used interchangeably. A data processing device can, in particular, be understood to be a data processing device that contains a processing circuit. The data processing device can thus, in particular, process data to perform arithmetic operations. This may also include operations to perform indexed accesses to a data structure, for example, a look-up table (LUT), as well as a data processing process implemented in hardware.

[0063] The data processing device may, in particular, contain one or more computers, one or more microcontrollers, and / or one or more integrated circuits, for example, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), and / or one or more systems on a chip (SoCs). The data processing device may also contain one or more processors, for example, one or more microprocessors, one or more central processing units (CPUs), one or more graphics processing units (GPUs), and / or one or more signal processors, in particular one or more digital signal processors (DSPs).The data processing device may also include a physical or virtual network of computers or other units of the aforementioned type.

[0064] In various embodiments, the data processing device includes one or more hardware and / or software interfaces and / or one or more storage units.

[0065] A storage unit can be volatile data storage, for example as dynamic random access memory (DRAM) or static random access memory (SRAM), or as non-volatile data storage, for example as read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or flash EEPROM, ferroelectric random access memory (FRAM), or magnetoresistive random access memory.It can be designed as MRAM (magnetoresistive random access memory) or as phase-change random access memory, PCRAM (phase-change random access memory).

[0066] According to a further aspect of the invention, an imaging device is provided. The imaging device comprises an imaging system and a data processing system. The data processing system is adapted to perform a computer-implemented method according to the invention for parameterizing an imaging system in order to determine at least one imaging parameter of the imaging system. The imaging system is configured to image a clothed person according to the at least one imaging parameter.

[0067] According to another aspect of the invention, a first computer program with first instructions is specified. When the first instructions are executed by a data processing system, the instructions cause the data processing system to perform a computer-implemented method according to the invention for parameterizing an imaging system.

[0068] According to a further aspect of the invention, a second computer program with second instructions is provided. When the second instructions are executed by a data processing system, the instructions cause the data processing system to perform a computer-implemented training method according to the invention.

[0069] According to a further aspect of the invention, a third computer program with third commands is specified. When the third commands are executed by an imaging device according to the invention, in particular the data processing system of the imaging device, the commands cause the imaging device to perform an imaging method according to the invention.

[0070] The first, second, and / or third instructions can each be provided as program code. This program code can be provided as binary code or assembly language, and / or as source code in a programming language such as C, and / or as a program script, such as Python.

[0071] According to a further aspect of the invention, a computer-readable storage medium is provided which stores a first computer program according to the invention and / or a second computer program according to the invention and / or a third computer program according to the invention.

[0072] The first computer program, the second computer program, the third computer program and the computer-readable storage medium are each computer program products containing the first instructions, the second instructions and / or the third instructions.

[0073] Above and below, the solution according to the invention is described with respect to both the claimed systems and the claimed methods. Features, advantages, or alternative embodiments can be assigned to the other claimed subject matter and vice versa. In other words, the claims and embodiments for the systems can be improved by features described or claimed in connection with the respective methods. In this case, the functional features of the method are implemented by physical units of the system.

[0074] Furthermore, the solution according to the invention is described above and below with respect to methods and systems for parameterizing an imaging system and with respect to methods and systems for providing a trained MLM. Features, advantages, or alternative embodiments can be assigned to the other claimed subject matter and vice versa. In other words, claims and embodiments for providing a trained MLM can be improved with features that are described or claimed in connection with the parameterization of an imaging system.In particular, the datasets used in the procedures and systems can have the same properties and characteristics as the corresponding datasets used in the procedures and systems to provide a trained MLM, and the trained MLMs provided by the respective procedures and systems can be used in the procedures and systems to parameterize an imaging system.

[0075] Further features and combinations of features of the invention will become apparent from the figures and their description, as well as from the claims. In particular, further embodiments of the invention need not necessarily include all features of any one of the claims. Further embodiments of the invention may have features or combinations of features not mentioned in the claims.

[0076] The invention is explained in more detail below with reference to specific embodiments and associated schematic drawings. In the figures, identical or functionally equivalent elements may be designated with the same reference numerals. The description of identical or functionally equivalent elements is not necessarily repeated with respect to different figures.

[0077] The figures show Fig. 1 a schematic representation of an exemplary embodiment of an imaging device according to the invention; Fig. 2 a schematic flowchart of an exemplary embodiment of a computer-implemented method according to the invention for parameterizing an imaging system; Fig. 3 a schematic flowchart of a further exemplary embodiment of a computer-implemented method according to the invention for parameterizing an imaging system; Fig. 4 a schematic flowchart of an exemplary embodiment of a computer-implemented training method according to the invention; Fig. 5 a schematic flowchart of a further exemplary embodiment of a computer-implemented training method according to the invention; Fig. 6 a schematic flowchart of a further exemplary embodiment of a computer-implemented training method according to the invention; Fig. 7 schematically a CNN; and Fig. 8 schematically an ANN.

[0078] In Fig. Figure 1 shows a schematic representation of an exemplary embodiment of an imaging device 1 according to the invention. The imaging device 1 comprises an imaging system and a data processing system 7, 9.

[0079] In the example of the Fig. 1 The imaging system is an X-ray imaging system and comprises a source unit 3 with an X-ray source, a detector unit 4 with an X-ray detector and a control system 7 which is set up to control the X-ray source and the X-ray detector to produce X-ray images depicting a clothed person 6.

[0080] The X-ray imaging system can, for example, comprise a patient table 5 on which the object 6 is arranged. The X-ray imaging system 1 comprises a data processing system 9 according to the invention. In the following, some functions and process steps are described that are executed by the control system 7, while other functions and process steps are described that are executed by the data processing system 9. It should be noted that the functions and process steps may be distributed differently in alternative embodiments.

[0081] For example, the control system 7 can adjust various imaging parameters of the X-ray imaging system, including exposure parameters such as the peak voltage of the X-ray source, the tube current of the X-ray source, and / or the X-ray pulse duration. The control system 7 can also adjust other imaging parameters, such as the filter material and / or the filter thickness of an X-ray filter, for example, a copper filter, by inserting or removing the corresponding X-ray filter from the beam path. Furthermore, the control system 7 can adjust other imaging parameters, such as the size of the collimator aperture of an X-ray collimator 10. For example, the control system 7 can insert or remove the X-ray collimator 10 from the beam path.Control system 7 can, for example, adjust further imaging parameters such as the gain factor of the X-ray detector. Control system 7 can, for instance, introduce or remove an anti-scattering grid from the beam path.

[0082] In some implementations, the X-ray imaging system includes a display device 8, wherein the control system 7 is configured to control the display device 8 to display X-ray images or processed X-ray images.

[0083] In some implementations, the X-ray imaging system 1 is designed as a fluoroscopy system, specifically a C-arm fluoroscopy system, in which the source unit 3 and the detector unit 4 are mounted opposite each other on a C-arm 2, which can be rotated about various axes. These movements are referred to as angular and orbital motion, respectively. In some embodiments, the patient table 5 and the C-arm 2 can be positioned relative to each other by corresponding translational movements of the C-arm 2 and / or the patient table 5, in addition to the rotational movement of the C-arm 2. Consequently, the position and / or orientation of the X-ray source with respect to the object 6 and the position and / or orientation of the X-ray detector with respect to the object 6 can be precisely adjusted to the desired image perspective.

[0084] The data processing system 7, 9 is adapted to carry out a computer-implemented method according to the invention for parameterizing the imaging system in order to determine at least one imaging parameter 13 of the imaging system, as schematically shown in the figures Fig. 2 and Fig. 3 shown. The imaging system is set up to image a clothed person 6 according to the at least one imaging parameter 13, after it has been configured accordingly.

[0085] According to the computer-implemented procedure for parameterizing the imaging system, image data 11 of the clothed person 6 are obtained, for example, a two-dimensional or two-and-a-half-dimensional image or a video of the clothed person 6. Body shape information of the person 6, for example, a body contour of the person 6 and / or the respective positions of characteristic points of the person 6's body, is determined by applying a trained MLM 12 to the image data 11. In the example of the Fig. 2. The MLM 12 is trained accordingly to predict body shape information directly based on the image data 11. In the example of the Fig. 3. The MLM 12 is trained to generate a body model 14 of person 6, and the body shape information is then determined based on the body model 14. In both cases, at least one imaging parameter 13 of the imaging system is determined based on the body shape information.

[0086] Fig. 4 to Fig. Figure 6 shows schematic flowcharts of respective exemplary embodiments of the computer-implemented training methods according to the invention for an MLM 12 for use in a computer-implemented method according to the invention for parameterizing an imaging system, such as with regard to Fig. 1 to Fig. 3 described.

[0087] Training data is obtained and the untrained or partially trained MLM 12 is monitored or unmonitored depending on the training data to predict the body shape information or the body model 14 of the person from which the body shape information can be derived by applying the MLM 12 to the image data 11.

[0088] In the example of the Fig. 4. For example, MLM 12 includes a CNN, and the training is supervised. Specifically, the training data contains a multitude of training datasets, each containing training image data 15 of a clothed person 6 and associated baseline truth data 16. An evaluation block 17 calculates a loss function based on the deviation of the output predicted by MLM 12 from the baseline truth data 16 based on the training image data 15. Based on the result, MLM 12 is updated, for example, by backpropagation.

[0089] In this way, the MLM 12 can, for example, be trained to output an image of person 6 in which the body contour is highlighted. It is also possible for the MLM 12 to output the positioning of characteristic points on person 6's body and / or secondary body information of person 6, such as weight or height.

[0090] In the example of the Fig. In 5, the MLM 12, for example, is a generator module of a GAN 21, and the training is unsupervised. The training data includes a variety of first training datasets, each of which contains training image data 18 of a clothed person. The training data includes a variety of second training datasets, each of which contains training body shape information or a training body model 19 of a person. An evaluation block 20 computes one or more loss functions, as is known for GANs 21 per se.

[0091] The example of Fig. 6 is based on the Fig. 5. GAN 21 is designed as a cycle-consistent adverse network. GAN 21 has a first generator 12 that predicts, for example, a body model 25 based on the training image data 18. A first discriminator 24 assesses whether the predicted body model 25 corresponds to the distribution of the training body models 19. A second generator 22, based on the predicted body model 25, in turn predicts image data 26 of the clothed person 6. A second discriminator 23 assesses whether the predicted image data 26 corresponds to the distribution of the training image data 18. The training of generators 12, 22 and discriminators 23, 24 can be carried out in a known manner, in particular in stages, as described, for example, in the publication by J. Zhu et al. mentioned at the beginning.

[0092] After the training is completed, the first generator 12 is able to correctly predict the body model 14 based on the image data 11 of the clothed person 6 and can therefore be considered an MLM 12 as with regard to Fig. 3 described can be used.

[0093] Fig. Figure 8 shows an embodiment of an artificial neural network, ANN, 800. The ANN 800 comprises nodes 820, ..., 832 and edges 840, ..., 842, where each edge 840, ..., 842 is a directed connection from a first node 820, ..., 832 to a second node 820, ..., 832. In general, the first node 820, ..., 832 and the second node 820, ..., 832 are distinct nodes 820, ..., 832. However, it is also possible for the first node 820, ..., 832 and the second node 820, ..., 832 to be identical. In Fig. For example, edge 840 is a directed connection from node 820 to node 823, and edge 842 is a directed connection from node 830 to node 832. An edge 840, ..., 842 from a first node 820, ..., 832 to a second node 820, ..., 832 is also called an incoming edge for the second node 820, ..., 832 and an outgoing edge for the first node 820, ..., 832.

[0094] In this example, the nodes 820, ..., 832 of ANN 800 can be arranged in layers 810, ..., 813, where the layers can have an intrinsic order introduced by the edges 840, ..., 842 between the nodes 820, ..., 832. In particular, the edges 840, ..., 842 can only exist between adjacent layers of nodes. In the example shown, there is an input layer 810 consisting only of nodes 820, ..., 822 with no incoming edges, an output layer 813 consisting only of nodes 831, 832 with no outgoing edges, and hidden layers 811, 812 between the input layer 810 and the output layer 813. In general, the number of hidden layers 811, 812 can be chosen arbitrarily. For a multilayer perceptron (MLP), this number is at least one. The number of nodes is 820, ..., 822 within the input layer 810 usually refers to the number of input values ​​of the artificial neural network 800, and the number of nodes 831, 832 within the output layer 813 usually refers to the number of output values ​​of the artificial neural network 800.

[0095] In particular, each node 820, ..., 832 of the artificial neural network 800 can be assigned a real number as its value. Here, x denotes... (n) iThe value of the i-th node 820, ..., 832 of the n-th layer 810, ..., 813. The values ​​of nodes 820, ..., 822 of the input layer 810 correspond to the input values ​​of the artificial neural network 800. The values ​​of nodes 831, 832 of the output layer 813 correspond to the output value of the artificial neural network 800. Furthermore, each edge 840, ..., 842 can have a weight, which is a real number. In particular, the weight is a real number within the interval [-1, 1] or within the interval [0, 1]. Here, w denotes (m,n) i,j the weight of the edge between the i-th node 820, ..., 832 of the m-th layer 810, ..., 813 and the j-th node 820, ..., 832 of the n-th layer 810, ..., 813. Furthermore, the abbreviation w (n) i,j for the weight w (n,n+1) i,jdefined. To calculate the output values ​​of neural network 800, the input values ​​are propagated through neural network 800. Specifically, the values ​​of nodes 820, ..., 832 of the (n+1)th layer 810, ..., 813 can be calculated based on the values ​​of nodes 820, ..., 832 of the nth layer 810, ..., 813 as xj(n+1)=f(∑ixi(n)wi,j(n)).

[0096] In this context, the function f is referred to as the transfer function or activation function. Well-known transfer functions include step functions, sigmoid functions (for example, the logistic function, the generalized logistic function, the hyperbolic tangent, the arctangent function), the error function, the smoothstep function, and rectifier functions. The transfer function is used, for example, for normalization. Specifically, the values ​​are propagated layer by layer through neural network 800, with the values ​​of input layer 810 being given by the input of neural network 800. The values ​​of the first hidden layer 811 can be calculated based on the values ​​of input layer 810 of neural network 800, the values ​​of the second hidden layer 812 can be calculated based on the values ​​of the first hidden layer 811, and so on.

[0097] To determine the values ​​w (m,n)i,j To define the edges, the neural network 800 must be trained with training data. The training data includes, in particular, training input data and training output data (denoted as t). i In a training step, the neural network 800 is applied to the training input data to generate computed output data. Specifically, the training data and the computed output data comprise a number of values ​​equal to the number of nodes in the output layer. A comparison between the computed output data and the training data is used to recursively adjust the weights within the neural network 800 (backpropagation algorithm). Specifically, the weights are modified according to the following formula. wi,j'(n)=wi,j(n)−γ δj(n) xi(n), where γ is a predefined learning rate, and the numbers δ (n) j can be calculated recursively as δj(n)=(∑kδk(n+1)wj,k(n+1))f'(xi(n)wi,j(n)) based on δ (n+1) , if the (n+1)th layer is not output layer 813, and δj(n)=(xj(n+1)−tj(n+1))f'(xi(n)wi,j(n)), if the (n+1)th layer is output layer 813, where f is the first derivative of the activation function, and t (n+1) j The comparison training value for the j-th node of the output layer is 813.

[0098] A convolutional neural network (CNN) is an ANN that uses a convolutional operation instead of general matrix multiplication in at least one of its layers. These layers are called convolutional layers. Specifically, a convolutional layer performs a dot product of one or more convolutional kernels with the convolutional layer's input data, where the entries of the one or more convolutional kernels are parameters or weights that can be adjusted through training. In particular, one can use the inner Frobenius product and the ReLU activation function. A convolutional neural network may include additional layers, such as pooling layers, fully connected layers, and / or normalization layers.

[0099] By using convolutional neural networks, input can be processed very efficiently. This is because a convolution operation based on different kernels can extract different image features, allowing the relevant image features to be determined during training by adjusting the weights of the convolution kernel. Furthermore, because weights are shared across convolutional kernels, fewer parameters need to be trained, preventing overfitting during the training phase and enabling faster training or more layers in the network, thus improving network performance.

[0100] Fig.Figure 8 shows an exemplary embodiment of a convolutional neural network 700. In the illustrated embodiment, the convolutional neural network 700 comprises an input node layer 710, a convolution layer 711, a pooling layer 713, a fully connected layer 714, and an output node layer 716, as well as hidden node layers 712 and 714. Alternatively, the convolutional neural network 700 can also include multiple convolution layers 711, multiple pooling layers 713, and / or multiple fully connected layers 715, as well as other types of layers. The order of the layers can be chosen arbitrarily; typically, fully connected layers 715 are used as the last layers before the output layer 716.

[0101] In particular, in a convolutional neural network 700, the nodes 720, 722, 724 of a node layer 710, 712, 714 can be viewed as a d-dimensional matrix or a d-dimensional image. Specifically, in the two-dimensional case, the value of the node 720, 722, 724 indexed by i and j in the nth node layer 710, 712, 714 can be denoted as x(n)[i, j]. However, the arrangement of the nodes 720, 722, 724 of a node layer 710, 712, 714 has no effect on the computations performed within the convolutional neural network 700, as these are determined solely by the structure and the weights of the edges.

[0102] A convolution layer 711 is a connecting layer between a front node layer 710 with node values ​​x(n-1) and a back node layer 712 with node values ​​x(n). A convolution layer 711 is characterized in particular by the structure and weights of the incoming edges that perform a convolution operation based on a certain number of kernels. Specifically, the structure and weights of the edges of the convolution layer 711 are chosen such that the values ​​x(n) of the nodes 722 of the back node layer 712 are computed as a convolution x(n) = K * x(n-1) based on the values ​​x(n-1) of the nodes 720 of the front node layer 710, where the convolution * in the two-dimensional case is defined as x(n)[i,j]=(K∗x(n−1))[i,j]=∑i'∑j'K[i',j']⋅x(n−1)[i−i',j−j'].

[0103] Here, the kernel K is a d-dimensional matrix, in this example a two-dimensional matrix that is usually small compared to the number of nodes 720, 722, for example a 3x3 matrix or a 5x5 matrix. This means, in particular, that the weights of the edges in the convolution layer 711 are not independent, but are chosen such that they yield the aforementioned convolution equation. Specifically, for a kernel that is a 3x3 matrix, there are only 9 independent weights, where each entry of the kernel matrix corresponds to an independent weight, regardless of the number of nodes 720, 722 in the front node layer 710 and the back node layer 712.

[0104] In general, convolutional neural networks use 700 node layers 710, 712, 714 with a multitude of channels, particularly due to the use of a multitude of kernels in the convolution layers 711. In these cases, the node layers can be viewed as (d+1)-dimensional matrices, where the first dimension indexes the channels. The effect of a convolution layer 711 is then defined in a two-dimensional example as xb(n)[i,j]=∑a(Ka,b∗xa(n−1))[i,j]=∑i'∑i'∑j'Ka,b[i',j']⋅xa(n−1)[i−i',j−j'], where xa(n) corresponds to the a-th channel of the preceding layer 710, xb(n) corresponds to the b-th channel of the following nodal layer 712 and K a,b corresponds to one of the kernels. If a convolution layer 711 acts on a preceding nodal layer 710 with A channels and outputs a subsequent nodal layer 712 with B channels, there exist AB-independent d-dimensional kernels K.a,b .

[0105] In general, 700 activation functions can be used in convolutional neural networks. In this embodiment, ReLU (rectified linear unit) is used, with R(z) = max(0, z), so that the effect of the convolution layer 711 in the two-dimensional example xb(n)[i,j]=R(∑a(Ka,b∗xa(n−1)[i,j])=R(∑a∑i'∑j'Ka,b[i',j']⋅xa(n−1)[i−i',j−j']) It is also possible to use other activation functions, such as ELU (Exponential Linear Unit), LeakyReLU, Sigmoid, Tanh, or Softmax.

[0106] In the illustrated embodiment, the input layer 710 comprises 36 nodes 720 arranged in a two-dimensional 6×6 matrix. The first hidden node layer 712 comprises 72 nodes 722 arranged as two-dimensional 6×6 matrices, each of which is the result of convolution of the values ​​of the input layer with a 3×3 kernel within the convolution layer 711. Equivalently, the nodes 722 of the first hidden node layer 712 can be interpreted as a three-dimensional 2×6×6 matrix, the first dimension corresponding to the channel dimension.

[0107] One advantage of using convolutional layers 711 is that a spatially local correlation of the input data can be exploited by enforcing a local connectivity pattern between the nodes of neighboring layers, in particular by connecting each node only to a small range of the nodes of the preceding layer.

[0108] A pooling layer 713 is a connecting layer between an preceding node layer 712 with node values ​​x(n-1) and a subsequent node layer 714 with node values ​​x(n). A pooling layer 713 can be characterized, in particular, by the structure and weights of the edges and the activation function, which perform a pooling operation based on a nonlinear pooling function f. For example, in the two-dimensional case, the values ​​x(n) of the nodes 724 of the subsequent node layer 714 can be calculated based on the values ​​x(n-1) of the nodes 722 of the anterior node layer 712 as follows. xb(n)[i,j]=f(xb(n−1)[id1,jd2],..., xb(n−1)[(i+1)d1−1,(j+1)d2−1]).

[0109] In other words, by using a pooling layer 713, the number of nodes 722, 724 can be reduced by replacing a number of d1-d2 neighboring nodes 722 in the preceding node layer 712 with a single node 722 in the subsequent node layer 714, which is calculated as a function of the values ​​of the aforementioned number of neighboring nodes. The pooling function f can, in particular, be the max function, the mean, or the L2 norm. Specifically, in a pooling layer 713, the weights of the incoming edges are fixed and are not changed by training.

[0110] The advantage of using a pooling layer 713 is that it reduces the number of nodes (722, 724) and the number of parameters. This leads to a reduction in computational overhead in the network and helps control overfitting.

[0111] In the illustrated embodiment, the pooling layer 713 is a max-pooling layer, in which four adjacent nodes are replaced by only one node, where the value is the maximum of the values ​​of the four adjacent nodes. The max-pooling is applied to each d-dimensional matrix of the preceding layer. In this embodiment, the max-pooling is applied to each of the two-dimensional matrices, thereby reducing the number of nodes from 72 to 18.

[0112] In general, the last layers of a convolutional neural network can be fully connected layers. A fully connected layer is a linking layer between a preceding node layer and a subsequent node layer. A fully connected layer can be characterized by the presence of a majority, in particular all, edges between the nodes of the preceding node layer and the nodes of the subsequent node layer, and the weight of each of these edges can be individually adjusted.

[0113] In this embodiment, the nodes 724 of the leading node layer 714 of the fully connected layer 715 are represented both as two-dimensional matrices and additionally as non-connected nodes, displayed as a line of nodes, the number of which has been reduced for better clarity. This process is also called flattening. In this embodiment, the number of nodes 726 in the subsequent node layer 716 of the fully connected layer 715 is less than the number of nodes 724 in the preceding node layer 714. Alternatively, the number of nodes 726 can also be equal to or greater than the number of nodes 724 in the preceding node layer 714.

[0114] Furthermore, in this embodiment, the softmax activation function is used within the fully connected layer 715. By applying the softmax function, the sum of the values ​​of all nodes 726 of the output layer 716 is equal to 1, and all values ​​of all nodes 726 of the output layer 716 are real numbers between 0 and 1. In particular, when using the convolutional neural network 700 to categorize input data, the values ​​of the output layer 716 can be interpreted as the probability that the input data falls into one of the various categories.

[0115] In particular, convolutional neural networks 700 can be trained based on the backpropagation algorithm. To prevent overfitting, regularization techniques can be used, such as omitting nodes 720, ..., 724, stochastic pooling, using artificial data, weight reduction based on the L1 or L2 norm, or max-norm constraints.

[0116] Regardless of the grammatical gender of a particular term, persons with male, female or other gender identities are included.

Claims

[1] Computer-implemented method for parameterizing an imaging system for imaging a clothed person (6), wherein - Image data (11) of the clothed person (6) will be obtained; - Body shape information of the person is determined by applying a trained machine learning model, MLM, (12) to the image data (11); and - at least one imaging parameter (13) of the imaging system is determined depending on the body shape information, wherein secondary body information of the person is estimated depending on the body shape information and the at least one imaging parameter (13) is determined depending on the secondary body information, wherein the secondary body information includes a body weight of the person and / or a material composition of the person's body. [2] Computer-implemented method according to claim 1, wherein the image data (11) - include a two-dimensional or two-and-a-half-dimensional image of the clothed person (6); and / or - include a video of the clothed person (6); and / or - include a two-and-a-half-dimensional or three-dimensional point cloud representing the clothed person (6). [3] Computer-implemented method according to any of the preceding claims, wherein the body shape information - describe a person's body contour; and / or - Specify the respective positions of characteristic points on the person's body. [4] Computer-implemented method according to one of the preceding claims, wherein by applying the MLM (12) to the image data (11) a body model (14) of the person is generated and the body shape information is determined depending on the body model (14). [5] Computer-implemented method according to any one of the preceding claims, wherein - the imaging system is an X-ray-based imaging system and includes at least one imaging parameter (13), at least one exposure setting of an X-ray source (3) of the imaging system and / or a detector gain of an X-ray detector (4) of the imaging system and / or a collimator position of a collimator (10) of the imaging system and / or a size of a collimator aperture of the collimator (10) and / or a shape of a collimator aperture of the collimator (10); or - the imaging system is a magnetic resonance imaging system and at least one imaging parameter (13) includes an imaging target area within a patient tube of the magnetic resonance imaging system. [6] Computer-implemented method according to one of the preceding claims, wherein the imaging system is configured at least partially automatically according to the determined at least one imaging parameter (13). [7] Imaging techniques for depicting a clothed person (6) wherein - at least one imaging parameter (13) of an imaging system is determined by performing a computer-implemented method according to any one of claims 1 to 6; and - the clothed person (6) is imaged using the imaging system configured according to at least one imaging parameter (13). [8] Imaging method according to claim 7, wherein - the image data (11) of the clothed person (6) are generated by means of a camera; and / or - the imaging system is configured according to which at least one imaging parameter (13) is configured; and / or - the imaging system configured according to at least one imaging parameter (13) to image the clothed person (6); and / or - by imaging the clothed person (6) a medical image data set is generated. [9] Data processing system (7, 9) adapted to carry out a computer-implemented method according to any one of claims 1 to 6. [10] Imaging device (1) comprising an imaging system and a data processing system (7, 9), wherein - the data processing system (7, 9) is adapted to perform a computer-implemented method according to one of claims 1 to 6 in order to determine at least one imaging parameter (13) of the imaging system; and - the imaging system is configured to image a clothed person (6) according to at least one imaging parameter (13). [11] computer program product - Commands which, when executed by a data processing system (7, 9), cause the data processing system (7, 9) to perform a computer-implemented method according to any one of claims 1 to 6; and / or - further commands which, when executed by an imaging device (1) according to claim 10, cause the imaging device (1) to perform an imaging procedure according to one of claims 7 or 8.

Citation Information

Patent Citations

  • Topogram Prediction from Surface Data in Medical Imaging

    US20190057521A1