Learning device, learning method, operating program for learning device, teacher data generation device
By using three-dimensional computer graphics data to generate pseudo-images, the learning device efficiently trains machine learning models to accurately determine imaging locations and directions, addressing the challenges of subject variability and reducing reshoots in medical imaging systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FUJIFILM CORP
- Filing Date
- 2022-02-14
- Publication Date
- 2026-06-02
AI Technical Summary
Existing medical imaging systems face challenges in accurately determining the imaging location and direction of subjects due to the variability in subject posture and appearance, leading to a high frequency of reshoots, and there is a need for efficient training of machine learning models to improve recognition accuracy.
A learning device and method that uses three-dimensional computer graphics data to generate pseudo-images mimicking various subject postures and appearances, combined with actual camera images, to train a machine learning model for precise imaging area and direction determination.
The approach enables efficient training of machine learning models to accurately determine imaging locations and directions, reducing the need for reshoots and enhancing the recognition accuracy of imaging systems.
Smart Images

Figure 0007869188000001 
Figure 0007869188000002 
Figure 0007869188000003
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a learning device, a learning method, an operation program of the learning device, a teacher data generation device, a machine learning model, and a medical imaging device.
Background Art
[0002] As a medical imaging device, for example, a radiographic imaging device is known. At the time of radiographic imaging, based on the imaging order acquired by the doctor who requested the imaging, a technician performs positioning to adjust the relative positional relationship between the radiographic imaging device and the subject (for example, Japanese Patent Application Laid-Open No. 2020-192440). In the imaging order, an imaging site such as the chest or abdomen and an imaging direction such as the front or back are defined. The technician positions the subject based on the content of the imaging order.
[0003] Japanese Patent Application Laid-Open No. 2020-192440 discloses a technique for assisting appropriate positioning by photographing a subject in a state of being positioned with respect to a radiographic imaging device with an optical camera and causing a technician to confirm a composite image obtained by combining the photographed camera image and a mark indicating an ideal positioning state.
Summary of the Invention
Problems to be Solved by the Invention
[0004] It is being considered to use camera images like the one described above to verify whether the subject's imaging location and direction match the imaging order. Specifically, it is being considered to use a camera image as input and have a machine learning model derive the subject's imaging location and direction as captured in the camera image. This would help to suppress errors such as imaging a different location than the one specified in the imaging order. In medical imaging, the subject is often unable to move freely, and the burden of reshoots can be significant, so there is a greater need to suppress reshoots compared to other applications. Therefore, when using such machine learning models in medical imaging, a machine learning model with high accuracy in recognizing the subject's imaging location and direction as captured in the camera image is desired.
[0005] To improve the recognition accuracy of a machine learning model, the more training data consisting of ground truth data for combinations of camera images, shooting location, and shooting direction, the better. In particular, the more diverse the camera images in the training data are for each combination of shooting location and shooting direction, the better the recognition accuracy of the machine learning model will be. Even with the same combination of shooting location and shooting direction, the posture or appearance of the subject captured in the camera image can vary. Even if the posture or appearance of the subject differs, the machine learning model needs to be able to recognize that it is the same combination of shooting location and shooting direction. To achieve this, it is necessary to collect a large number of diverse camera images for each combination of shooting location and shooting direction, each with different subject postures and appearances, to use as training data.
[0006] However, collecting a large number of such camera images is extremely time-consuming, so there was a need for a way to efficiently improve the recognition accuracy of machine learning models.
[0007] The technology disclosed herein provides a learning device, learning method, learning program, training data generation device, and machine learning model that enable efficient training of a machine learning model for deriving the shooting location and shooting direction of a subject captured in a camera image, compared to using only camera images as training data. [Means for solving the problem]
[0008] To achieve the above objective, the learning device of this disclosure is a learning device that takes a camera image taken with an optical camera of a subject positioned in front of a medical imaging device as input and trains a machine learning model that derives the imaging area and imaging direction of the subject captured in the camera image, and comprises a processor and memory connected to or built into the processor, wherein the processor generates multiple pseudo-images of the human body, which are generated based on a human body model composed of three-dimensional computer graphics data, and which mimic a subject positioned in front of a medical imaging device, for each combination of imaging area and imaging direction, and trains the machine learning model using multiple training data consisting of the generated pseudo-images and correct data for each combination.
[0009] Three-dimensional computer graphics data may include modeling parameters for modifying at least one of the posture and appearance of a human body model.
[0010] The modeling parameters may include at least one of the following: body shape information representing the physique of the human body model, gender, posture information, skin color, hair color, hairstyle, and clothing.
[0011] The processor can generate a pseudo-image by rendering a human body model from a set viewpoint, and the viewpoint may be changeable by rendering parameters.
[0012] The viewpoint information used to set the viewpoint may include the focal length of a virtual camera virtually placed at the viewpoint and the shooting distance, which is the distance from the virtual camera to the human body model.
[0013] When the training data consisting of pseudo-images and ground truth data is used as the first training data, the processor may train the machine learning model using the first training data plus the second training data consisting of camera images captured by an optical camera and ground truth data.
[0014] Furthermore, a machine learning model for body size output may be trained using multiple training data sets for body size output, which consist of pseudo-images and ground truth data representing the body size of a human body model, to derive body size information representing the body size of a subject from camera images as input.
[0015] A medical imaging device may include at least one of the following: a radiography device and an ultrasound device.
[0016] The learning method described herein is a learning method that uses a computer to train a machine learning model that derives the imaging area and imaging direction of a subject captured in a camera image, using a camera image of a subject positioned in front of a medical imaging device as input. The method involves generating multiple pseudo-images of a human body based on a human body model composed of three-dimensional computer graphics data, each pseudo-image mimicking a subject positioned in front of a medical imaging device, for each combination of imaging area and imaging direction. The machine learning model is then trained using multiple training data sets consisting of the generated pseudo-images and the correct data for each combination.
[0017] The program for operating a learning device is a program for operating a computer as a learning device that takes camera images taken with an optical camera of a subject positioned in front of a medical imaging device as input, and trains a machine learning model that derives the imaging area and direction of the subject as captured in the camera images, wherein the program generates multiple pseudo-images of the human body based on a human body model composed of three-dimensional computer graphics data, each pseudo-image that mimics a subject positioned in front of a medical imaging device, for each combination of imaging area and direction, and uses multiple training data consisting of the generated pseudo-images and the correct data for each combination to train the machine learning model.
[0018] The training data generation device of this disclosure is a training data generation device that takes camera images taken with an optical camera of a subject positioned in front of a medical imaging device as input and generates training data used to train a machine learning model that derives the imaging area and imaging direction of the subject captured in the camera image, and comprises a processor and memory connected to or built into the processor, wherein the processor uses three-dimensional computer graphics data that constitutes a human body model for generating pseudo-images of the human body, and the three-dimensional computer graphics data is accompanied by parameters for changing at least one of the posture and appearance of the human body model, and by changing the parameters, generates a plurality of pseudo-images for each combination of imaging area and imaging direction in which at least one of the posture and appearance of the human body model differs, and generates a plurality of training data consisting of the plurality of generated pseudo-images and the correct data for the combination.
[0019] The machine learning model of the present disclosure is a machine learning model that takes, as input, a camera image of a subject positioned with respect to a medical imaging device captured by an optical camera, and derives the imaging region and imaging direction of the subject captured in the camera image. The machine learning model is composed of a plurality of teacher data generated for each combination, and is composed of a pseudo image of the human body generated based on a human body model constituted by three-dimensional computer graphics data, the pseudo image simulating the subject positioned with respect to the medical imaging device, and correct answer data of the combination of the imaging region and the imaging direction. of It is trained using
[0020] The medical imaging device of the present disclosure includes a machine learning model.
Advantages of the Invention
[0021] According to the technology of the present disclosure, a machine learning model for deriving the imaging region and imaging direction of a subject captured in a camera image can be efficiently trained as compared with the case of using only the camera image as teacher data.
Brief Description of the Drawings
[0022] [Figure 1] It is a diagram showing a schematic configuration of a learning device and a radiation imaging system. [Figure 2] It is a diagram for explaining an imaging order. [Figure 3] It is a diagram for explaining a method of determining irradiation conditions. [Figure 4] It is a diagram for explaining the functions of a console. [Figure 5] It is a diagram for explaining a learned model. [Figure 6] It is a diagram showing the hardware configuration of a learning device. [Figure 7] It is a diagram for explaining the functions of a learning device. [Figure 8] It is a diagram showing an example of 3DCG data. [Figure 9] It is a diagram showing the functions of a teacher data generation unit. [Figure 10]This figure shows an example of a pseudo-image. [Figure 11] This figure shows an example of a pseudo-image. [Figure 12] This figure shows an example of a pseudo-image. [Figure 13] This figure shows an example of a pseudo-image. [Figure 14] This figure shows an example of a pseudo-image. [Figure 15] This figure shows an example of a pseudo-image. [Figure 16] This is the main flowchart illustrating the processing procedure of the learning device. [Figure 17] This flowchart shows the procedure for generating training data. [Figure 18] This is a flowchart of the learning procedure. [Figure 19] This figure shows a modified version in which pseudo-images and camera images are mixed together as training data. [Figure 20] This figure shows a modified example with an added matching section. [Figure 21] This is a diagram illustrating a pre-trained model for physique output. [Figure 22] This diagram illustrates the training data used to train a machine learning model for physique output. [Figure 23] This diagram illustrates a training data generation unit applied to an ultrasound imaging device. [Modes for carrying out the invention]
[0023] Figure 1 is a diagram showing a schematic of the overall configuration of the learning device 40 and radiography system 10 of this disclosure. The radiography system 10 is an example of a medical imaging device related to the technology of this disclosure, and is also an example of a radiography device. The radiography system 10 obtains a radiographic image XP of the subject H by photographing the subject H using radiation R. When radiography is performed using the radiography system 10, the operator, a radiologic technologist (hereinafter simply referred to as "technologist") RG, positions the subject H relative to the radiography system 10.
[0024] The radiography system 10 in this example has an imaging support function that assists the technician RG in confirming whether the patient H is properly positioned according to the imaging order 31. As will be described in detail later, the imaging support function is performed using a trained model LM generated by training the machine learning model LM0. The learning device 40 has a learning unit 52 for training the machine learning model LM0. The learning unit 52 is used to train the machine learning model LM0 in order to generate the trained model LM provided to the radiography system 10.
[0025] Here, although the trained model LM is also a machine learning model, for convenience it is distinguished from the machine learning model LM0 which is the target of training by the learning device 40. The trained model LM is a machine learning model that has been trained at least once by the learning device 40 and is put into operation in the radiography system 10. Note that the machine learning model LM0 may be an untrained machine learning model or a trained model LM which is the target of additional training. Below, the outline of the radiography system 10 will be described, followed by a description of the learning device 40.
[0026] As shown in Figure 1, the radiography system 10 comprises a radiation source 11, a radiation source control device 12, an electronic cassette 13, and a console 14. The electronic cassette 13 is an example of a radiographic image detector that detects the radiographic image XP of the subject H by receiving radiation R transmitted through the subject H, and is a portable radiographic image detector. The radiography system 10 in this example also has an optical camera 15. This configuration is for realizing the above-mentioned imaging support function, and the optical camera 15 is used to capture the state of the subject H positioned relative to the radiography system 10.
[0027] Positioning for radiography involves adjusting the relative positions of the subject H, the electronic cassette 13, and the radiation source 11. For example, this is done as follows: Technician RG first aligns the electronic cassette 13 with the area of the subject H being photographed. In the example in Figure 1, the electronic cassette 13 is positioned facing the chest of the subject H. Then, the position of the radiation source 11 is adjusted so that the electronic cassette 13, positioned at the area of the subject H being photographed, faces the radiation source 11. Also in the example in Figure 1, the radiation source 11 faces the back of the subject H's chest, and the direction of radiation (the direction from which the radiation R is emitted) is the back of the subject H. When photographing the chest from the front, Technician RG positions the front of the subject H's chest facing the radiation source 11. By performing this positioning, radiation R is emitted from the back of the subject H, allowing for the acquisition of a radiographic image XP of the subject H's chest.
[0028] In the example shown in Figure 1, the electronic cassette 13 is set on a standing imaging table 25 for imaging the subject H in an upright position. The electronic cassette 13 may also be set on a supine imaging table or the like for imaging the subject H in a supine position. Furthermore, since the electronic cassette 13 is portable, it can also be removed from the imaging table and used independently.
[0029] The radiation source 11 comprises a radiation tube 11A that generates radiation R and an irradiation field limiter 11B that limits the irradiation field, which is the area irradiated by radiation R. The radiation tube 11A has, for example, a filament that emits thermionic electrons and a target that emits radiation when the thermionic electrons emitted from the filament collide with it. The irradiation field limiter 11B has, for example, four lead plates that shield radiation R placed on each side of a rectangle, so that a rectangular irradiation aperture that transmits radiation R is formed in the center. In the irradiation field limiter 11B, the size of the irradiation aperture changes by moving the position of the lead plates. This adjusts the size of the irradiation field. In addition to the radiation tube 11A, the radiation source 11 may also incorporate an irradiation field display light source (not shown) for visualizing the irradiation field by projecting visible light onto the subject H through the irradiation aperture.
[0030] In the example shown in Figure 1, the radiation source 11 is a ceiling-suspended type and is attached to an extendable support column 22. The vertical height of the radiation source 11 can be adjusted by extending or retracting the support column 22. The support column 22 is attached to a ceiling-mounted device (not shown) that runs on rails placed on the ceiling and can move horizontally along the rails. Furthermore, the radiation source 11 can rotate around the focal point of the radiation tube 11A as the center of rotation. The direction in which the radiation source 11 irradiates with radiation R can be adjusted by these various displacement mechanisms.
[0031] The radiation source control device 12 controls the radiation source 11. The radiation source control device 12 is equipped with an operation panel (not shown). By operating the operation panel, the technician RG sets the radiation irradiation conditions and the size of the irradiation aperture of the irradiation field limiter 11B. The radiation irradiation conditions include the tube voltage (unit: kV), tube current (unit: mA), and radiation irradiation time (unit: mS) applied to the radiation source 11.
[0032] The radiation source control device 12 includes a voltage generator that generates a voltage to be applied to the radiation tube 11A, and a timer. The radiation source control device 12 operates the radiation source 11 to generate radiation R according to the irradiation conditions by controlling the voltage generator and the timer. An irradiation switch 16 is also connected to the radiation source control device 12 via a cable or the like. The irradiation switch 16 is operated by the technician RG when radiation irradiation is started. When the irradiation switch 16 is operated, the radiation source control device 12 generates radiation in the radiation tube 11A. As a result, radiation R is irradiated toward the irradiation field.
[0033] As described above, the electronic cassette 13 detects the radiographic image XP based on the radiation R that was irradiated from the radiation source 11 and transmitted through the imaging area of the subject H. The electronic cassette 13, for example, has a wireless communication unit and a battery, and can operate wirelessly. The electronic cassette 13 wirelessly transmits the detected radiographic image XP to the console 14.
[0034] The optical camera 15 is an optical digital camera comprising a CMOS (Complementary Metal Oxide Semiconductor) type image sensor or a CCD (Charge Coupled Device) type image sensor, and performs imaging based on visible light, for example. The optical camera 15 is capable of taking still images and videos, for example.
[0035] The optical camera 15 is used to photograph the subject H in a position relative to the electronic cassette 13 and the radiation source 11. For example, the optical camera 15 is mounted on the outer periphery of the irradiation field limiter 11B of the radiation source 11 and positioned near the irradiation aperture. The optical camera 15 is mounted in a position where its optical axis is parallel to the irradiation axis of the radiation R from the radiation source 11. The optical camera 15 generates a camera image CP, which is an optical image in visible light, by optically photographing the area including the irradiation field of radiation R. Since the radiation source 11 is positioned at the area of the subject H being photographed, the camera image CP taken in that state will depict the area of the subject H being photographed. In this example, the camera image CP is, for example, a color still image.
[0036] In this example, the optical camera 15 is attached to the outer periphery of the irradiation field limiter 11B, but the optical camera 15 does not have to be attached to the outer periphery of the radiation source 11; it may be built into the radiation source 11.
[0037] The optical camera 15 is connected to the console 14 by wire or wireless connection. The console 14 functions as a control device for the optical camera 15, controlling the shooting operation of the optical camera 15, such as the timing of the shots. Technician RG, for example, inputs instructions for shooting with the optical camera 15 into the console 14 after positioning the subject H.
[0038] Console 14 is connected via network N to the RIS (Radiology Information System) and PACS (Picture Archiving and Communication System) installed in the radiography system 10.
[0039] The RIS is a device that manages imaging orders 31 for the radiography system 10. For example, in a medical facility, physicians in departments such as internal medicine and surgery request radiography from the radiology department, which is responsible for radiography. The imaging order 31 is issued from the physician in the department to the radiology department. The RIS manages the imaging orders 31 from the physician in the department. If there are multiple radiography systems 10, the multiple imaging orders 31 managed by the RIS are assigned to the multiple radiography systems 10 according to the content of the imaging order and the operating status of the radiography systems 10. The console 14 receives the imaging orders 31 transmitted from the RIS.
[0040] Figure 2 shows an example of an imaging order 31. The contents of the imaging order 31 include an order ID (Identification Data) issued for each order, a subject ID issued for each subject H, the imaging procedure, and the purpose of imaging (not shown). Here, the imaging procedure is an imaging method defined by at least a combination of imaging site and imaging direction. In the imaging order 31, the physician specifies the imaging procedure, including the imaging site and imaging direction, according to the purpose of the examination for the subject H. For example, if the purpose of the examination is the diagnosis of lung cancer, the imaging procedure is specified as the chest imaging site and the imaging direction as either dorsal or anterior. In Figure 2, in imaging order 31 with order ID "N0001", "chest / dorsal" is specified as the imaging procedure. In imaging order 31 with order ID "N0002", "chest / anterior" is specified as the imaging procedure. In imaging order 31 with order ID "N0003", "abdomen / anterior" is specified as the imaging procedure. In imaging order 31 with order ID "N0004", the imaging technique specified is "both knees / frontal". In imaging order 31 with order ID "N0005", the imaging technique specified is "right knee / lateral".
[0041] Furthermore, as shown in Figure 3, the imaging procedure information, along with the subject H's physique information, is used by technician RG to determine the irradiation conditions of the radiation source 11. For example, if the imaging procedure is "chest / back," technician RG considers the subject H's physique information (primarily body thickness) to estimate the chest thickness and determines the tube voltage, tube current, and irradiation time. Generally, the thicker the body, the lower the transmittance of radiation R, so a higher radiation dose of R is set. The irradiation dose is determined by the mAs value, which is the product of the tube current and the irradiation time.
[0042] Returning to Figure 1, the PACS stores the radiographic images XP taken by the radiography system 10. The console 14 transmits the radiographic images XP received from the electronic cassette 13 to the PACS, associating them with the imaging order 31. In the PACS, the radiographic images XP are stored in an image file format, such as one conforming to the DICOM (Digital Imaging and Communication in Medicine) standard. The radiographic images XP stored in the PACS are made available for viewing by physicians in the requesting department that issued the imaging order 31.
[0043] Console 14 is composed of a computer, such as a personal computer or workstation. Console 14 has an order receiving function for receiving imaging orders 31, a setting function for configuring various settings of the electronic cassette 13, and a function for displaying the radiographic image XP received from the electronic cassette 13 on the display 14C. .child In addition to these basic functions, console 14 is equipped with the aforementioned shooting support functions.
[0044] As shown in Figure 1 and also enlarged in Figure 4, the console 14 includes an imaging order reception unit 14A and an imaging technique determination unit 14B as components related to the imaging support function. The imaging order reception unit 14A displays the imaging order 31 received from the RIS on the display 14C. The imaging technique determination unit 14B derives the imaging technique for the positioned subject H by performing image analysis on the camera image CP using a trained model LM. The imaging technique determination unit 14B outputs a determination result 32 including the imaging technique derived by the trained model LM. The console 14 displays a positioning confirmation screen 36 on the display 14C, which includes the imaging order 31 received by the imaging order reception unit 14A and the determination result 32 from the imaging technique determination unit 14B.
[0045] On the positioning confirmation screen 36, for example, the imaging order 31 and the judgment result 32 are displayed side by side for easy comparison. For the imaging order 31, only the part related to the imaging procedure is extracted, and a message 31A such as "The imaging order is 'chest / back'" is displayed. Also, the judgment result 32 displays the imaging procedure determined by the imaging procedure determination unit 14B (in the example in Figure 4, "chest / back"). ) This includes the camera image CP. In the judgment result 32, the imaging procedure is displayed in the form of a message such as, for example, "The imaging procedure determined from the camera image is 'chest / back'."
[0046] Through the positioning confirmation screen 36, technician RG can visually compare the imaging order 31 with the imaging procedure included in the judgment result 32 to confirm whether the positioning of the subject H conforms to the imaging order 31.
[0047] As shown in Figure 5, a pre-trained model LM, for example, is a convolutional neural network (CNN) suitable for image analysis. The pre-trained model LM includes, for example, an encoder 37 and a classification unit 38.
[0048] Encoder 37 is composed of a Convolutional Neural Network (CNN) and extracts multiple types of feature maps representing the features of the camera image CP by performing convolution and pooling operations on the camera image CP. As is well known, in a CNN, convolution is a spatial filtering process that uses multiple types of filters, for example, each with a size such as 3x3. In the case of a 3x3 filter, filter coefficients are assigned to each of the 9 cells. In convolution, the center cell of these filters is aligned with the target pixel of the camera image CP, and the sum of the products of the pixel values of the target pixel and the 8 surrounding pixels (a total of 9 pixels) is output. The output sum of products represents the feature quantity of the target region to which the filter was applied. Then, for example, by shifting the target pixel one pixel at a time and applying the filter to all pixels of the camera image CP, a feature map with a feature quantity equivalent to the number of pixels in the camera image CP is output. By applying multiple types of filters with different filter coefficients, multiple types of feature maps are output. The number of feature maps corresponding to the number of filters is also called the number of channels.
[0049] This convolution process is repeated while reducing the size of the camera image CP. SizeThe process of reducing the size of the image is called pooling. Pooling is performed by applying processes such as decimation and averaging of adjacent pixels. Through pooling, the size of the camera image CP is reduced in stages, such as to 1 / 2, 1 / 4, and 1 / 8. A multi-channel feature map is output for each size of the camera image CP. When the size of the camera image CP is large, the camera image CP depicts fine morphological features of the subject, whereas when the size of the camera image CP is small (i.e., the resolution is low), fine morphological features are abstracted from the camera image CP, and only the general morphological features of the subject are depicted. Therefore, the feature map when the camera image CP is large represents the microscopic features of the subject depicted in the camera image CP, and the feature map when the size is small represents the macroscopic features of the subject depicted in the camera image CP. The encoder 37 extracts multi-channel feature maps that represent the macroscopic and microscopic features of the camera image CP by applying such convolution and pooling processes to the camera image CP.
[0050] The trained model LM is, for example, a classification model that derives the single most likely imaging technique from among several imaging techniques as indicated by the camera image CP. Therefore, a classification unit 38 is provided to derive a single imaging technique based on the features of the camera image CP extracted by the encoder 37. The classification unit 38 includes, for example, multiple perceptrons, each having one output node for multiple input nodes. The perceptrons are also assigned weights indicating the importance of the multiple input nodes. In each perceptron, the sum of the products obtained by multiplying each input value input to the multiple input nodes by their respective weights is output as the output value from the output node. Such perceptrons can be formulated using an activation function such as a sigmoid function, for example.
[0051] In the output section, the outputs and inputs of multiple perceptrons are coupled, forming a multi-layered neural network with multiple hidden layers between the input and output layers. As a method for coupling multiple perceptrons between layers, for example, a fully coupled configuration is employed in which all output nodes of the previous layer are coupled to one input node of the next layer.
[0052] In the input layer of the classification unit 38, all features included in the feature map of the camera image CP are input. In the perceptrons that make up each layer of the classification unit 38, the features are input as input values to the input nodes. Then, for each input node, the sum of the values obtained by multiplying the features by their weights is output as an output value from the output node, and this output value is passed to the input node of the perceptron of the next layer. In the final output layer of the classification unit 38, based on the output values of multiple perceptrons, the probability of each of several types of imaging procedures is output using a softmax function or the like. Then, based on these probabilities, the single imaging procedure with the highest probability is derived. Figure 5 shows an example in which the imaging procedure "chest / back" is derived based on the camera image CP.
[0053] This example is just one illustration; the trained model LM can be any other type of classification model that can derive the shooting technique based on the camera image CP.
[0054] Next, the learning device 40 will be described using Figures 6 and 7. Figure 6 shows the hardware configuration of the learning device 40. The learning device 40 is composed of a computer such as a personal computer or workstation. The learning device 40 includes a display 41, an input device 42, a CPU 43, a memory 44, a storage device 46, and a communication unit 47. These are interconnected via a data bus 48.
[0055] The display 41 is a display unit that displays various operation screens equipped with GUI (Graphical User Interface) operation functions. The input device 42 is an input operation unit including a touch panel or keyboard.
[0056] The storage device 46 consists of, for example, an HDD (Hard Disk Drive) and an SSD (Solid State Drive), and is either built into the learning device 40 or connected externally to the learning device 40. External connection is made via cable or network. The storage device 46 stores control programs such as the operating system, various application programs, and various data associated with these programs. One of the various application programs is an operation program AP that makes the computer function as a learning device 40. The various data includes the machine learning model LM0 that is the target of the learning process, the trained model LM after the learning process is completed, the training data TD used for learning, three-dimensional computer graphics data (hereinafter referred to as 3DCG data) 56, and parameter specification information 61.
[0057] Memory 44 is work memory for the CPU 43 to execute processing. The CPU 43 loads programs stored in storage device 46 into memory 44 and executes processing according to the programs, thereby comprehensively controlling each part of the learning device 40.
[0058] The communication unit 47 communicates with the console 14 via the network N. The communication unit 47 is used for transferring trained model LMs, for example, by sending trained model LMs from the learning device 40 to the console 14, and by sending trained model LMs from the console 14 to the learning device 40 in order to further train the trained model LMs on the console 14.
[0059] Figure 6 shows the hardware configuration of the computer for realizing the learning device 40, and the hardware configuration of the console 14 described above is similar. That is, the configuration of the console 14, such as the imaging order receiving unit 14A and the imaging procedure judgment unit 14B, is realized through the cooperation of a processor such as the CPU 43, memory such as the memory 44 which is built into or connected to the CPU 43, and a program executed by the CPU 43.
[0060] As shown in Figure 7, the learning device 40 has a training data generation unit 51 and a learning unit 52. Each of these processing units is implemented by a CPU 43. The training data generation unit 51 generates multiple simulated images SP of the human body, which are generated based on a human body model 56A composed of 3DCG data 56, for each imaging procedure, which is a combination of imaging area and imaging direction. The simulated image SP is a two-dimensional image that mimics a subject H positioned in front of a radiography system 10, which is an example of a medical imaging device.
[0061] More specifically, the training data generation unit 51 generates a pseudo-image SP based on three-dimensional computer graphics data (hereinafter referred to as 3DCG data) 56 and parameter specification information 61. The parameter specification information 61 includes modeling parameter specification information 61M (hereinafter referred to as M-parameters) specified when modeling a three-dimensional human body model 56A that mimics the human body, and rendering parameter specification information 61R (hereinafter referred to as R-parameters) specified when setting a viewpoint and rendering the human body model 56A. Details of the modeling of the human body model 56A and the method of generating the pseudo-image SP by rendering the human body model 56A will be described later.
[0062] The training data generation unit 51 then generates training data TD, which consists of the generated pseudo-image SP and the correct data AD, which is a combination of the imaging area and imaging direction. The training data TD shown in Figure 7 is an example where the imaging technique, which is the combination of imaging area and imaging direction, is "chest / back". In other words, in the training data TD shown in Figure 7, the pseudo-image SP shows the back of the chest of a human body and is an image that indicates the imaging technique as "chest / back", and correspondingly the imaging technique in the correct data AD is also "chest / back".
[0063] Furthermore, the learning unit 52 includes the main processing unit 52A, the evaluation unit 52B, and the update unit 52C. The main processing unit 52A reads the machine learning model LM0, which is the target of the learning process, from the storage device 46 and inputs a pseudo-image SP to the loaded machine learning model LM0. Then, it causes the machine learning model LM0 to execute a process to derive the shooting technique indicated by the input pseudo-image SP. The basic configuration of the machine learning model LM0 is the same as that of the trained model LM described in Figure 5, but in the state before learning, the values such as the filter coefficients in the encoder 37 and the perceptron weights in the classification unit 38 are different from those of the trained model LM.
[0064] The main processing unit 52A outputs the imaging technique derived by the machine learning model LM0 as output data OD to the evaluation unit 52B. The evaluation unit 52B compares the output data OD with the correct data AD included in the training data TD and evaluates the difference between the two as the loss using a loss function. Similar to the trained model LM described above, in the machine learning model LM0 of this example, the multiple imaging techniques shown by the pseudo-image SP are each output as probabilities. Therefore, not only is there a loss if the imaging technique in the output data OD is incorrect, but even if it is correct, if the output value output as the probability is lower than the target value, the difference between the target value and the output value is evaluated as a loss. The evaluation unit 52B outputs the evaluated loss as the evaluation result to the update unit 52C.
[0065] The update unit 52C updates values such as the filter coefficients in the encoder 37 and the perceptron weights in the classification unit 38 so that the loss included in the evaluation result is reduced. The series of processes from input of training data TD to update are repeated until the termination timing arrives. The termination timing may be, for example, when the training of the planned number of training data TD has been completed, or when the loss falls below the target value.
[0066] The method for generating pseudo-images SP performed by the training data generation unit 51 will be explained in detail below using Figures 8 to 15. First, the structure of the 3DCG data 56 will be explained using Figure 8. In this example, the 3DCG data 56 is data for constructing a human body model 56A. In addition to the structure data of the human body model 56A, the 3DCG data 56 includes modeling parameters (called M-parameters) 56B. The human body model 56A is three-dimensional (X, Y, and Z directions) data that mimics the human body. The M-parameters 56B are parameters that can be changed when modeling the human body model 56A, and are parameters for changing at least one of the posture and appearance of the human body model 56A. The M-parameters 56B are parameters specified by the M-parameter specification information 61M.
[0067] Specifically, M-parameter 56B includes various items such as body size information, gender, posture information, skin color, hair color, hairstyle, and clothing. Body size information includes various items such as height, weight, sitting height, inseam length, head circumference, neck circumference, shoulder width, chest circumference, waist circumference, and hand length, wrist circumference, hand width, foot length, foot width, thigh circumference, and calf circumference. Posture information includes items indicating basic postures such as standing, lying down, and sitting, as well as items such as whether or not the limbs are flexed. For limb flexion, for example, the part to be flexed, such as the right hand, left hand, right knee, left knee, and both knees, as well as the direction and angle of flexion, can be specified. Furthermore, M-parameter 56B may also allow specifying internal or external rotation of the knee joint. In this case, the direction of rotation (medial or lateral) and the angle of rotation may also be specified.
[0068] As shown in Figure 4, in the human body model 56A, the flexible joints 57 are pre-defined, and the areas where joints 57 are defined can be flexed. By specifying these M-parameters 56B, it is possible to change the posture and appearance of the human body during the modeling of the human body model 56A. By finely setting the joints 57, it is possible to generate pseudo-images SP with slightly different postures even for the same shooting area. In addition, the number and location of joints 57 in the human body model 56A may be arbitrarily changed. This makes it possible to make the human body model 56A assume complex postures by setting a large number of joints 57. On the other hand, by setting a small number of joints 57, it is possible to reduce the number of M-parameters 56B, which makes it possible to speed up the modeling process.
[0069] As shown in Figure 9, the training data generation unit 51 has a modeling unit 51A and a rendering unit 51B. The modeling unit 51A models a human body model 56A having the posture and appearance of the human body specified by the M-parameter specification information 61M, based on 3DCG data 56 and M-parameter specification information 61M. In Figure 9, the M-parameter specification information 61M includes multiple M-parameter sets MPS. For each M-parameter set MPS, a human body model 56A with a different posture and appearance is generated.
[0070] As shown in Figure 10, as an example, in the M-parameter set MPS1, the body size information specifies a height of 170 cm and a weight of 70 kg, and the gender is specified as male. The posture information specifies an upright position with no flexion of the limbs. Based on this M-parameter set MPS1, the modeling unit 51A generates 3DCG data 56 having a human body model 56A_MPS1 having the posture and appearance specified by the M-parameter set MPS1.
[0071] Furthermore, as shown in Figure 8 with respect to M-parameter 56B, it is possible to change the skin color, hairstyle, and hair color of the human body model 56A by specifying these in the M-parameter set MPS. It is also possible to change the clothing of the human body model 56A by specifying the clothing in the M-parameter set MPS. Thus, M-parameter 56B includes information related to appearance, such as body size and skin color, as well as posture information. Therefore, by changing body size information or information related to appearance, such as skin color, without changing the posture information, it is possible to generate multiple human body models 56A with the same posture but different appearances, or to generate human body models 56A with different postures without changing the information related to appearance. In addition, when the shooting area is the chest, the hairstyle and hair color are often not visible in the chest camera image CP, but when the shooting area is the head, they are visible in the head camera image CP. Therefore, being able to change the hairstyle and hair color is useful when generating a pseudo-image SP of the head.
[0072] The rendering unit 51B can generate a two-dimensional pseudo-image SP by rendering the human body model 56A (human body model 56A_MPS1 in the example of Figures 9 and 10) modeled by the modeling unit 51A from a set viewpoint. In this example, rendering is a process that virtually performs the process of acquiring a camera image CP of the subject H using an optical camera 15, using 3DCG data 56. That is, a virtual camera 15V corresponding to the optical camera 15 is set in the three-dimensional space where the human body model 56A is placed, and a pseudo-image SP is acquired by virtually photographing the human body model 56A which is modeled after the subject H.
[0073] In Figure 9, the R-parameter specification information 61R contains multiple R-parameter sets RPS. A pseudo-image SP is generated for each R-parameter set RPS. The R-parameter set RPS is information that specifies the viewpoint position, etc., when the rendering unit 51B renders the human body model 56A.
[0074] The R-parameter set RPS1 shown in Figures 9 and 10 includes, as an example, imaging procedure information and virtual camera information. The imaging procedure information includes the imaging site and imaging direction. If "chest / back" is specified as the imaging procedure information, the position of the virtual camera 15V is roughly set to a position where the chest of the human body model 56A can be imaged from the back.
[0075] Virtual camera information specifies the position of virtual camera 15V in more detail, and includes settings such as shooting distance, focal length, installation height, and viewing direction. The shooting distance is the distance from virtual camera 15V to human body model 56A. Since virtual camera 15V is a virtual representation of the optical camera 15 attached to the radiation source 11, the shooting distance is set in accordance with the SID (Source to image receptor distance), which is the distance between the radiation source 11 and the electronic cassette 13. The SID is a value that is changed as appropriate depending on the imaging procedure, etc., and in the example in Figure 9, the shooting distance is 90 cm The focal length is set to the following. The focal length is information used to define the field of view of the virtual camera 15V. By defining the shooting distance and focal length, the shooting range SR of the virtual camera 15V (see Figures 11 and 12, etc.) is defined. The shooting range SR is set to be the size that includes the irradiation field of the radiation source 11. In this example, the focal length is set to 40 mm. The installation height is the position in the Z direction where the virtual camera 15V is installed, and is set according to the area being photographed. In this example, since the area being photographed is the chest, it is set to a height that is located at the chest of the standing human body model 56A. The viewpoint direction is set according to the posture of the human body model 56A and the shooting direction of the photography procedure. In the example shown in Figure 10, the human body model 56A is standing and the shooting direction is the back, so it is set to the -Y direction. In this way, the rendering unit 51B can change the viewpoint from which to generate the pseudo-image SP according to the R-parameter set RPS.
[0076] The rendering unit 51B renders the human body model 56A based on the R-parameter set RPS. This generates a pseudo-image SP. In the example in Figure 10, the rendering unit 51B renders the human body model 56A_MPS1, which was modeled by the modeling unit 51A based on the specification of the M-parameter set MPS1, based on the specification of the R-parameter set RPS1, thereby generating the pseudo-image SP1-1.
[0077] Furthermore, in Figure 10, a virtual cassette 13V, which is a virtual electronic cassette 13, is placed in the three-dimensional space where the human body model 56A is positioned. This is shown for convenience to clarify the position of the area being photographed. However, it is also possible to model the virtual cassette 13V in addition to the human body model 56A during the modeling process and then perform rendering including the modeled virtual cassette 13V.
[0078] Figure 11 shows an example where the imaging technique generates a pseudo-image SP of the "chest / front". The M-parameter set MPS1 in Figure 11 is the same as in the example in Figure 10, and the human body model 56A_MPS1 is the same as in the example in Figure 10. The R-parameter set RPS2 in Figure 11 differs from the R-parameter set RPS1 in Figure 10 in that the imaging direction has been changed from "back" to "front". Along with the change in imaging direction, the viewpoint direction of the virtual camera information has been changed to the "Y direction". In other respects, the R-parameter set RPS2 in Figure 11 is the same as the R-parameter set RPS1 in Figure 10.
[0079] Based on this R-parameter set RPS2, the position of the virtual camera 15V and other parameters are set, and rendering is performed. This captures the "chest" of the human body model 56A_MPS1 from the "front," and generates a pseudo-image SP1-2 corresponding to the "chest / front" imaging technique.
[0080] Figure 12 shows an example where the imaging technique generates a pseudo-image SP of the "abdomen / frontal" view. The M-parameter set MPS1 in Figure 12 is the same as in the examples in Figures 10 and 11, and the human body model 56A_MPS1 is the same as in the examples in Figures 10 and 11. The R-parameter set RPS3 in Figure 12 differs from the R-parameter set RPS1 in Figure 10 in that the imaging site has been changed from "chest" to "abdomen". The imaging direction is "frontal", which is different from "back" in Figure 10, but is the same as in the example in Figure 11. In other respects of the R-parameter set RPS3 in Figure 12, it is the same as the R-parameter set RPS2 in Figure 11.
[0081] Based on this R-parameter set RPS3, the position of the virtual camera 15V and other parameters are set, and the human body model 56A_MPS1 is rendered. This generates pseudo-images SP1-3 corresponding to the "chest / frontal" imaging technique of the human body model 56A_MPS1.
[0082] Figure 13 shows an example where the imaging technique generates a pseudo-image SP of "both knees / front view". Unlike the examples in Figures 10 to 13, the M-parameter set MPS2 in Figure 13 includes "sitting" and "both" as posture information. knees The "flexion present" specification is given. As a result, the human body model 56A_MPS2 is in a seated position, both sides knees The model is set to a bent posture. In Figure 13, the R-parameter set RPS4 specifies "both knees" as the imaging site and "front" as the imaging direction in the imaging procedure information. In the virtual camera information, the installation height is specified as 150 cm, which is the height at which both knees of the seated human body model 56A_MPS2 can be photographed from the front, and the imaging direction is specified as downward in the Z direction, i.e., "-Z direction".
[0083] Based on this R-parameter set RPS4, the position of the virtual camera 15V and other parameters are set, and the human body model 56A_MPS2 is rendered. This generates a pseudo-image SP2-4 corresponding to the "both knees / front view" shooting technique of the human body model 56A_MPS2.
[0084] Figure 14, like Figure 11, shows an example where the imaging technique generates a pseudo-image SP of the "chest / frontal" view. In the example in Figure 14, the R-parameter set RPS2 is the same as in the example in Figure 11. The difference is the body size information in the M-parameter set MPS4 in Figure 14. Compared to the M-parameter set MPS1 in Figure 11, the values for weight, as well as chest circumference and waist circumference (not shown), are larger. In other words, the human body model 56A_MPS4 modeled based on the M-parameter set MPS4 in Figure 14 is fatter than the human body model 56A_MPS1 in Figure 11. Other aspects are the same as in the example in Figure 11. As a result, for the fatter human body model 56A_MPS4, a pseudo-image SP4-2 corresponding to the "chest / frontal" imaging technique is generated.
[0085] Figure 15, similar to Figure 13, shows an example where the imaging technique generates a pseudo-image SP of "both knees / front view". The difference from the example in Figure 13 is that the skin color is specified as brown in the M-parameter set MPS5. Otherwise, it is the same as the example in Figure 13. As a result, for the human body model 56A_MPS5 with brown skin color, the pseudo-image SP5-4 of "both knees / front view" is generated.
[0086] The training data generation unit 51 generates multiple pseudo-images SP for each combination of imaging area and imaging direction based on the parameter specification information 61. For example, the training data generation unit 51 generates multiple pseudo-images SP for each imaging technique, such as "chest / front" and "chest / back," using different parameter sets (M-parameter set MPS and R-parameter set RPS).
[0087] Next, the operation of the learning device 40 with the above configuration will be explained using the flowcharts shown in Figures 16 to 18. As shown in the main flowchart of Figure 16, the learning device 40 generates training data TD based on 3DCG data (step S100) and trains the machine learning model LM0 using the generated training data TD (step S200).
[0088] Figure 17 shows the details of step S100 for generating training data TD. As shown in Figure 17, in step S100, first the training data generation unit 51 of the learning device 40 acquires 3DCG data 56 (step S101). Next, the training data generation unit 51 acquires parameter specification information 61. In this example, in steps S101 and S102, the training data generation unit 51 reads the 3DCG data 56 and parameter specification information 61 stored in the storage device 46 from the storage device 46.
[0089] The parameter specification information 61 records, for example, multiple M-parameter sets MPS and multiple R-parameter sets RPS. The training data generation unit 51 generates a pseudo-image SP by performing modeling and rendering for each combination of one M-parameter set MPS and one R-parameter set RPS, as illustrated in Figures 10 to 15 (step S103). The training data generation unit 51 generates training data TD by combining the generated pseudo-image SP with the ground truth data AD (step S104). The generated training data TD is stored in the storage device 46.
[0090] In step S105, the training data generation unit 51 executes the processes in steps S103 and S104 if there are any unentered parameter sets (combinations of M-parameter set MPS and R-parameter set RPS) among the parameter sets included in the acquired parameter specification information 61 (YES in step S105). The training data generation unit 51 generates multiple pseudo-images SP using different parameter sets (M-parameter set MPS and R-parameter set RPS) for each imaging technique such as "chest / front" and "chest / back". If there are no unentered parameter sets (NO in step S105), the unit executes the processes in steps S103 and S104. )、 The training data generation process is now complete.
[0091] Figure 18 shows the details of step S200, in which the machine learning model LM0 is trained using the generated training data TD. As shown in Figure 18, in step S200, first, in the learning unit 52 of the learning device 40, the main processing unit 52A acquires the machine learning model LM0 from the storage device 46 (step S201). Then, the main processing unit 52A acquires the training data TD generated by the training data generation unit 51 from the storage device 46 and inputs the pseudo-images SP contained in the training data TD into the machine learning model LM0 one by one (step S202). Then, the main processing unit 52A causes the machine learning model LM0 to derive output data OD including the shooting technique (step S203). Next, the evaluation unit 52B compares the correct data AD with the output data OD and evaluates the output data OD (step S204). The update unit 52C updates the values of the machine learning model LM0, such as the filter coefficients and perceptron weights, based on the evaluation result of the output data OD (step S205). Until the termination timing arrives (YES in step S206), the learning unit 52 repeats the series of processes from step S202 to step S205. When the termination timing arrives, such as when the learning of the planned number of training data TDs is completed, or when the loss falls below the target value (NO in step S206), the learning unit 52 terminates the learning and outputs the trained model LM (step S207).
[0092] As described above, the learning device 40 according to this embodiment generates multiple pseudo-images SP of a human body, each for each combination of imaging area and imaging direction. These pseudo-images SP are generated based on a human body model 56A composed of 3DCG data and mimic a subject H positioned in front of a radiography system 10, which is an example of a medical imaging device. The machine learning model LM0 is then trained using multiple training data sets TD, which consist of the generated pseudo-images SP and the correct data AD of the combinations. The pseudo-images SP that make up the training data sets TD are generated based on a human body model 56A composed of 3DCG data. Therefore, it is easier to increase the number of training data sets TD that represent imaging techniques compared to when camera images CP, which represent imaging techniques and are combinations of imaging areas and imaging directions, are acquired by an optical camera 15. As a result, the trained model LM, which derives the imaging area and imaging direction of the subject H captured in the camera images CP, can be trained more efficiently compared to when only camera images CP are used as training data TD.
[0093] Furthermore, the 3DCG data in the above example includes M-parameters 56B for changing at least one of the posture and appearance of the human body model 56A. Therefore, for example, by simply changing the M-parameters 56B using parameter specification information 61, it is possible to generate multiple pseudo-images SP in which at least one of the posture and appearance of the human body model 56A differs. Thus, it is easy to generate a variety of pseudo-images SP in which at least one of the posture or appearance of the human body model 56A differs. The posture and appearance of actual subjects H also vary from person to person. Collecting camera images CP of various subjects H or camera images CP of subjects H in various postures is extremely time-consuming. In light of these circumstances, the method of easily changing the posture or appearance of the human body model 56A by changing the M-parameters 56B is extremely effective.
[0094] Furthermore, the M-parameter 56B includes at least one of the following: body size information, gender, posture information, skin color, hair color, hairstyle, and clothing of the human body model 56A. In the technology of this disclosure, by making it possible to specify appearance-related information and posture information in detail, it is possible to easily collect images of human bodies with appearances or postures that are difficult to collect with actual camera images CP. For example, even if the shooting area and shooting direction are the same, it is possible to train the machine learning model LM0 using pseudo-images SP of human body models 56A that mimic subjects H with diverse appearances as training data TD. This improves the accuracy of deriving the shooting area and shooting direction from the camera image CP in the trained model LM compared to the case where it is not possible to specify body size information, etc.
[0095] In the example above, body measurements were explained as including height, weight, sitting height, inseam length, head circumference, head width, neck circumference, shoulder width, chest circumference (bust), waist circumference, arm length, wrist circumference, hand width, foot length, foot width, thigh circumference, and calf circumference. However, it is sufficient to include at least one of these. Of course, the more body measurements included, the greater the diversity of the human body model 56A, which is preferable. Also, although the example explained used numerical values for height and weight, evaluation information such as slim, standard, and chubby is also acceptable.
[0096] Furthermore, in this embodiment, the rendering unit 51B can generate a pseudo-image SP by rendering the human body model 56A from a set viewpoint, and the viewpoint can be changed by the R-parameter set RPS which defines the rendering parameters. Therefore, compared to the case where such rendering parameters are not present, it becomes easier to change the viewpoint and to generate a variety of pseudo-image SPs.
[0097] The viewpoint information used to set the viewpoint includes the focal length of the virtual camera 15V, which is virtually placed at the viewpoint, and the shooting distance, which is the distance from the virtual camera 15V to the human body model 56A. Since the shooting distance and focal length are changeable, it is easy to change the shooting range (SR) of the virtual camera 15V.
[0098] In the above embodiment, medical for The example described here is an application of the technology of this disclosure to a radiography system 10, which is an example of an imaging device and a radiography device. In radiography, the categories are "chest / front," "abdomen / front," and "both." knees There are many types of imaging orders that specify combinations of imaging area and imaging direction, such as "front view." This invention is particularly effective when used with medical imaging equipment that has such a wide variety of imaging order types.
[0099] In the above example, batch processing may be performed to generate multiple pseudo-images SP sequentially using the parameter specification information 61 shown in Figure 9. The parameter specification information 91 includes M-parameter specification information 61M and R-parameter specification information 61R. As shown in Figure 9, it is possible to specify multiple M-parameter sets MPS in the M-parameter specification information 61M, and multiple R-parameter sets RPS in the R-parameter specification information 61R. By specifying a pair of M-parameter set MPS and R-parameter set RPS, one pseudo-image SP is generated. Therefore, by using the parameter specification information 91, it is possible to perform batch processing to generate multiple pseudo-images SP sequentially. Such batch processing is possible when generating pseudo-images SP using 3DCG data 56. By performing batch processing, it is possible to efficiently generate multiple pseudo-images SP.
[0100] Furthermore, although the above example describes an example where the optical camera 15 takes still images, it may also take video. If the optical camera 15 is to take video, the technician RG may input a shooting instruction before the positioning of the subject H begins, causing the optical camera 15 to start shooting video and transmit the video captured by the optical camera 15 to the console 14 in real time. In this case, for example, the frame images that make up the video are used as input data for the trained model LM.
[0101] "Variation 1" In the example above, we explained an example where only the pseudo-image SP is used as the training data TD. However, as shown in Figure 19, the machine learning model LM0 may also be trained by mixing the camera image CP with the pseudo-image SP. Here, if the training data TD, which consists of the pseudo-image SP and the ground truth data AD, is used as the first training data, then the training data consisting of the camera image CP and the ground truth data AD is used as the second training data TD2. The learning unit 52 trains the machine learning model LM0 using both the first training data TD and the second training data TD2. By training using the camera image CP that is used when the trained model LM is in operation, the accuracy of deriving the shooting area and shooting direction is improved compared to the case where only the pseudo-image SP is used.
[0102] "Variation 2" In the above example, as shown in Figure 4, during the operation phase of the trained model LM, the imaging procedure determination unit 14B outputs the imaging procedure included in the imaging order 31 ("chest / back," etc.) and the imaging procedure derived by the trained model LM ("chest / back," etc.) directly to the display 14C. However, as shown in Figure 20, a comparison unit 14D may be provided in the imaging procedure determination unit 14B so that the comparison unit 14D compares the imaging procedure included in the imaging order 31 with the imaging procedure derived by the trained model LM. In this case, the comparison unit 14D may output a comparison result indicating whether the two imaging procedures match or not, and display the comparison result on the positioning confirmation screen 36. In the example in Figure 20, the comparison result is displayed in the form of a message such as "The determined imaging procedure matches the imaging order." Furthermore, if the matching result indicates that the imaging procedures of both parties do not match, a message (not shown) such as "The determined imaging procedure does not match the imaging order" will be displayed on the positioning confirmation screen 36 as the matching result. Such a message functions as an alert to make the technician RG aware of a positioning error. In addition, voice, warning sounds, and warning lights may also be used as alerts.
[0103] In the above embodiment, the learning device 40 has a teacher data generation unit 51, and the learning device 40 also functions as a teacher data generation device. However, the teacher data generation unit 51 may be separated from the learning device 40 and operated as an independent device.
[0104] "Second Embodiment" The second embodiment shown in Figures 21 and 22 considers a pre-trained model LMT for body size output that takes a camera image CP as input and derives body size information representing the body size of the subject H captured in the camera image CP. In the second embodiment, in order to generate such a pre-trained model LMT for body size output, for This is an example of using a pseudo-image SP to train the machine learning model LMT0 for physique output, which is the basis for the trained model LMT. As shown in Figure 21, the trained model LMT for physique output includes, for example, an encoder 37 and a regression unit 68. The encoder 37 is the same as the trained model LM that derives the imaging technique shown in Figure 5. The trained model LMT for physique output is a regression model that estimates, for example, a numerical value of body thickness as physique information of the subject H captured in the camera image CP from the feature map of the camera image CP. For the regression unit 68, for example, a linear regression model or a support vector machine can be used.
[0105] As shown in Figure 22, the learning unit 52 trains the machine learning model LMT0 for physique output using multiple training data TDT for physique output using pseudo-image SP. The training data TDT for physique output consists of pseudo-image SP and ground truth data ADT representing the physique of the human body model 56A.
[0106] As shown in Figure 3, the body thickness of subject H is used as basic information when determining the irradiation conditions. Therefore, if body thickness can be derived from the camera image CP, convenience in radiography will be improved.
[0107] "Third Embodiment" The third embodiment shown in Figure 23 is an example in which the technology of this disclosure is applied to an ultrasound imaging device as a medical imaging device. The ultrasound imaging device has a probe 71 equipped with an ultrasound transmitter and a receiving unit. Ultrasound imaging is performed by bringing the probe 71 into contact with the imaging site of the subject H. In ultrasound imaging, the imaging site is specified by the imaging order, and the imaging direction (orientation of the probe 71) may also be determined according to the purpose of imaging. If the imaging site and imaging direction are inappropriate, it may not be possible to obtain a suitable ultrasound image.
[0108] In particular, unlike radiography, ultrasound imaging can be performed not only by physicians but also by nurses and caregivers, and in the future, it is anticipated that patients H, who are unfamiliar with ultrasound imaging, may operate the probe 71 themselves to perform ultrasound imaging. In this case, the use of camera images CP for imaging guidance is also being considered. In such cases, it is expected that a trained model LM, which derives the imaging site and direction from the camera image CP as input, will be used in many situations. That is, in ultrasound imaging, the camera image CP is an image taken using the optical camera 15 of the state in which the probe 71 is applied to the imaging site of the patient H, that is, the state of the patient H positioned relative to the probe 71. In order to ensure that the derivation accuracy of such a trained model LM is accurate enough to withstand the use of imaging guidance, it is important to train the trained model LM using a wide variety of training data TD. By using the technology disclosed herein, it is possible to efficiently train the model using training data TD of a wide variety of pseudo-images SP.
[0109] As shown in Figure 23, even when the technology of this disclosure is used for ultrasound imaging, the training data generation unit 51 generates a pseudo-image SP by rendering the positioning of the human body model 56A and the virtual probe 71 at a viewpoint set by the virtual camera 15V. As shown in Figure 7, the machine learning model LM is trained using the training data TD of the generated pseudo-image SP.
[0110] In each of the above embodiments, the hardware structure of the processing unit that performs various processes, such as the training data generation unit 51 and the learning unit 52, is a processor of the following types.
[0111] Various types of processors include CPUs, programmable logic devices (PLDs), and dedicated electrical circuits. A CPU, as is well known, is a general-purpose processor that executes software (programs) and functions as various processing units. A PLD, such as an FPGA (Field Programmable Gate Array), is a processor whose circuit configuration can be changed after manufacturing. Dedicated electrical circuits are processors with circuit configurations specifically designed to perform particular processing, such as an ASIC (Application Specific Integrated Circuit).
[0112] A single processing unit may be composed of one of these various processors, or it may be composed of a combination of two or more processors of the same or different types (for example, multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, multiple processing units may be composed of a single processor. This is also possible. Examples of configuring multiple processing units with a single processor include, firstly, a configuration in which one processor is composed of a combination of one or more CPUs and software, and this processor functions as multiple processing units. Secondly, a configuration using a processor that realizes the functions of the entire system, including multiple processing units, on a single IC chip, as exemplified by a System on a Chip (SoC). Thus, various processing units are configured as hardware structures using one or more of the above-mentioned various processors.
[0113] Furthermore, the hardware structure of these various processors is, more specifically, an electrical circuit composed of circuit elements such as semiconductor devices.
[0114] The technology disclosed herein is not limited to the embodiments described above, and various configurations can be adopted as long as they do not deviate from the gist of the technology disclosed herein. Furthermore, the technology disclosed herein extends not only to programs but also to computer-readable storage media for non-temporarily storing programs.
[0115] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
Claims
1. A learning device that takes a camera image taken with an optical camera of a subject positioned in front of a medical imaging device as input, and trains a machine learning model to derive the imaging area and imaging direction of the subject as captured in the camera image, It comprises a processor and memory connected to or built into the processor, The aforementioned processor, A simulated image of a human body generated based on a human body model composed of three-dimensional computer graphics data, wherein multiple simulated images are generated for each combination of the imaging area and imaging direction, each simulating the subject positioned relative to the medical imaging device. A learning device that trains a machine learning model using multiple training data sets, each consisting of the generated pseudo-images and the correct data for the combinations.
2. The learning device according to claim 1, wherein the three-dimensional computer graphics data is accompanied by modeling parameters for changing at least one of the posture and appearance of the human body model.
3. The learning device according to claim 2, wherein the modeling parameters include at least one of body size information representing the physique of the human body model, gender, posture information, skin color, hair color, hairstyle, and clothing.
4. The learning device according to any one of claims 1 to 3, wherein the processor is capable of generating the pseudo-image by rendering the human body model from a set viewpoint, and the viewpoint is changeable by rendering parameters.
5. The learning device according to claim 4, wherein the viewpoint information for setting the viewpoint includes the focal length of a virtual camera virtually installed at the viewpoint and the shooting distance, which is the distance from the virtual camera to the human body model.
6. When the training data consisting of the pseudo-image and the correct data is designated as the first training data, The aforementioned processor, The learning device according to any one of claims 1 to 5, wherein the machine learning model is trained using, in addition to the first training data, second training data consisting of camera images captured by an optical camera and the correct answer data.
7. Furthermore, the learning device according to any one of claims 1 to 6, which uses a plurality of training data for body size output, consisting of the pseudo-image and the correct data of body size information representing the body size of the human body model, to train a machine learning model for body size output that derives body size information representing the body size of the subject from the camera image as input.
8. The learning device according to any one of claims 1 to 7, wherein the medical imaging device includes at least one of a radiography device and an ultrasound imaging device.
9. A learning method for training a machine learning model using a computer, which takes a camera image taken with an optical camera of a subject positioned in front of a medical imaging device as input, and derives the imaging area and imaging direction of the subject as captured in the camera image, A simulated image of a human body generated based on a human body model composed of three-dimensional computer graphics data, wherein multiple simulated images are generated for each combination of the imaging area and imaging direction, each simulating the subject positioned relative to the medical imaging device. A learning method for training a machine learning model using multiple training data sets, each consisting of the generated pseudo-images and the correct data for the combinations.
10. An operating program for a learning device that causes a computer to function as a learning device, which takes a camera image taken with an optical camera of a subject positioned in front of a medical imaging device as input, and trains a machine learning model to derive the imaging area and imaging direction of the subject as captured in the camera image, A simulated image of a human body generated based on a human body model composed of three-dimensional computer graphics data, wherein multiple simulated images are generated for each combination of the imaging area and imaging direction, each simulating the subject positioned relative to the medical imaging device. An operating program for a learning device that causes a computer to function as a learning device for training the machine learning model using a plurality of training data consisting of the generated pseudo-images and the correct data for the combinations.
11. A training data generation device that generates training data for training a machine learning model that derives the imaging area and imaging direction of the subject captured in a camera image, taking a camera image of the subject positioned in front of a medical imaging device as input, It comprises a processor and memory connected to or built into the processor, The aforementioned processor, Three-dimensional computer graphics data comprising a human body model for generating a simulated image of a human body that mimics the subject positioned in relation to the medical imaging device, wherein the three-dimensional computer graphics data includes parameters for changing at least one of the posture and appearance of the human body model, By changing the aforementioned parameters, a plurality of pseudo-images in which at least one of the posture and appearance of the human body model differs are generated for each combination of the shooting area and shooting direction. A training data generation device that generates multiple training data sets, each consisting of a plurality of generated pseudo-images and correct data for the combination thereof.