Learning device, learning method, recording medium, teacher data generation device, machine learning model, and medical imaging device

By generating a combination of pseudo-images of human body models based on 3D computer graphics data and camera images, and training a machine learning model, the problem of low efficiency in identifying the imaging parts and directions of subjects in medical imaging devices is solved, thereby improving the recognition accuracy and reducing labor and time costs.

CN116829074BActive Publication Date: 2026-03-24FUJIFILM CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2026-03-24

Smart Images

  • Figure CN116829074B_ABST
    Figure CN116829074B_ABST
Patent Text Reader

Abstract

A learning device that causes a machine learning model to learn, the machine learning model deriving a photographic site of a subject and a photographic direction of the subject reflected in a camera image, using a camera image obtained by photographing the subject in a state positioned with respect to a photographic device for medical use as input, the learning device having a processor and a memory connected to or built into the processor, the processor performing processing of generating a plurality of pseudo images of a human body, i.e., pseudo images simulating a subject in a state positioned with respect to a virtual dark box photographic device for medical use, from a human body model generated from three-dimensional computer graphics data, for each combination of the photographic site and the photographic direction, and causing the machine learning model to learn using a plurality of teacher data composed of the generated pseudo images and correct data for the combination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning device, a learning method, a working procedure of the learning device, a teacher data generation device, a machine learning model, and a medical imaging device. Background Technology

[0002] As a medical imaging device, radiographic devices are known, for example. During radiographic imaging, the technician adjusts the relative position of the radiographic device to the patient based on imaging instructions obtained from the physician who commissioned the imaging (e.g., Japanese Patent Application Publication No. 2020-192440). These instructions specify the imaging site, such as the chest or abdomen, and the imaging direction, such as frontal or back views. The technician positions the patient according to these instructions.

[0003] Japanese Patent Application Publication No. 2020-192440 discloses the following technology: using an optical camera to photograph a subject in a position relative to a radiographic device, and allowing a technician to confirm a composite image formed by combining the photographed camera image with a marker representing the ideal position, thereby providing appropriate positioning assistance. Summary of the Invention

[0004] The technical problem to be solved by the invention

[0005] Research is underway to utilize camera images, as described above, to verify whether the photographic areas and directions of the subject conform to the photographic instructions. Specifically, research is being conducted to use camera images as input and derive the photographic areas and directions of the subject as reflected in the camera images into a machine learning model. This can suppress errors such as photographing areas that differ from those specified in the photographic instructions. In the case of medical photography, subjects often have limited mobility, and re-photographing can be burdensome; therefore, the need to suppress re-photographing is higher compared to other applications. Thus, when using such a machine learning model in medical photography, a machine learning model with high accuracy in recognizing the photographic areas and directions of the subject reflected in camera images is desired.

[0006] To improve the recognition accuracy of machine learning models, the larger the amount of teacher data, consisting of correct combinations of camera images and shooting locations and directions, the better. In particular, the more diverse the camera images in the teacher data for each combination of shooting locations and directions, the higher the recognition accuracy of the machine learning model. Even with the same combination of shooting locations and directions, the subject's posture or appearance reflected in the camera images can vary. Even when the subject's posture or appearance differs, the machine learning model needs to recognize it as the same combination of shooting locations and directions. Therefore, for the camera images used in the teacher data, it is necessary to collect a large number of diverse camera images of subjects with different postures and appearances for each combination of shooting locations and directions.

[0007] However, collecting large amounts of such camera images is very labor-intensive and time-consuming, so there is a need for methods to efficiently improve the recognition accuracy of machine learning models.

[0008] The present invention provides a learning device, a learning method, a working procedure for the learning device, a teacher data generation device, a machine learning model, and a medical imaging device that enable a machine learning model to efficiently learn the photographic parts and photographic directions of the subject as reflected in the camera images, compared to the case where only camera images are used as teacher data.

[0009] means for solving technical problems

[0010] To achieve the above objectives, the learning device of the present invention enables a machine learning model to learn. The machine learning model takes as input camera images obtained by taking pictures of a subject positioned relative to a medical imaging device using an optical camera, and derives the imaging parts and imaging directions of the subject reflected in the camera images. The learning device has a processor and a memory connected to or built into the processor. The processor performs the following processing: for each combination of imaging parts and imaging directions, it generates multiple pseudo-images of the human body generated based on a human body model composed of three-dimensional computer graphics data, i.e., pseudo-images simulating the subject's position relative to the medical imaging device; and uses multiple teacher data composed of the generated pseudo-images and combined correct data to enable the machine learning model to learn.

[0011] Three-dimensional computer graphics data may include modeling parameters for changing at least one of the pose and appearance of the human body model.

[0012] Modeling parameters may include at least one of the following: physical information representing the human body model, gender, posture information, skin color, hair color, hairstyle, and clothing.

[0013] The processor can generate pseudo-images by rendering a human model from a set viewpoint, which can be changed through rendering parameters.

[0014] The viewpoint information used to set the viewpoint may include the focal length of a virtual camera virtually set at the viewpoint and the photographic distance, which is the distance from the virtual camera to the human model.

[0015] When the teacher data consisting of pseudo-images and correct data is used as the first teacher data, the processor can process it as follows: in addition to the first teacher data, a second teacher data consisting of camera images taken with an optical camera and correct data is used to enable the machine learning model to learn.

[0016] Multiple physique outputs consisting of pseudo-images and correct data representing the physique of a human model can also be used with teacher data, enabling the physique outputs that derive physique information representing the subject's physique from camera images as input to be learned by a machine learning model.

[0017] Medical imaging devices may include at least one of radiographic devices and ultrasound imaging devices.

[0018] The learning method of the present invention uses a computer to enable a machine learning model to learn. The machine learning model takes camera images obtained by taking pictures of a subject in a position relative to a medical imaging device as input to derive the imaging parts and imaging directions of the subject reflected in the camera images. For each combination of imaging parts and imaging directions, multiple pseudo-images of the human body generated based on a human body model composed of three-dimensional computer graphics data are generated, i.e., pseudo-images of the subject in a position relative to the medical imaging device. Multiple teacher data composed of the generated pseudo-images and the correct data of the combination are used to enable the machine learning model to learn.

[0019] A learning device operating procedure enables a computer to function as a learning device, wherein the learning device enables a machine learning model to learn, the machine learning model taking as input camera images obtained by taking pictures of a subject in a position relative to a medical imaging device using an optical camera to derive the imaging parts and imaging directions of the subject reflected in the camera images, wherein multiple pseudo-images of the human body generated based on a human body model composed of three-dimensional computer graphics data are generated for each combination of imaging parts and imaging directions, i.e., pseudo-images of the subject in a position relative to the medical imaging device, and the machine learning model learns using multiple teacher data composed of the generated pseudo-images and combined correct data.

[0020] The teacher data generation device of the present invention generates teacher data, which is used to enable a machine learning model to learn. The machine learning model takes as input camera images obtained by taking pictures of a subject in a position relative to a medical imaging device using an optical camera to derive the imaging parts and imaging directions of the subject reflected in the camera images. The learning device has a processor and a memory connected to or built into the processor. The processor performs the following processing: using three-dimensional computer graphics data, which constitutes a human body model for generating pseudo-images of the human body and includes parameters for changing at least one of the posture and appearance of the human body model, by changing the parameters, generating multiple pseudo-images of the human body model with different postures and appearances for each combination of imaging parts and imaging directions, and generating multiple teacher data consisting of the generated multiple pseudo-images and the correct data of the combination.

[0021] The machine learning model of this invention takes camera images obtained by taking pictures of a subject in a position relative to a medical imaging device with an optical camera as input to derive the imaging parts and imaging directions of the subject reflected in the camera images. The machine learning model is trained using multiple teacher data, which consists of correct data consisting of combinations of pseudo images and imaging parts and imaging directions, and is generated for each combination. The pseudo images are pseudo images of the human body generated based on a human body model composed of three-dimensional computer graphics data, and are pseudo images simulating the position of the subject relative to the medical imaging device.

[0022] The medical imaging device of the present invention has a machine learning model.

[0023] Invention Effects

[0024] According to the technology of the present invention, compared with the case of using only camera images as teacher data, machine learning models that derive the photographic parts and photographic directions of the subject reflected in the camera images can be learned more efficiently. Attached Figure Description

[0025] Figure 1 It is a diagram showing the general structure of the learning device and the radiographic system.

[0026] Figure 2 This is a diagram illustrating the shooting instructions.

[0027] Figure 3 This is a diagram illustrating the method for determining irradiation conditions.

[0028] Figure 4 This is a diagram illustrating the functionality of the console.

[0029] Figure 5 This is a diagram illustrating the completed learning process of the model.

[0030] Figure 6 This is a diagram showing the hardware structure of the learning device.

[0031] Figure 7 This is a diagram illustrating the function of the learning device.

[0032] Figure 8 This is a diagram representing an example of 3DCG data.

[0033] Figure 9 This is a diagram representing the functions of the teacher data generation department.

[0034] Figure 10 This is a diagram representing an example of a pseudo-image.

[0035] Figure 11 This is a diagram representing an example of a pseudo-image.

[0036] Figure 12 This is a diagram representing an example of a pseudo-image.

[0037] Figure 13 This is a diagram representing an example of a pseudo-image.

[0038] Figure 14 This is a diagram representing an example of a pseudo-image.

[0039] Figure 15 This is a diagram representing an example of a pseudo-image.

[0040] Figure 16 This is the main flowchart representing the processing sequence of the learning device.

[0041] Figure 17 This is a flowchart showing the order in which teacher data is generated.

[0042] Figure 18 It is a flowchart representing the learning sequence.

[0043] Figure 19 This is a diagram representing a variation where teacher data, pseudo-images, and camera images coexist.

[0044] Figure 20 This is a diagram showing a modified example with an added verification section.

[0045] Figure 21 This is a diagram illustrating the physical output using the learned model.

[0046] Figure 22 This is a graph illustrating teacher data used in machine learning models for physical output.

[0047] Figure 23 This is a diagram illustrating the teacher data generation section applicable to ultrasonic imaging devices. Detailed Implementation

[0048] Figure 1 This is a schematic diagram showing the overall structure of the learning device 40 and the radiography system 10 of the present invention. The radiography system 10 is an example of a medical imaging device according to the technology of the present invention, and is an example of a radiography device. The radiography system 10 obtains a radiographic image XP of the subject H by taking a picture of the subject H using radiation R. Furthermore, when performing radiography using the radiography system 10, the radiographer (hereinafter referred to as the technician) RG, as the operator, positions the subject H relative to the radiography system 10.

[0049] The radiography system 10 in this example has a radiographic assistance function that assists the technician RG in confirming whether the subject H is properly positioned according to the radiographic instructions 31. Details will be described later. The radiographic assistance function uses a learned model LM generated by training the machine learning model LMO. The learning device 40 has a learning unit 52 for training the machine learning model LMO. The learning unit 52 is used to train the machine learning model LMO to generate the learned model LM provided to the radiography system 10.

[0050] Here, the learned model LM is also a machine learning model, but for convenience, it is distinguished from the machine learning model LMO that becomes the object of learning by the learning device 40. The machine learning model that has been learned by the learning device 40 at least once and used in the radiography system 10 is called the learned model LM. In addition, the machine learning model LMO can be an unlearned machine learning model or a learned model LM that becomes the object of additional learning. Hereinafter, after explaining the general overview of the radiography system 10, the learning device 40 will be described.

[0051] like Figure 1 As shown, the radiography system 10 includes a radiation source 11, a radiation source control device 12, an electronic cassette 13, and a control console 14. The electronic cassette 13 is an example of a portable radiation image detector that detects the radiation image XP of the subject H by receiving the radiation R transmitted through the subject H. Furthermore, the radiography system 10 in this example includes an optical camera 15. The optical camera 15 is a structure used to realize the aforementioned radiographic assistance functions, and is used to capture the state of the subject H positioned relative to the radiography system 10.

[0052] Positioning during radiography involves adjusting the relative positions of the patient H, the electronic cassette 13, and the radiation source 11. As an example, the technician RG first aligns the electronic cassette 13 with the area of ​​the patient H to be radiographed. Figure 1In this example, the electronic cassette 13 is configured to face the chest of the subject H. Furthermore, the position of the radiation source 11 is adjusted so that the electronic cassette 13, positioned at the radiographic site of the subject H, is opposite the radiation source 11. And, in Figure 1 In this example, the radiation source 11 is positioned opposite the back of the patient H's chest, and the direction of the radiation R, i.e., the imaging direction, is the back of the patient H. In the case of imaging the chest from the front, the technician RG positions the front of the patient H's chest opposite the radiation source 11. By positioning the patient in this way, the radiation R is irradiated from the back of the patient H, thereby enabling the imaging of the patient H's chest using radiation image XP.

[0053] exist Figure 1 In this example, the electronic cassette 13 is mounted on a standing radiography table 25 for photographing the subject H in a standing position. The electronic cassette 13 can also be mounted on a supine radiography table or similar device for photographing the subject H in a supine position. Furthermore, since the electronic cassette 13 is portable, it can also be detached from the radiography table and used as a standalone device.

[0054] The radiation source 11 includes a radiation tube 11A that generates radiation R and an irradiation field limiter 11B that defines the area where radiation R is irradiated, i.e., the irradiation field. The radiation tube 11A, for example, has a filament that emits thermionic electrons and a target that emits radiation by colliding with the thermionic electrons emitted from the filament. The irradiation field limiter 11B, for example, consists of four lead plates arranged on each side of a quadrilateral to shield radiation R, forming a quadrilateral irradiation opening in the center through which radiation R is transmitted. The size of the irradiation opening is changed by moving the positions of the lead plates in the irradiation field limiter 11B, thereby adjusting the size of the irradiation field. In addition to the radiation tube 11A, the radiation source 11 may also include an irradiation field display light source (not shown) for projecting visible light onto the subject H through the irradiation opening, thereby visualizing the irradiation field.

[0055] exist Figure 1 In this example, the radiation source 11 is a ceiling-suspended type, mounted on a retractable support column 22. Regarding the radiation source 11, its vertical height can be adjusted by extending and retracting the support column 22. Furthermore, the support column 22 is mounted on a ceiling-mounted traveling device (not shown) that travels on a track disposed on the ceiling, allowing it to move horizontally along the track. Moreover, the radiation source 11 can rotate about the focal point of the radiation tube 11A. The direction of the irradiated radiation R can be adjusted by various displacement mechanisms.

[0056] The radiation source control device 12 controls the radiation source 11. An operation panel (not shown) is provided in the radiation source control device 12. Technician RG operates the operation panel to set the radiation irradiation conditions and the size of the irradiation opening of the irradiation field limiter 11B. The radiation irradiation conditions include the tube voltage (kV), tube current (mA), and radiation irradiation time (ms) applied to the radiation source 11.

[0057] The radiation source control device 12 includes a voltage generating unit that generates a voltage applied to the radiation tube 11A and a timer. By controlling the voltage generating unit and the timer, the radiation source control device 12 activates the radiation source 11A to generate radiation R corresponding to the irradiation conditions. Furthermore, an irradiation switch 16 is connected to the radiation source control device 12 via a cable or the like. The irradiation switch 16 is operated by technician RG when irradiation begins. When the irradiation switch 16 is operated, the radiation source control device 12 causes the radiation tube 11A to generate radiation. Thus, radiation R is irradiated into the irradiation field.

[0058] As described above, the electronic cassette 13 detects a radiographic image XP based on radiation R irradiated and transmitted from the radiation source 11 to the radiographic site of the subject H. As an example, the electronic cassette 13 has a wireless communication unit and a battery, enabling it to operate wirelessly. The electronic cassette 13 wirelessly transmits the detected radiographic image XP to the control console 14.

[0059] The optical camera 15 is an optical digital camera comprising an CMOS (Complementary Metal Oxide Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and performs visible light-based photography as an example. As an example, the optical camera 15 is capable of both still image photography and video recording.

[0060] The optical camera 15 is used to photograph the subject H in a position relative to the electronic cassette 13 and the radiation source 11. Therefore, as an example, the optical camera 15 is mounted on the outer periphery of the irradiation field limiter 11B of the radiation source 11, positioned near the irradiation opening. Furthermore, the optical camera 15 is mounted with its optical axis parallel to the irradiation axis of the radiation source 11. The optical camera 15 generates a camera image CP as a visible light-based optical image by optically photographing the area containing the irradiation field of the radiation R. Since the radiation source 11 is positioned at the photographic site of the subject H, the photographic site of the subject H is depicted in the camera image CP taken in this state. In this example, the camera image CP is, for example, a color still image.

[0061] In this example, the optical camera 15 is mounted on the outer periphery of the irradiation field limiter 11B, but the optical camera 15 may not be mounted on the outer periphery of the radiation source 11, but may be built into the radiation source 11.

[0062] The optical camera 15 is connected to the control console 14 via wired or wireless means. The control console 14 functions as a control device for the optical camera 15, controlling its photographic actions such as timing. For example, the technician RG inputs the photographic instructions for the optical camera 15 into the control console 14 while the subject H is positioned.

[0063] The console 14 is connected via network N to the RIS (Radiology Information System) and PACS (Picture Archiving and Communication System) located in the radiography system 10.

[0064] The RIS (Radio Instruction System) is a device that manages imaging commands 31 for the radiography system 10. For example, in a medical facility, physicians from departments such as internal medicine and surgery commission a radiology department to perform radiography. The physicians from the internal medicine departments issue imaging commands 31 to the radiology department. The RIS manages these imaging commands 31 from the physicians. In the presence of multiple radiography systems 10, the multiple imaging commands 31 managed by the RIS are distributed to the multiple radiography systems 10 according to the content of the imaging commands and the operating status of the radiography systems 10. The control console 14 receives the imaging commands 31 sent from the RIS.

[0065] Figure 2 An example of imaging instruction 31 is shown. Imaging instruction 31 includes an instruction ID (Identification Data) for each instruction, a subject ID for each subject H, the imaging technique, and the imaging purpose (not shown). Here, the imaging technique refers to an imaging method defined at least according to the combination of the imaging site and the imaging direction. In imaging instruction 31, the physician specifies the imaging technique, including the imaging site and the imaging direction, according to the examination purpose of the subject H. For example, in the case of diagnosing lung cancer, the imaging site is specified as the chest, and the imaging direction is specified as either the back or the front. Figure 2In the photographic instruction 31 with instruction ID "N0001", "chest / back" was specified as the photographic technique. In the photographic instruction 31 with instruction ID "N0002", "chest / front" was specified as the photographic technique. In the photographic instruction 31 with instruction ID "N0003", "abdomen / front" was specified as the photographic technique. In the photographic instruction 31 with instruction ID "N0004", "both knees / front" was specified as the photographic technique. In the photographic instruction 31 with instruction ID "N0005", "right knee / side view" was specified as the photographic technique.

[0066] And, as Figure 3 As shown, information from the radiographic technique, along with the subject H's physical information, is used by technician RG to determine the irradiation conditions of radiation source 11. For example, when the radiographic technique is "chest / back," technician RG considers the subject H's physical information (primarily body thickness), estimates the chest thickness, and determines the tube voltage, tube current, and irradiation time. Generally, the greater the body thickness, the lower the transmittance of radiation R, therefore the higher the radiation dose of radiation R is set. The radiation dose is limited by the mAs value, which is the product of tube current and irradiation time.

[0067] Return to Figure 1 The PACS stores the radiographic images XP captured by the radiography system 10. The console 14 sends the radiographic images XP received from the electronic cassette 13 to the PACS in a state associated with the imaging instruction 31. In the PACS, the radiographic images XP are saved, for example, as image files converted to a format conforming to the DICOM (Digital Imaging and Communication in Medicine) standard. The radiographic images XP saved in the PACS are made available for viewing by physicians and other personnel in the clinic of the client who issued the imaging instruction 31.

[0068] The control console 14 is, for example, a computer such as a personal computer or a workstation. The control console 14 has a command receiving function for receiving photographic instructions 31, a setting function for making various settings of the electronic cassette 13, and a function for displaying the radiographic images XP received from the electronic cassette 13 on the monitor 14C. In addition to these basic functions, the control console 14 also has the aforementioned photographic assistance functions.

[0069] Apart from Figure 1 In addition, there is also Figure 4As shown in the enlarged view, the console 14, as a structure related to the photography assistance function, includes a photography instruction receiving unit 14A and a photography technique determination unit 14B. The photography instruction receiving unit 14A displays the photography instruction 31 received from the RIS on the display 14C. The photography technique determination unit 14B performs image analysis on the camera image CP using a learned model LM to derive the photography technique for the located subject H. The photography technique determination unit 14B outputs a determination result 32 containing the photography technique derived from the learned model LM. The console 14 displays a location confirmation screen 36 on the display 14C, containing the photography instruction 31 received by the photography instruction receiving unit 14A and the determination result 32 of the photography technique determination unit 14B.

[0070] On the positioning confirmation screen 36, for example, the photography instruction 31 and the judgment result 32 are displayed side by side for comparison. Regarding the photography instruction 31, only the part related to the photography technique is extracted, such as displaying a message 31A like "The photography instruction is 'chest / back'". Furthermore, the judgment result 32 includes the photography technique determined by the photography technique judgment unit 14B (in... Figure 4 In the example, it is "chest / back" and camera image CP. In judgment result 32, as an example, the photographic technique is displayed in the form of a message such as "The photographic technique determined according to the camera image is 'chest / back'."

[0071] By confirming the location of the image 36, technician RG can visually verify the photographic techniques contained in the photographic instructions 31 and the judgment results 32 to confirm whether the subject H's location status conforms to the photographic instructions 31.

[0072] like Figure 5 As shown, a learned model LM may be used, for example, as a convolutional neural network (CNN) suitable for image analysis. The learned model LM may have an encoder 37 and a classifier 38.

[0073] Encoder 37, constructed using a CNN, extracts various feature maps representing the features of the camera image CP by performing convolution and pooling processes on the CP. As is well known, in CNNs, convolution is, for example, spatial filtering using multiple filters of similar 3×3 size. In the case of a 3×3 filter, filter coefficients are assigned to 9 cells. During convolution, the center cell of such a filter is aligned with the pixel of interest in the camera image CP, and the output is the sum of the product of the pixel of interest and the pixel values ​​of the 8 surrounding pixels, totaling 9 pixels. The output product sum represents the feature quantity of the region of interest to which the filter was applied. Furthermore, for example, by applying the filter to all pixels of the camera image CP while staggering the pixels of interest by one pixel, a feature map with a feature quantity equivalent to the number of pixels in the camera image CP is output. By applying multiple filters with different filter coefficients, multiple feature maps are output. The number of feature maps corresponding to the number of filters is also called the number of channels.

[0074] This convolution is repeated while reducing the size of the camera image CP. The process of reducing the size of the camera image CP is called pooling. Pooling is performed by interleaving adjacent pixels and averaging. Through pooling, the size of the camera image CP is reduced to 1 / 2, 1 / 4, and 1 / 8 in stages. For each size of the camera image CP, a feature map with multiple channels is output. When the size of the camera image CP is large, the fine morphological features of the subject are depicted in the camera image CP. Conversely, when the size of the camera image CP is small (i.e., low resolution), the fine morphological features are discarded from the camera image CP, and only the general morphological features of the subject are depicted. Therefore, the feature map when the camera image CP is large represents the microscopic features of the subject depicted in the camera image CP, and the feature map when the size is small represents the macroscopic features of the subject depicted in the camera image CP. Encoder 37 performs this convolution and pooling process on the camera image CP, thereby extracting feature maps with multiple channels representing the macroscopic and microscopic features of the camera image CP.

[0075] As an example, the learned model LM is a classification model that derives the most probable photographic technique from multiple photographic techniques as represented by the camera image CP. Therefore, the classification unit 38 is configured to derive a photographic technique based on the features of the camera image CP extracted by the encoder 37. The classification unit 38, for example, has multiple perceptrons, each having one output node for multiple input nodes. Furthermore, the perceptrons are assigned weights representing the importance of the multiple input nodes. In each perceptron, the output node outputs the sum of the values ​​obtained by multiplying the input values ​​to the multiple input nodes by their respective weights, i.e., the product sum, as the output value. As an example, such a perceptron is formulated using an activation function such as a sigmoid function.

[0076] In the output section, by combining the outputs and inputs of multiple perceptrons, a multi-layered neural network with multiple intermediate layers between the input and output layers is formed. One method for combining multiple perceptrons between layers is to combine all output nodes of the previous layer with a single input node of the next layer.

[0077] In the input layer of classification unit 38, all feature quantities contained in the feature map of the camera image CP are input. In the perceptrons constituting each layer of classification unit 38, the feature quantities are input as input values ​​to the input nodes. Furthermore, the sum of the values ​​obtained by multiplying the feature quantities by their weights for each input node is output as the output value from the output nodes, and this output value is passed to the input nodes of the perceptrons in the next layer. In the output layer of the final layer in classification unit 38, based on the output values ​​of multiple perceptrons, the probabilities of various photographic techniques are output using a softmax function, etc. Furthermore, the photographic technique with the highest probability is derived based on these probabilities. Figure 5 The image shows an example of a photographic technique such as deriving a “chest / back” view from a camera image CP.

[0078] This example is just one illustration. As long as the learned model LM can derive the classification model of photographic techniques from the camera image CP, it can also be in other ways.

[0079] Next, use Figure 6 and Figure 7 The learning device 40 will be described below. Figure 6 The hardware structure of the learning device 40 is shown in the figure. The learning device 40 is composed of a computer such as a personal computer or a workstation. The learning device 40 includes a display 41, an input device 42, a CPU 43, a memory 44, a storage device 46, and a communication unit 47. They are interconnected via a data bus 48.

[0080] Display 41 is a display unit that displays various operation screens with GUI (Graphical User Interface) based operation functions. Input device 42 is an input operation unit including a touch panel or keyboard, etc.

[0081] Storage device 46, for example, is composed of HDD (Hard Disk Drive) and SSD (Solid State Drive), and is either built into learning device 40 or externally connected to learning device 40. External connection is via cable or network. Storage device 46 stores control programs such as the operating system, various application programs, and various data associated with these programs. One of the various application programs is the operating program AP that enables the computer to function as learning device 40. Various data include the machine learning model LM0, which is the object of learning processing; the completed learning model LM; teacher data TD used for learning; three-dimensional computer graphics data (hereinafter referred to as 3DCG data) 56; and parameter specification information 61.

[0082] Memory 44 is the working memory used by CPU 43 for processing. CPU 43 loads the program stored in memory device 46 into memory 44 and executes the program, thereby centrally controlling the various parts of learning device 40.

[0083] The communication unit 47 communicates with the network N via the console 14. For example, the communication unit 47 is used for sending and receiving the learned model LM from the learning device 40 to the console 14, and sending the learned model LM from the console 14 to the learning device 40 in order to add learning to the learned model LM of the console 14.

[0084] In addition, Figure 6 The hardware structure of the computer used to implement the learning device 40 is shown in the figure, but the hardware structure of the console 14 described above is also the same. That is, the structure of the camera instruction receiving unit 14A and the camera technique determination unit 14B of the console 14 is implemented by the cooperation of a processor such as CPU 43, a memory such as memory 44 built into or connected to CPU 43, and a program executed by CPU 43.

[0085] like Figure 7 As shown, the learning device 40 includes a teacher data generation unit 51 and a learning unit 52. These processing units are implemented by a CPU 43. The teacher data generation unit 51 generates multiple pseudo-images SP of the human body based on a human body model 56A composed of 3DCG data 56 for each photographic technique that is a combination of photographic location and photographic direction. The pseudo-images SP are two-dimensional images of the subject H simulating the state of positioning relative to a radiographic system 10, which is an example of a medical imaging device.

[0086] More specifically, the teacher data generation unit 51 generates a pseudo-image SP based on three-dimensional computer graphics data (hereinafter referred to as 3DCG data) 56 and parameter specification information 61. The parameter specification information 61 includes modeling parameter (hereinafter referred to as M-parameter) specification information 61M specified when modeling the three-dimensional human body model 56A simulating the human body, and rendering parameter (hereinafter referred to as R-parameter) specification information 61R specified when rendering the human body model 56A from a set viewpoint. Details regarding the method of generating the pseudo-image SP through modeling and rendering the human body model 56A will be described later.

[0087] Furthermore, the teacher data generation unit 51 generates teacher data TD, which consists of the generated pseudo-image SP and correct data AD, which is a combination of the photographic location and the photographic direction. Figure 7 The teacher data TD shown is an example of a photographic technique where the combination of the photographic part and the photographic direction is "chest / back". That is, in Figure 7 In the teacher data TD shown, the pseudo-image SP represents the back of the human chest, which is an image representing "chest / back" as a photographic technique. Correspondingly, the photographic technique of the correct data AD is also "chest / back".

[0088] Furthermore, the learning unit 52 includes a main processing unit 52A, an evaluation unit 52B, and an update unit 52C. The main processing unit 52A reads the machine learning model LMO, which is the object of the learning process, from the storage device 46, and inputs the pseudo-image SP into the read machine learning model LMO. It then causes the machine learning model LMO to perform processing to derive the photographic technique shown in the input pseudo-image SP. The basic structure of the machine learning model LMO is similar to that in... Figure 5 The learned model LM is the same as described above, but the filter coefficients in encoder 37 and the weights of the perceptron in classifier 38 are different from those in the learned model LM before the machine learning model LM0 is learned.

[0089] The main processing unit 52A outputs the photographic techniques derived from the machine learning model LM0 as output data OD to the evaluation unit 52B. The evaluation unit 52B compares the output data OD with the correct data AD contained in the teacher data TD, and evaluates the difference between the two as a loss using a loss function. Similar to the learned model LM, in this example's machine learning model LMO, multiple photographic techniques shown in the pseudo-image SP are output with probabilities. Therefore, the loss is self-evident when the photographic technique in the output data OD is incorrect; even when it is correct, if the output value, as its probability output, is lower than the target value, the difference between the target value and the output value is also evaluated as a loss. The evaluation unit 52B outputs the evaluated loss as an evaluation result to the update unit 52C.

[0090] The update unit 52C updates the filter coefficients in the encoder 37 and the weights of the perceptron in the classification unit 38, etc., to reduce the loss included in the evaluation result. This series of processes from input to update of the teacher data TD is repeated until the end timer arrives. The end timer may be, for example, when the learning of a predetermined number of teacher data TDs has been completed, or when the loss is lower than the target value.

[0091] The following uses Figures 8-15 This section provides a detailed explanation of the method for generating pseudo-images SP executed by the teacher data generation unit 51. First, using... Figure 8 The structure of 3DCG data 56 will be explained below. In this example, 3DCG data 56 is used to construct the human body model 56A. In addition to the structural data of the human body model 56A, 3DCG data 56 also includes modeling parameters (called M-parameters) 56B. The human body model 56A is three-dimensional (X, Y, and Z directions) data simulating the human body. M-parameters 56B are parameters that can be changed when modeling the human body model 56A; they are used to change at least one of the posture and appearance of the human body model 56A. M-parameters 56B are parameters specified by M-parameter specification information 61M.

[0092] Specifically, as an example, M-parameter 56B includes various items such as physical information, gender, posture information, skin color, hair color, hairstyle, and clothing. Physical information includes height, weight, sitting height, crotch length, head circumference, neck circumference, shoulder width, chest circumference, waist circumference, hand length, wrist circumference, hand width, foot length, foot width, thigh circumference, and calf circumference. Posture information includes items indicating basic postures such as standing, lying, and sitting, as well as items indicating whether the hands and feet are bent. Regarding the bending of the hands and feet, it allows specifying, for example, the bending location, direction, and angle of the right hand, left hand, right knee, left knee, and both knees. Furthermore, M-parameter 56B can also specify internal or external rotation of the knee joint. In this case, it can also specify the direction and angle of rotation for the inner or outer sides.

[0093] like Figure 4As shown, in the human body model 56A, bendable joints 57 are pre-defined, allowing the parts with joints 57 to bend. By specifying such M-parameters 56B, the posture and appearance of the human body can be changed when modeling the human body model 56A. By finely setting the joints 57, even for the same photographic area, pseudo-images SP with slightly different postures can be generated. Furthermore, the number and position of joints 57 in the human body model 56A can be arbitrarily changed. Therefore, by setting a larger number of joints 57 in the human body model 56A, complex poses can be achieved. On the other hand, by setting a smaller number of joints 57, the number of M-parameters 56B can be reduced, thus speeding up the modeling process.

[0094] like Figure 9 As shown, the teacher data generation unit 51 includes a modeling unit 51A and a rendering unit 51B. The modeling unit 51A models a human body model 56A with the pose and appearance of the human body specified by the M-parameter specification information 61M, based on 3DCG data 56 and M-parameter specification information 61M. Figure 9 In the M-parameter specification 61M, multiple M-parameter sets (MPS) are included. For each M-parameter set (MPS), a human body model 56A with different poses and appearances is generated.

[0095] like Figure 10 As shown, in the M-parameter group MPS1, the height is specified as 170cm and the weight as 70kg as body information, and the gender as male. The posture information is specified as standing with arms and legs not bent. Based on this M-parameter group MPS1, the modeling unit 51A generates 3DCG data 56, which has a human body model 56A_MPS1, having the posture and appearance specified by the M-parameter group MPS1.

[0096] And, as Figure 8As shown in M-parameter 56B, by specifying skin tone, hairstyle, and hair color in the M-parameter group MPS, the skin tone, hairstyle, and hair color of the human body model 56A can be changed. Furthermore, by specifying clothing in the M-parameter group MPS, the clothing of the human body model 56A can also be changed. Thus, M-parameter 56B contains physique information and appearance-related information such as skin tone, as well as pose information. Therefore, by changing physique information or appearance-related information such as skin tone without changing pose information, it is possible to generate multiple human body models 56A with the same pose but different appearances, or to generate human body models 56A with different poses without changing appearance-related information. Furthermore, when the camera focus is on the chest, hairstyle and hair color are usually not reflected in the camera image CP of the chest, but when the camera focus is on the head, they are reflected in the camera image CP of the head. Therefore, it is effective to change hairstyle and hair color when generating a pseudo-image SP of the head.

[0097] Rendering unit 51B can use the human body model 56A (modeled by modeling unit 51A) to render the human body model 56A (in the modeling unit 51A). Figure 9 and Figure 10 In this example, a two-dimensional pseudo-image SP is generated by rendering the human model 56A (MPS1) from a set viewpoint. According to this example, the rendering is performed virtually using 3DCG data 56 as follows: a camera image CP of the subject H is acquired using an optical camera 15. That is, a virtual camera 15V corresponding to the optical camera 15 is set within the three-dimensional space where the human model 56A is configured, and a virtual image of the simulated subject H's human model 56A is captured, thereby obtaining the pseudo-image SP.

[0098] exist Figure 9 In the R-parameter specification information 61R, multiple R-parameter groups RPS are included. A pseudo-image SP is generated for each R-parameter group RPS. The R-parameter group RPS specifies information such as the viewpoint position when the rendering unit 51B renders the human body model 56A.

[0099] As an example, Figure 9 and Figure 10 The R-parameter group RPS1 shown contains photographic technique information and virtual camera information. The photographic technique information includes the photographic body part and the photographic direction. With "chest / back" specified as the photographic technique information, the position of the virtual camera 15V is roughly set at a position that allows it to photograph the chest of the human body model 56A from the back.

[0100] The virtual camera information provides more detailed information, such as the position of the virtual camera 15V, and sets the shooting distance, focal length, height, and viewpoint direction. The shooting distance is the distance from the virtual camera 15V to the human model 56A. Since the virtual camera 15V is a virtual camera installed on the optical camera 15 of the radiation source 11, the shooting distance is set according to the SID (Source to image receptor distance), which is the distance between the radiation source 11 and the electronic cassette 13. The SID is a value that is appropriately changed according to the shooting method, etc. Figure 9 In this example, the shooting distance is set to 90cm. Focal length is information used to define the field of view of the virtual camera 15V. The shooting range SR of the virtual camera 15V is defined by limiting the shooting distance and focal length (see reference). Figure 11 and Figure 12 (etc.). The imaging range SR is set to include the size of the irradiation field of the radiation source 11. In this example, the focal length is set to 40mm. The height is set to the Z-direction position of the virtual camera 15V, depending on the part being photographed. In this example, since the part being photographed is the chest, it is set to the height of the chest of the standing mannequin 56A. The viewpoint direction is set according to the posture of the mannequin 56A and the direction of the photographic technique. Figure 10 In the example shown, the human model 56A is standing, and the camera direction is from the back, so it is set to the -Y direction. Thus, the rendering unit 51B can change the viewpoint for generating the pseudo-image SP according to the R-parameter group RPS.

[0101] The rendering unit 51B renders the human model 56A based on the R-parameter group RPS. This generates a pseudo-image SP. Figure 10 In the example, the rendering unit 51B renders the human body model 56A_MPS1 modeled by the modeling unit 51A according to the specification of the M-parameter group RPS1, thereby generating the pseudo image SP1-1.

[0102] Furthermore, in Figure 10 In the diagram, within the three-dimensional space containing the human body model 56A, a virtual camera obscura 13V, serving as a virtual electronic camera obscura 13, is also present. This is illustrated to clearly define the location of the photographic parts. However, during modeling, the virtual camera obscura 13V can also be modeled in addition to the human body model 56A, and rendered including the modeled virtual camera obscura 13V.

[0103] Figure 11 This represents an example of generating a pseudo-image SP with the photographic technique of "chest / front". Figure 11 The M-parameter set MPS1 and Figure 10 The same example applies to human model 56A_MPS1 and... Figure 10The examples are the same. Figure 11 R-parameter set RPS2 and Figure 10 The R-parameter group RPS1 differs, with the shooting direction changing from "back" to "front". Along with this change in shooting direction, the viewpoint direction of the virtual camera information changes to the "Y direction". Figure 11 In the R-parameter set RPS2, other aspects are similar to Figure 10 The R-parameter group RPS1 is the same.

[0104] Based on the R-parameter set RPS2, the position of the virtual camera 15V is set, and rendering is performed. Thus, by photographing the "chest" of the human model 56A_MPS1 from a "frontal" perspective, pseudo-images SP1-2 corresponding to the "chest / frontal" photographic technique are generated.

[0105] Figure 12 This represents an example of a pseudo-image SP generated using the photographic technique of "abdomen / front". Figure 12 The M-parameter set MPS1 and Figure 10 and Figure 11 The same example applies to human model 56A_MPS1 and... Figure 10 and Figure 11 The examples are the same. Figure 12 R-parameter group RPS3 and Figure 10 Unlike the R-parameter group RPS1, the imaging site has changed from "chest" to "abdomen". The imaging direction is "frontal", which is different from... Figure 10 The "back" is different, but similar to Figure 11 The examples are the same. In Figure 12 In the R-parameter group RPS3, other aspects are similar to Figure 11 The R-parameter group RPS2 is the same.

[0106] Based on the R-parameter set RPS3, the position of the virtual camera 15V is set, and the human body model 56A_MPS1 is rendered. This generates pseudo-images SP1-3 corresponding to the "chest / frontal" photographic technique of the human body model 56A_MPS1.

[0107] Figure 13 This represents an example of a pseudo-image SP generated using the photographic technique of "double knees / front view". Figure 13 The M-parameter set MPS2 and Figures 10-13 Unlike the previous example, this one specifies "seated" and "knees bent" as pose information. Therefore, human model 56A_MPS2 is modeled as a seated posture with knees bent. Regarding... Figure 13In the R-parameter group RPS4, the shooting location is specified as "both knees" and the shooting direction is specified as "frontal" in the shooting technique information. In the virtual camera information, the height of the knees of the human model 56A_MPS2, which is capable of shooting from the front in a seated posture, is set to 150cm, and the shooting direction is set to downward in the Z direction, i.e., "-Z direction".

[0108] Based on the R-parameter set RPS4, the position of the virtual camera 15V is set, and the human body model 56A_MPS2 is rendered. This generates a pseudo-image SP2-4 corresponding to the "double knees / frontal view" photographic technique of the human body model 56A_MPS2.

[0109] Figure 14 and Figure 11 Similarly, an example of a pseudo-image SP representing a photographic technique called "chest / front view" is used. In Figure 14 In the example, the R-parameter set RPS2 and Figure 11 The examples are the same. The difference lies in... Figure 14 The physical information of the M-parameter group MPS4, and Figure 11 Compared to the M-parameter group MPS1, except for weight, the values ​​for chest and waist circumference (not shown) increased. That is, according to... Figure 14 The human body model 56A_MPS4 model with M-parameter set MPS4 is compared to Figure 11 The human body model 56A_MPS1 is fat. Other aspects are similar to... Figure 11 The same applies to the example. Thus, for the obese human model 56A_MPS4, a pseudo-image SP4-2 corresponding to the "chest / front" photographic technique is generated.

[0110] Figure 15 and Figure 13 Similarly, an example of a pseudo-image SP representing a photographic technique of "both knees / front view" is given. (And...) Figure 13 The difference in the example is that, in the M-parameter group MPS5, skin color is specified as brown. Other aspects are the same as... Figure 13 The same applies to the example. Thus, for the human model 56A_MPS5 with brown skin, a pseudo-image SP5-4 of "both knees / front view" is generated.

[0111] The teacher data generation unit 51 generates multiple pseudo-images SP for each combination of photographic part and photographic direction based on the parameter specification information 61. For example, the teacher data generation unit 51 uses different parameter groups (M-parameter group MPS and R-parameter group RPS) to generate multiple pseudo-images SP for each photographic technique such as "chest / front" and "chest / back".

[0112] Next, use Figures 16-18The flowchart shown illustrates the function of the learning device 40 with the above-described structure. Figure 16 As shown in the main flowchart, the learning device 40 generates teacher data TD based on 3DCG data (step S100), and uses the generated teacher data TD to enable the machine learning model LMO to learn (step S200).

[0113] Figure 17 This describes the details of step S100, which generates teacher data TD. For example... Figure 17 As shown, in step S100, firstly, the teacher data generation unit 51 of the learning device 40 acquires 3DCG data 56 (step S101). Next, the teacher data generation unit 51 acquires parameter specification information 61. In this example, in steps S101 and S102, the teacher data generation unit 51 reads the 3DCG data 56 and parameter specification information 61 stored in the storage device 46 from the storage device 46.

[0114] For example, parameter specification information 61 records multiple M-parameter groups MPS and multiple R-parameter groups RPS. For example, in... Figures 10-15 As illustrated, the teacher data generation unit 51 models and renders each combination of an M-parameter group MPS and an R-parameter group RPS, generating a pseudo-image SP (step S103). The teacher data generation unit 51 generates teacher data TD by combining the generated pseudo-image SP with the correct data AD (step S104). The generated teacher data TD is stored in the storage device 46.

[0115] In step S105, if the teacher data generation unit 51 finds any uninputted parameter groups in the parameter groups (a combination of M-parameter group MPS and R-parameter group RPS) included in the acquired parameter specification information 61 (if "Yes" is true in step S105), it executes the processing steps S103 and S104. The teacher data generation unit 51 generates multiple pseudo-images SP using different parameter groups (M-parameter group MPS and R-parameter group RPS) for each photographic technique such as "chest / front" and "chest / back". If no uninputted parameter groups are found (if "No" is true in step S105), the teacher data generation process ends.

[0116] Figure 18 This describes the detailed steps S200 of using the generated teacher data TD to enable the machine learning model LMO to learn. (For example...) Figure 18As shown, in step S200, firstly, in the learning unit 52 of the learning device 40, the main processing unit 52A retrieves the machine learning model LMO from the storage device 46 (step S201). Then, the main processing unit 52A retrieves the teacher data TD generated by the teacher data generation unit 51 from the storage device 46, and inputs the pseudo-images SP contained in the teacher data TD one by one into the machine learning model LMO (step S202). Next, the main processing unit 52A causes the machine learning model LMO to output output data OD containing photographic techniques (step S203). Then, the evaluation unit 52B compares the correct data AD with the output data OD and evaluates the output data OD (step S204). The update unit 52C updates the filter coefficients and perceptron weights of the machine learning model LMO based on the evaluation result of the output data OD (step S205). During the period until the end timer arrives ("Yes" in step S206), the learning unit 52 repeats the series of processes up to steps S202 to S205. When the learning of the predetermined number of teacher data TDs is completed, and the end timer (no in step S206) is reached when the loss is lower than the target value, the learning unit 52 ends the learning and outputs the learned model LM (step S207).

[0117] As described above, the learning device 40 of this embodiment generates multiple pseudo-images SP of the human body generated based on the human body model 56A composed of 3DCG data for each combination of imaging location and imaging direction. These pseudo-images SP simulate the state of the subject H relative to the radiography system 10, which is an example of a medical imaging device. Furthermore, the machine learning model LMO learns using multiple teacher data TDs composed of the generated pseudo-images SP and the combined correct data AD. Pseudo-images SP constituting teacher data TDs are generated based on the human body model 56A composed of 3DCG data. Therefore, compared to the case where the camera image CP representing the imaging technique as a combination of imaging location and imaging direction is acquired by the optical camera 15, it is easier to increase the number of teacher data TDs representing the imaging technique. Therefore, compared to the case where only the camera image CP is used as teacher data TD, it is possible to efficiently learn the learning model LM that derives the imaging location and imaging direction of the subject H reflected in the camera image CP.

[0118] Furthermore, the 3DCG data in the above example includes M-parameters 56B for changing at least one of the pose and appearance of the human model 56A. Therefore, for example, by simply changing M-parameters 56B according to parameter specification information 61, it is possible to generate multiple pseudo-images SP that differ in at least one of the pose and appearance of the human model 56A. Thus, it is easy to generate diverse pseudo-images SP that differ in at least one of the pose or appearance of the human model 56A. The actual pose or appearance of the subject H also varies from person to person. Collecting diverse camera images CP of subjects H or camera images CP of subjects H in various poses is very labor-intensive and time-consuming. Given this reality, the method of easily changing the pose or appearance of the human model 56A by changing M-parameters 56B is very effective.

[0119] Furthermore, M-parameter 56B includes at least one of the following: physical information, gender, posture information, skin color, hair color, hairstyle, and clothing of the human body model 56A. In the technology of this invention, by precisely specifying appearance-related information and posture information, it is possible to easily collect human body images of appearances or postures that are difficult to collect in actual camera images CP. Therefore, for example, even when the photographic location and photographic direction are the same, pseudo-images SP of the human body model 56A simulating a variety of appearances of the subject H can be used as teacher data TD for the machine learning model LMO to learn. Thus, compared to cases where physical information cannot be specified, the accuracy of deriving the photographic location and photographic direction from the camera image CP can be improved after the model LM has been learned.

[0120] In the examples above, the physical information included height, weight, sitting height, crotch length, head circumference, head width, neck circumference, shoulder width, chest circumference, waist circumference, hand length and wrist circumference, hand width, foot length and foot width, thigh circumference, and calf circumference. However, including at least one of these is sufficient. Of course, the more items in the physical information, the more diversity of the anatomical model 56A can be ensured, and therefore this is preferred. Furthermore, the example used numerical values ​​to define height and weight, etc., but in addition to numerical values, evaluation information such as slender, standard, and robust can also be included.

[0121] Furthermore, in this embodiment, the rendering unit 51B can generate pseudo-images SP by rendering the human model 56A from a set viewpoint, and the viewpoint can be changed by the R-parameter group RPS that limits the rendering parameters. Therefore, compared with the case without such rendering parameters, viewpoint changing becomes easier, and it is easier to generate a variety of pseudo-images SP.

[0122] The viewpoint information used to set the viewpoint can include the focal length of the virtual camera 15V, which is virtually set at the viewpoint, and the shooting distance, which is the distance from the virtual camera 15V to the human model 56A. Since the shooting distance and focal length can be changed, the shooting range SR of the virtual camera 15V can be easily changed.

[0123] In the above embodiments, an example of applying the technology of the present invention to a radiographic system 10, which is an example of a medical imaging apparatus and a radiographic imaging apparatus, has been described. In radiographic imaging, there are many types of imaging instructions that specify combinations of imaging sites and imaging directions, such as "chest / front," "abdomen / front," and "both knees / front." Thus, this invention is particularly effective in the case of medical imaging apparatuses that use a wide variety of imaging instructions.

[0124] In the examples above, it is also possible to use Figure 9 The parameter specification information 61 shown executes batch processing to continuously generate multiple pseudo-images SP. The parameter specification information 91 includes M-parameter specification information 61M and R-parameter specification information 61R. Furthermore, as... Figure 9 As shown, multiple M-parameter groups (MPS) can be specified in the M-parameter specification information 61M, and multiple R-parameter groups (RPS) can be specified in the R-parameter specification information 61R. By specifying groups of M-parameter groups (MPS) and R-parameter groups (RPS), one pseudo-image SP is generated. Therefore, by using parameter specification information 91, batch processing for continuously generating multiple pseudo-image SPs can be performed. Such batch processing is possible when generating pseudo-image SPs using 3DCG data 56. Batch processing allows for the efficient generation of multiple pseudo-image SPs.

[0125] Furthermore, while the above example illustrates still image photography using optical camera 15, moving image photography is also possible. In the case of moving image photography using optical camera 15, the technician RG can input photography instructions before the subject H begins positioning, causing optical camera 15 to begin moving image photography, and the moving images captured by optical camera 15 can be sent to console 14 in real time. In this case, for example, the frame images constituting the moving images are used as input data for the learned model LM.

[0126] "Variation Example 1"

[0127] The above example illustrates the use of only the pseudo-image SP as teacher data TD, but it can also be illustrated as follows: Figure 19As shown, in addition to the pseudo-image SP, the camera image CP is used concurrently to enable the machine learning model LMO to learn. Here, in the case where teacher data TD, composed of the pseudo-image SP and the correct data AD, is used as the first teacher data, the teacher data composed of the camera image CP and the correct data AD is used as the second teacher data TD2. The learning unit 52 uses both the first teacher data TD and the second teacher data TD2 to enable the machine learning model LMO to learn. By using the camera image CP, which is used when the learned model LM is applied, the accuracy of deriving the shooting location and shooting direction is improved compared to the case where only the pseudo-image SP is used.

[0128] "Variation Example 2"

[0129] In the above examples, such as Figure 4 As shown, during the application phase after learning the model LM, the photography technique determination unit 14B outputs the photography techniques ("chest / back", etc.) contained in the photography instruction 31 and the photography techniques ("chest / back", etc.) exported from the learned model LM to the display 14C as is. However, as Figure 20 As shown, a verification unit 14D can also be provided in the photography technique determination unit 14B, which verifies the photography techniques contained in the photography instruction 31 and the photography techniques derived from the learned model LM. In this case, the verification unit 14D can also output a verification result showing whether the two photography techniques are consistent, and display the verification result on the positioning confirmation screen 36. Figure 20 In the example, the verification result is displayed as a message such as "The determined photography technique is consistent with the photography instructions." Furthermore, if the verification result is that the two photography techniques are inconsistent, a message such as "The determined photography technique is inconsistent with the photography instructions" (not shown) is displayed on the positioning confirmation screen 36 as the verification result. This message serves as an alarm to alert technician RG to the positioning error. Moreover, as an alarm, voice, warning sound, and warning light can be used.

[0130] In the above embodiment, the learning device 40 has a teacher data generation unit 51, and the learning device 40 also functions as a teacher data generation device, but the teacher data generation unit 51 can also be separated from the learning device 40 and used as an independent device.

[0131] "Second Implementation Method"

[0132] Figure 21 and Figure 22The second embodiment shown considers using a camera image CP as input to derive a learned physique output model (LMT) representing physique information of the subject H as reflected in the camera image CP. The second embodiment uses a pseudo-image SP in the learning of the physique output machine learning model LMTO, which forms the basis of the learned physique output model LMT, in order to generate such a model. For example... Figure 21 As shown, the body output uses a learned LMT model, for example, which includes an encoder 37 and a regressor 68. Encoder 37 and... Figure 5 The learned model LM for the derived photographic technique is the same. The learned model LMT for body output is a regression model that infers, for example, the value of body thickness from the feature map of the camera image CP as the body information of the subject H reflected in the camera image CP. As the regression part 68, for example, a linear regression model, support vector machine, etc. are used.

[0133] like Figure 22 As shown, the learning unit 52 uses multiple physique output teacher data TDTs using pseudo-images SP to enable the physique output machine learning model LMTO to learn. The physique output teacher data TDTs consist of pseudo-images SPs and correct data ADTs representing physique information of the human body model 56A.

[0134] As in Figure 3 As shown, the body thickness of the subject H is used as the basis for determining the irradiation conditions. Therefore, if the body thickness can be derived from the camera image CP, the convenience in radiography is improved.

[0135] "Third Implementation Method"

[0136] Figure 23 The third embodiment shown is an example of applying the technology of the present invention to an ultrasound imaging apparatus as a medical imaging device. The ultrasound imaging apparatus has a probe 71, which includes a transmitting part and a receiving part for ultrasonic waves. Ultrasonic imaging is performed by bringing the probe 71 into contact with the imaging site of the subject H. In ultrasound imaging, the imaging site is also specified by imaging instructions, and sometimes the imaging direction (the direction of the probe 71) is also determined according to the imaging purpose. Sometimes, if the imaging site and imaging direction are inappropriate, a suitable ultrasound image cannot be obtained.

[0137] In particular, unlike radiography, ultrasound imaging can be performed by nurses and caregivers in addition to physicians, and it is also envisioned that in the future, patients H who are unfamiliar with radiography may operate the probe 71 themselves to perform ultrasound imaging. In this case, the use of camera images CP for imaging guidance has also been studied. In this case, it is believed that the use of camera images CP as input to derive a learned model LM of the imaging site and imaging direction will also increase. That is, in ultrasound imaging, camera images CP are images captured by optical camera 15 of the state of the probe 71 against the imaging site of the patient H, that is, the state of the patient H relative to the positioning of the probe 71. As for the accuracy of deriving such a learned model LM, in order to ensure that it can withstand the accuracy used for imaging guidance, it is important to use a variety of teacher data TD to learn the learned model LM. By using the technology of the present invention, the learning of teacher data TD based on a variety of pseudo-images SP can be performed efficiently.

[0138] like Figure 23 As shown, when the technology of the present invention is applied to ultrasound imaging, the teacher data generation unit 51 also generates a pseudo-image SP by rendering the state in which the human body model 56A and the virtual probe 71 are positioned from a viewpoint set by the virtual camera 15V. Figure 7 As shown, the machine learning model LM is trained using the teacher data TD generated from the pseudo-image SP.

[0139] In the above embodiments, for example, the hardware structure of the processing unit that performs various processes such as the teacher data generation unit 51 and the learning unit 52 is as shown below, using various processors.

[0140] Various processors include CPUs, programmable logic devices (PLDs), and application-specific circuits (ASICs). As is well known, a CPU is a general-purpose processor that executes software (programs) and functions as a processing unit for various tasks. A PLD is a processor like a field-programmable gate array (FPGA) whose circuit structure can be modified after manufacturing. An application-specific circuit (ASIC) is a processor with a circuit structure specifically designed to perform specific processes.

[0141] A processing unit can consist of one of these various processors, or it can consist of a combination of two or more processors of the same or different types (e.g., multiple FPGAs or a combination of a CPU and an FPGA). Furthermore, multiple processing units can also be composed of a single processor.

[0142] As examples of a single processor comprising multiple processing units, firstly, there exists a configuration where a single processor is composed of one or more CPUs combined with software, functioning as multiple processing units. Secondly, there exists a configuration, such as a System-on-Chip (SoC), where a single IC chip implements the overall functionality of a system including multiple processing units. Thus, various processing units are constructed using one or more of these different processors as their hardware structure.

[0143] Moreover, more specifically, the hardware structure of these various processors is a circuit composed of circuit elements such as semiconductor components.

[0144] The technology of this invention is not limited to the embodiments described above. Various structures can be adopted as long as they do not depart from the spirit of this invention. Moreover, in addition to programs, the technology of this invention also relates to computer-readable storage media that do not temporarily store programs.

[0145] The descriptions and illustrations shown above are detailed explanations of the parts involved in the technology of this invention, and are merely examples of the technology of this invention. For example, the descriptions of the structure, function, effect, and effect described above are examples of the structure, function, effect, and effect of the parts involved in the technology of this invention. Therefore, it is natural that, without departing from the spirit of the technology of this invention, unnecessary parts may be deleted, new elements may be added, or substitutions may be made to the descriptions and illustrations shown above. Furthermore, in order to avoid complexity and facilitate the understanding of the parts involved in the technology of this invention, descriptions related to technical common sense that do not require special explanation have been omitted in the descriptions and illustrations shown above, based on the technology that enables the implementation of this invention.

Claims

1. A learning device that enables a machine learning model to learn, the machine learning model taking as input camera images obtained by photographing a subject in a position relative to a medical imaging device, the subject's imaging location and imaging direction as reflected in the camera images. The learning device has a processor and a memory connected to or built into the processor. The processor performs the following processing: For each combination of the imaging location and imaging direction, multiple pseudo-images of the human body are generated based on a human body model composed of three-dimensional computer graphics data; that is, pseudo-images of the subject simulating their position relative to the medical imaging device. The machine learning model learns using multiple sets of teacher data consisting of the generated pseudo-images and the combined correct data.

2. The learning device according to claim 1, wherein, The three-dimensional computer graphics data includes modeling parameters for changing at least one of the posture and appearance of the human body model.

3. The learning device according to claim 2, wherein, The modeling parameters include at least one of the following: physical information, gender, posture information, skin color, hair color, hairstyle, and clothing, representing the physique of the human model.

4. The learning device according to claim 1, wherein, The processor can generate the pseudo-image by rendering the human model from a set viewpoint, which can be changed by rendering parameters.

5. The learning device according to claim 4, wherein, The viewpoint information used to set the viewpoint includes the focal length of a virtual camera virtually set at the viewpoint and the photographic distance, which is the distance from the virtual camera to the human model.

6. The learning device according to claim 1, wherein, In the case where the teacher data consisting of the pseudo-image and the correct data is used as the first teacher data, The processor performs the following processing: In addition to the first teacher data, a second teacher data consisting of camera images taken with an optical camera and the correct data is used to enable the machine learning model to learn.

7. The learning device according to claim 1, wherein, Furthermore, a plurality of physical output teacher data, consisting of the pseudo-images and correct data representing physical information of the human body model, are used to enable a physical output machine learning model that takes the camera images as input to derive physical information representing the subject's physique.

8. The learning device according to any one of claims 1 to 7, wherein, The medical imaging device includes at least one of a radiographic imaging device and an ultrasound imaging device.

9. A learning method that uses a computer to enable a machine learning model to learn, the machine learning model taking as input camera images obtained by photographing a subject in a position relative to a medical imaging device, and deriving the imaging location and imaging direction of the subject as reflected in the camera images, wherein... For each combination of the imaging location and imaging direction, multiple pseudo-images of the human body are generated based on a human body model composed of three-dimensional computer graphics data; that is, pseudo-images of the subject simulating their position relative to the medical imaging device. The machine learning model learns using multiple sets of teacher data consisting of the generated pseudo-images and the combined correct data.

10. A computer-readable recording medium storing an operating program for a learning device, the operating program of which enables a computer to function as a learning device, the learning device enabling a machine learning model to learn, the machine learning model taking as input camera images obtained by photographing a subject in a position relative to a medical imaging device, the subject's imaging location and imaging direction as reflected in the camera images, wherein... For each combination of the imaging location and imaging direction, multiple pseudo-images of the human body are generated based on a human body model composed of three-dimensional computer graphics data; that is, pseudo-images of the subject simulating their position relative to the medical imaging device. The machine learning model learns using multiple sets of teacher data consisting of the generated pseudo-images and the combined correct data.

11. A teacher data generation apparatus for generating teacher data, said teacher data being used to enable a machine learning model to learn, said machine learning model taking as input camera images obtained by photographing a subject positioned relative to a medical imaging device, and deriving the imaging location and imaging direction of the subject as reflected in the camera images, wherein... The teacher data generation device has a processor and a memory connected to or built into the processor. The processor performs the following processing: Using three-dimensional computer graphics data, the three-dimensional computer graphics data constitutes a human body model for generating pseudo-images of the human body and includes parameters for changing at least one of the posture and appearance of the human body model. By changing the parameters, multiple pseudo-images are generated for each combination of the photographed part and photographic direction, where at least one of the pose and appearance of the human model is different. Generate multiple sets of teacher data consisting of the generated multiple pseudo-images and the combined correct data.

12. A computer-readable recording medium storing a machine learning model that takes as input camera images obtained by an optical camera capturing a subject positioned relative to a medical imaging device, the machine learning model derives the imaging location and imaging direction of the subject as reflected in the camera images. The machine learning model was trained using multiple teacher data sets, which consist of correct data combining pseudo-images and the photographic location and direction, and generated for each combination. The pseudo-images are pseudo-images of the human body generated based on a human body model composed of three-dimensional computer graphics data, and are pseudo-images of the subject simulating the position relative to the medical imaging device.

13. A medical imaging device comprising a computer-readable recording medium storing the machine learning model of claim 12.

Citation Information

Patent Citations

  • Radiography system and method for operating the same

    JP2020192440A

  • Radiography system and method for operating radiography system

    CN109381208A

  • Radiologic imaging support system, radiologic imaging support method, and program

    WO2020250917A1