Machine learning system, learning method, trained model, program and recording medium
The machine learning system efficiently trains on composite medical images to enhance the recognition of regions of interest in ultrasound images by combining medical images with graphics, addressing the interference issue and improving recognition accuracy.
Patent Information
- Application Number
- JP2023503803
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-01
- Filing Date
- 2022-02-28
- Publication Date
- 2026-01-21
- Estimated Expiration
- 2042-02-28
AI Technical Summary
Existing machine learning systems for medical images, particularly ultrasound images, struggle to accurately recognize regions of interest due to the interference of user-specified graphics, and training with or without graphics leads to suboptimal performance.
A machine learning system that generates composite medical images by combining selected medical images with graphics, allowing the learning model to be trained efficiently using these composite images, including variations in graphic position and size, and using both graphic-synthesized and non-synthesized images for training.
Enables efficient and accurate training of the learning model to distinguish between graphics and regions of interest, improving the recognition of medical images with superimposed graphics.
Smart Images

Figure 0007803922000001 
Figure 0007803922000002 
Figure 0007803922000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a machine learning system, a recognizer, a learning method, and a program, and more particularly to a machine learning system, a recognizer, a learning method, and a program for medical images on which graphics are superimposed. [Background technology]
[0002] In recent years, highly accurate automatic recognition has become possible through machine learning using deep learning (Non-Patent Document 1). This technology is also applied to ultrasound diagnostic devices, which can recognize various information from ultrasound images using recognizers trained by deep learning and present it to the user.
[0003] In deep learning, supervised learning is known, in which learning is performed using training data and correct answer data that indicates the correct answer to the training data.
[0004] For example, Patent Document 1 describes a learning data generation system that generates learning data to be used in machine learning. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] A. Krizhevsky, I. Sutskever, and G. Hinton. ImageNetclassification with deep convolutional neural networks. In NIPS, 2012 [Patent documents]
[0006] [Patent Document 1] Japanese Patent Publication No. 2020-192006 Summary of the Invention [Problem to be solved by the invention]
[0007] Here, in the case of medical images (e.g., ultrasound images), it is considered that machine learning of a recognizer is performed using correct answer data containing information indicating the position and type of organs, and the recognizer that has undergone machine learning is used as a recognizer to present the position and type of organs to the user.
[0008] Medical images, especially ultrasound images, are often displayed with a user-specified graphic superimposed on the acquired medical image. Examples of user-specified graphic include a rectangle enclosing the position of an organ in the medical image, an arrow indicating a region of interest, a circle, a triangle, and a line representing the length of an organ in the medical image.
[0009] When a recognizer is trained using such medical images with superimposed graphics as training data, the recognizer may perceive the graphics as features and learn, which may prevent it from properly recognizing medical images without graphics.Furthermore, when a recognizer is trained using medical images without graphics as training data, the recognizer may not properly recognize medical images with graphics because it has not been trained on medical images with graphics.
[0010] To improve the accuracy of recognition of medical images that incorporate graphics, the recognizer must be trained using a large amount of training data to understand that the graphics and the medical image recognition target are unrelated.
[0011] However, preparing a large number of medical images containing various shapes, which are necessary for machine learning, is time-consuming and labor-intensive, and not easy.
[0012] The technology described in Patent Document 1 does not mention using the synthesized image as learning data for a recognizer that recognizes the region of interest or type of the image.
[0013] The present invention has been made in consideration of these circumstances, and its purpose is to provide a machine learning system, a learning method, and a program that can easily train a large number of synthesized medical images. [Means for solving the problem]
[0014] To achieve the above-mentioned objective, one aspect of the present invention is a machine learning system comprising an image database in which a plurality of medical images are stored, a graphic database in which graphics to be superimposed on the medical images are stored, a processor, and a learning model, wherein the processor performs a selection reception process that accepts the selection of medical images stored in the image database and the selection of graphics stored in the graphic database, a graphic synthesis process that synthesizes the selected medical images and graphics to generate a composite image, and a learning process that causes the learning model to learn using the composite image.
[0015] According to this aspect, medical images stored in the image database and figures stored in the figure database are selected, and a composite image is generated from the medical images and figures. This composite image is then used to train the learning model, making it possible to more efficiently and easily train a large number of composite medical images.
[0016] Preferably, in the learning process, the medical images accepted in the selection acceptance process are used at least twice, the first time using medical images that have not undergone graphic synthesis processing, and the second time using synthesized images that have undergone graphic synthesis processing to train the learning model.
[0017] According to this aspect, the learning model is trained using medical images that have not undergone graphic synthesis processing and composite images that have undergone graphic synthesis processing, thereby allowing the learning model to be trained more efficiently.
[0018] Preferably, the processor performs a size reception process for receiving a size of a graphic to be synthesized in the graphic synthesis process, and in the graphic synthesis process, the processor synthesizes the graphic with the medical image based on the size received in the size reception process.
[0019] Preferably, the graphic database stores at least one of maximum and minimum values of the graphic size and area information that can be combined with the medical image, in association with the graphic.
[0020] Preferably, the image database stores information relating to the region of interest in association with the medical image, and the graphic synthesis process uses the information relating to the region of interest to perform synthesis so that at least a part of the graphic is superimposed on the region of interest.
[0021] According to this aspect, the figure is synthesized so as to be superimposed on the region of interest, so that the learning model can be properly trained to distinguish between the figure and the region of interest, and the region of interest can be efficiently recognized.
[0022] Preferably, the image database stores information about the region of interest in association with the medical image, and in the graphic synthesis process, synthesis is performed using the information about the region of interest, excluding the region of interest from the graphics.
[0023] According to this aspect, the figure is synthesized excluding the area of interest, so that the learning model can properly learn to distinguish between the figure and the area of interest, and can efficiently recognize the area of interest (and type).
[0024] Preferably, in the graphic synthesis process, a plurality of synthesized images are generated by changing the position and size of the selected graphic, and in the learning process, the learning model is made to learn using the plurality of synthesized images.
[0025] According to this aspect, multiple composite images are generated by changing the position and size of the graphics, so that composite images can be generated more efficiently and easily.
[0026] Preferably, the medical image is an image obtained from an ultrasound diagnostic device.
[0027] Another aspect of the present invention is a recognizer that is trained using the machine learning system described above and recognizes regions of interest from medical images with graphics added.
[0028] Another aspect of the present invention is a learning method that uses a machine learning system that includes an image database in which a plurality of medical images are stored, a graphic database in which graphics to be superimposed on the medical images are stored, a processor, and a learning model, in which the processor performs a selection reception process in which a selection of medical images stored in the image database and a selection of graphics stored in the graphic database are received, a graphic synthesis process in which the selected medical images and graphics are synthesized to generate a synthesized image, and a learning process in which the learning model is trained using the synthesized image.
[0029] Another aspect of the present invention is a program that executes a learning method using a machine learning system that includes an image database in which a plurality of medical images are stored, a graphic database in which graphics to be superimposed on the medical images are stored, a processor, and a learning model, and causes the processor to execute a selection reception process that accepts a selection of a medical image stored in the image database and a selection of a graphic stored in the graphic database, a graphic synthesis process that synthesizes the selected medical image and graphic to generate a composite image, and a learning process that causes the learning model to learn using the composite image. [Effects of the Invention]
[0030] According to the present invention, medical images stored in an image database and figures stored in a figure database are selected, a composite image is generated from the medical image and the figure, and the learning model is trained using the composite image, thereby making it possible to train a large number of composite medical images more efficiently and easily. [Brief explanation of the drawings]
[0031] [Figure 1] FIG. 1 is a block diagram illustrating an example of the configuration of a machine learning system. [Figure 2] FIG. 2 is a block diagram showing the main functions that a processor implements in a machine learning system. [Figure 3]FIG. 3 is a diagram illustrating a specific data flow in the machine learning system. [Figure 4] FIG. 4 is a diagram showing an example of an ultrasound image. [Figure 5] FIG. 5 is a diagram showing a specific example of the storage configuration of the image database. [Figure 6] FIG. 6 is a diagram showing a specific example of the storage configuration of the graphic database. [Figure 7] FIG. 7 is a diagram showing an example of a composite image obtained by the graphic composition unit. [Figure 8] FIG. 8 is a diagram showing an example of a composite image obtained by the graphic composition unit. [Figure 9] FIG. 9 is a diagram showing an example of a composite image obtained by the graphic composition unit. [Figure 10] FIG. 10 is a diagram showing an example of a composite image obtained by the graphic composition unit. [Figure 11] FIG. 11 is a functional block diagram showing the main functions of the learning unit and the learning model. [Figure 12] FIG. 12 is a flow diagram illustrating a learning method performed using a machine learning system. [Figure 13] FIG. 13 is a diagram illustrating a specific data flow in the machine learning system 10 of this embodiment. [Figure 14] FIG. 14 is a functional block diagram showing the main functions of the learning unit and the learning model. [Figure 15] FIG. 15 is a flow diagram illustrating a learning method performed using a machine learning system. [Figure 16] FIG. 16 is a diagram conceptually illustrating an example of learning data. [Figure 17] FIG. 17 is a schematic diagram showing the overall configuration of an ultrasonic endoscope system. [Figure 18] FIG. 18 is a block diagram illustrating an embodiment of an ultrasound processor device. DETAILED DESCRIPTION OF THE INVENTION
[0032] Hereinafter, preferred embodiments of a machine learning system, a recognizer, a learning method, and a program according to the present invention will be described with reference to the accompanying drawings.
[0033] First Embodiment FIG. 1 is a block diagram showing an example of the configuration of a machine learning system according to this embodiment.
[0034] The machine learning system 10 is configured by a personal computer or a workstation. The machine learning system 10 includes a communication unit 12, an image database (referred to as an image DB in the figure) 14, a graphic database (referred to as a graphic DB in the figure) 16, an operation unit 20, a processor 22, a random access memory (RAM) 24, a read-only memory (ROM) 26, a display unit 28, and a learning model 30. Each unit is connected via a bus 32. Note that, although an example of the machine learning system 10 connected to the bus 32 has been described in this example, the example of the machine learning system 10 is not limited to this. For example, part or all of the machine learning system 10 may be connected via a network. Here, the network includes various communication networks such as a local area network (LAN), a wide area network (WAN), and the Internet. Furthermore, although the machine learning system 10 is described below in this example, the present invention is also applicable to a machine learning device.
[0035] The communication unit 12 is an interface that performs communication processing with an external device via wire or wirelessly, and exchanges information with the external device.
[0036] The ROM 26 permanently stores programs, data, etc., such as the computer's boot program and BIOS (Basic Input / Output System). The RAM 24 temporarily stores programs, data, etc. loaded from the ROM 26 or a separately connected storage device, and also provides a work area used by the processor 22 to perform various processes.
[0037] The operation unit 20 is an input interface that accepts various operation inputs to the machine learning system 10. The operation unit 20 uses a keyboard, a mouse, or the like that is connected to a computer by wire or wirelessly.
[0038] The processor 22 is configured with a CPU (Central Processing Unit). It reads various programs stored in a ROM 26 or a hard disk device (not shown) and executes various processes. The RAM 24 is used as a working area for the processor 22. The RAM 24 is also used as a storage unit that temporarily stores the read programs and various data. In the machine learning system 10, the processor 22 may be configured with a GPU (Graphics Processing Unit).
[0039] The display unit 28 is an output interface that displays information necessary for the machine learning system 10. The display unit 28 may be any of various monitors, such as a liquid crystal monitor, that can be connected to a computer.
[0040] The image database 14 is a database that stores a plurality of medical images, which are, for example, ultrasound images acquired by an ultrasound endoscope system described below.
[0041] The graphic database 16 is a database that stores a variety of graphics. The graphics are used as an aid for diagnosing (or observing) a region of interest when a doctor observes an ultrasound image, for example.
[0042] The image database 14 and the graphic database 16 are configured on a recording medium or a cloud. The image database 14 and the graphic database 16 will be described in detail later.
[0043] The learning model 30 is composed of a convolutional neural network (CNN) and performs machine learning to recognize the position and type of a region of interest from a medical image. When a medical image is input, the learning model 30 outputs at least one of the position and type of a region of interest as an estimation result. The learning model 30 installed in the machine learning system 10 is an untrained model, and the machine learning system 10 causes the learning model 30 to perform machine learning to estimate the position and type of a region of interest. Here, a region of interest is, for example, an organ, a lesion, etc., and is an area that a doctor focuses on when observing a medical image. Furthermore, the type of region of interest is, specifically, the name of an organ, the name of a lesion, the level of the lesion, etc., and refers to the classification of the region of interest.
[0044] Here, an example has been described in which the machine learning system 10 is configured using a single personal computer or workstation, but the machine learning system 10 may also be configured using multiple personal computers.
[0045] FIG. 2 is a block diagram showing the main functions that the processor 22 implements in the machine learning system 10.
[0046] The processor 22 mainly comprises a selection unit 22A, a graphic synthesis unit 22B, and a learning unit 22C.
[0047] The selection unit 22A performs a selection receiving process to receive a selection of medical images stored in the image database 14. The selection unit 22A also receives a selection of figures stored in the figure database 16. For example, a user selects medical images stored in the image database 14 and figures stored in the figure database 16 via the operation unit 20, and the selection unit 22A receives the selection. The selection unit 22A also performs a size receiving process to receive the size of the figures to be combined in the figure combining process. For example, the selection unit 22A also receives a selection of the size of the figures to be combined in conjunction with receiving the selection of the figures.
[0048] The graphic synthesis unit 22B performs a graphic synthesis process to synthesize the selected medical image and graphic to generate a composite image. The graphic synthesis unit 22B generates a composite image by superimposing a graphic on the medical image in various ways. For example, the graphic synthesis unit 22B can generate a composite image by superimposing a graphic on a region of interest of the medical image. The graphic synthesis unit 22B can also generate a composite image by superimposing a graphic so as to exclude the region of interest of the medical image. The graphic synthesis unit 22B can also superimpose a graphic on the medical image (randomly) regardless of the region of interest. The graphic synthesis unit 22B can determine the position at which the graphic is to be superimposed using, for example, a random number table, and then randomly superimpose the graphic on the medical image to synthesize it. If the size of the graphic is accepted by the selection unit 22A, the graphic is synthesized on the medical image based on the size accepted in the size acceptance process.
[0049] Furthermore, the graphic composition unit 22B may generate multiple composite images by changing the position and size of the selected graphic. That is, the graphic composition unit 22B may generate multiple variations of the composite image by changing the position at which the selected graphic is superimposed on the medical image or the size of the selected graphic. This makes it possible to easily generate more training data. The graphic composition unit 22B determines the composition mode of the composite image either by user specification or automatically, and generates the composite image.
[0050] The learning unit 22C performs a learning process and causes the learning model 30 to learn using the synthetic image. Specifically, the learning unit 22C optimizes the parameters of the learning model 30 based on the error between the estimated result output by the learning model 30 when the synthetic image is input and the correct data. The learning unit 22C will be described in detail later.
[0051] 3 is a diagram illustrating a specific data flow in the machine learning system 10. In the following description, an ultrasound image is used as an example of a medical image.
[0052] The selection unit 22A selects an ultrasound image P1 from the image database 14 in response to a user selection. For example, the user selects an ultrasound image P1 via the operation unit 20. The selection unit 22A also selects one or more figures Q from the figure database 16 in response to the user selection. Specific examples of the image database 14 and the figure database 16 will be described below.
[0053] FIG. 4 is a diagram showing an example of an ultrasound image P1 stored in the image database 14. As shown in FIG.
[0054] Ultrasound image P1 has regions of interest that depict organs or lesions. Specifically, ultrasound image P1 has regions of interest T1, T2, T3, and T4. Region of interest T1 is an image of organ A, region of interest T2 is an image of organ B, region of interest T3 is an image of lesion C, and region of interest T4 is an image of lesion D.
[0055] FIG. 5 is a diagram showing a specific example of the storage configuration of the image database 14. As shown in FIG.
[0056] The image database 14 stores a plurality of ultrasound images P. Each ultrasound image P is assigned an image ID. The image ID of ultrasound image P1 is 001. The image database 14 also stores the positions of the regions of interest T1 to T4 of the ultrasound image P1 in XY coordinates, associated with the image ID of the ultrasound image P1. The sizes of the regions of interest T1 to T4 are also stored in terms of the number of pixels. The types of the regions of interest T1 to T4 are also indicated. Note that some or all of the positions, sizes, and types of the regions of interest stored in the image database 14 in association with the image IDs are used as correct answer data F in machine learning, corresponding to the estimation results output by a learning model 30, which will be described later. Note that, although omitted in FIG. 5, the image database 14 also stores a plurality of ultrasound images P in addition to the ultrasound image P1.
[0057] FIG. 6 is a diagram showing a specific example of the storage configuration of the graphic database 16. As shown in FIG.
[0058] As shown in FIG. 6 , the graphic database 16 stores a graphic ID, a graphic, a minimum graphic value, a maximum graphic value, and a compositing area, all of which are associated with each other. Specifically, an arrow has a graphic ID of 001, a minimum graphic value of 32px, a maximum graphic value of 512px, and a compositing area that covers the entire ultrasound image. A line has a graphic ID of 002, a minimum graphic value of 8px, a maximum graphic value of 960px, and a compositing area that covers the region of interest in the ultrasound image. A rectangle has a graphic ID of 003, a minimum graphic value of 64px, a maximum graphic value of 256px, and a compositing area that covers the entire region. A circle has a graphic ID of 004, a minimum graphic value of 16px, a maximum graphic value of 256px, and a compositing area that covers the entire region. A triangle has a graphic ID of 005, a minimum graphic value of 128px, a maximum graphic value of 512px, and a compositing area outside the region of interest. The graphic database 16 also stores other graphics in addition to the graphics described above.
[0059] 3, the ultrasound image P1 and the selected figure Q selected by the selection unit 22A are input to the figure synthesis unit 22B. The figure synthesis unit 22B generates a composite image C using the ultrasound image P1 and the figure Q. A specific example of the composite image C will be described below.
[0060] 7 to 10 are diagrams showing examples of composite images synthesized by the graphic synthesis unit 22B.
[0061] In the example shown in FIG. 7, an ultrasound image P1, a graphic (arrow) Q1, and a graphic (line) Q2 are selected, and a composite image C1 is generated by combining these.
[0062] In the composite image C1, a figure Q1 is superimposed near the region of interest T3 to indicate the region of interest based on information (X3, Y3) about the position of the region of interest T3 and / or information (Z3px) about the size of the region of interest. In addition, in the composite image C1, a figure Q2 is superimposed to indicate both ends of the region of interest T2 based on information (X2, Y2) about the position of the region of interest T2 and / or information (Z2px) about the size of the region of interest. In this way, the composite image C1 is generated by superimposing the figures Q1 and Q2 on the ultrasound image P1, just as a doctor would add them when actually observing the ultrasound image P1.
[0063] In the example shown in FIG. 8, an ultrasound image P1 and a figure (rectangle) Q3 are selected, and a composite image C2 is generated by combining these.
[0064] In the composite image C2, a figure Q3 is superimposed based on information about the positions and / or sizes of the regions of interest T1 to T4 (see FIG. 5) so that the regions of interest T3 and T4 are contained within the rectangle, and a part of the figure Q3 crosses the regions of interest T1 and T2. In this way, the composite image C2 is generated by superimposing the figure Q3 on the ultrasound image P1 so that the figure Q3 can be used as an aid when a doctor observes the regions of interest T3 and T4.
[0065] In the example shown in FIG. 9, an ultrasound image P1 and a figure (rectangle) Q5 are selected, and a composite image C3 is generated by combining these.
[0066] In the composite image C3, the position of superimposition is determined using a random number table or the like, and the figure Q5 is randomly superimposed. In reality, when a doctor attaches a figure to an ultrasound image, he or she points to the region of interest (for example, by surrounding the region of interest with a rectangle or by indicating the region of interest with an arrow), but in the case shown in FIG. 9, the figure Q5 is superimposed regardless of the regions of interest T1 to T4. In other words, the figure Q5 is superimposed on the ultrasound image P1 in a form different from that attached by the doctor (independent of the position of the region of interest). By using such a composite image C3 as training data, the region of interest can also be effectively recognized.
[0067] In the example shown in FIG. 10, an ultrasound image P1, a figure (arrow) Q6, a figure (line) Q7, a figure (line) Q8, and a figure (arrow) Q9 are selected, and a composite image C4 is generated by combining these.
[0068] In the composite image C4, the figures Q6 and Q8 are superimposed so as to exclude the regions of interest T1 to T4 based on information relating to the positions and / or sizes of the regions of interest T1 to T4 (see FIG. 5). Furthermore, in the composite image C4, the figures Q7 and Q9 are arranged randomly. In this way, the figures Q6 to Q9 are superimposed randomly on the composite image C4, and the figures Q6 to Q9 are superimposed on the ultrasound image P1 in a form different from that when a doctor adds figures to an ultrasound image (irrespective of the positions of the regions of interest). By using such a composite image C3 as training data, the regions of interest can also be recognized effectively.
[0069] 3, the composite image C synthesized by the graphic synthesis unit 22B is input to the learning model 30. In addition, the position and type of the region of interest of the ultrasound image P1 selected by the selection unit 22A are input to the learning unit 22C as correct answer data F. Below, machine learning of the learning model 30 by the learning unit 22C will be described.
[0070] 11 is a functional block diagram showing the main functions of the learning unit 22C and the learning model 30. The learning unit 22C includes an error calculation unit 54 and a parameter update unit 56. Furthermore, correct answer data F is input to the learning unit 22C.
[0071] The learning model 30 is a recognizer that performs image recognition of the position and type of a region of interest within the ultrasound image P. The learning model 30 has a multi-layer structure and holds a multiplicity of weighting parameters. The learning model 30 changes from an unlearned model to a trained model by updating the weighting parameters from their initial values to optimal values.
[0072] This learning model 30 includes an input layer 52A, an intermediate layer 52B, and an output layer 52C. The input layer 52A, the intermediate layer 52B, and the output layer 52C each have a structure in which a plurality of "nodes" are connected by "edges." A synthetic image C, which is the learning target, is input to the input layer 52A.
[0073] The intermediate layer 52B extracts features from the image input from the input layer 52A. The intermediate layer 52B includes multiple sets of convolutional layers and pooling layers, and a fully connected layer. The convolutional layer performs a convolution operation using a filter on nearby nodes in the previous layer to obtain a feature map. The pooling layer reduces the feature map output from the convolutional layer to create a new feature map. The fully connected layer connects all nodes in the previous layer (here, the pooling layer). The convolutional layer is responsible for feature extraction, such as edge extraction from the image, and the pooling layer is responsible for providing robustness to the extracted features so that they are not affected by translation, etc. Note that the intermediate layer 52B is not limited to cases where a convolutional layer and a pooling layer are combined into one set, but may also include cases where convolutional layers are consecutive and a normalization layer.
[0074] The output layer 52C is a layer that outputs the recognition results of the position and type of the region of interest in the ultrasound image P based on the features extracted by the intermediate layer 52B.
[0075] The trained learning model 30 outputs the recognition results of the position and type of the region of interest.
[0076] The coefficients of the filters applied to each convolution layer of the learning model 30 before learning, the offset values, and the weights of the connections to the next layer in the fully connected layer are set to arbitrary initial values.
[0077] The error calculation unit 54 obtains the recognition result output from the output layer 52C of the learning model 30 and the correct answer data F for the input image, and calculates the error between them. Possible methods for calculating the error include softmax cross entropy or mean squared error (MSE). Note that the correct answer data F for the input image (synthetic image C1) is, for example, data indicating the positions and types of the attention regions T1 to T4.
[0078] The parameter update unit 56 adjusts the weight parameters of the learning model 30 using the error backpropagation method based on the error calculated by the error calculation unit 54.
[0079] This parameter adjustment process is repeated, and learning is repeated until the difference between the output of the learning model 30 and the correct answer data F becomes small.
[0080] The learning unit 22C uses a data set of at least the synthetic image C1 and the ground truth data F to optimize each parameter of the learning model 30. The learning by the learning unit 22C may use a mini-batch method in which a certain number of data sets are extracted, batch processing of learning is performed using the extracted data sets, and this is repeated.
[0081] Next, we will explain a learning method performed using the machine learning system 10. Note that each step of the learning method is performed by the processor 22 executing a program.
[0082] FIG. 12 is a flow diagram illustrating a learning method performed using the machine learning system 10.
[0083] First, the selection unit 22A accepts the selection of an ultrasound image (medical image) (selection accepting step: step S10). For example, the user checks the ultrasound image displayed on the display unit 28 and selects the ultrasound image via the operation unit 20. Next, the selection unit 22A accepts the selection of a figure (selection step: step S11). For example, the user checks the figure displayed on the display unit 28 and selects the figure via the operation unit 20. In this case, the size of the selected figure may also be selected, and the figures may be superimposed and synthesized at the selected size. Thereafter, the figure synthesis unit 22B synthesizes the selected ultrasound image and the selected figure to generate a synthesized image (figure synthesis step: step S12). Thereafter, the learning unit 22C causes the learning model 30 to perform machine learning using the synthesized image (learning step: step S13).
[0084] As described above, according to this embodiment, medical images stored in the image database 14 and figures stored in the figure database 16 are selected, and a composite image is generated from the medical images and figures. In this embodiment, the learning model 30 is trained using the composite image, so a large number of composite medical images can be trained efficiently and easily.
[0085] <Second embodiment> Next, a second embodiment will be described. In this embodiment, a selected ultrasound image is used for machine learning at least twice. For example, in this embodiment, an ultrasound image without a graphic superimposed thereon is used for machine learning the first time, and an ultrasound image with a graphic superimposed thereon is used for machine learning the second time.
[0086] 13 is a diagram illustrating a specific data flow in the machine learning system 10 of this embodiment. Note that the same reference numerals are used to denote parts that have already been explained in FIG. 3, and explanations thereof will be omitted.
[0087] In the machine learning system 10 of this embodiment, a composite image C and an ultrasound image P1 are input to the learning model 30. Here, the learning model 30 outputs an estimation result for each input image. Specifically, the learning model 30 receives the composite image C and outputs an estimation result, and receives the ultrasound image P1 and outputs an estimation result.
[0088] 14 is a functional block diagram showing the main functions of the learning unit 22C and the learning model 30 of this embodiment. Note that the same reference numerals are used to denote parts that have already been explained in FIG. 11, and explanations thereof will be omitted.
[0089] In this embodiment, machine learning is performed using the composite image C1 and the ultrasound image P1.
[0090] Specifically, a composite image C1 and an ultrasound image P1 are input to the input layer 52A. When the composite image C1 is input, the learning model 30 outputs an estimation result for the composite image C1. When the ultrasound image P1 is input, the learning model 30 outputs an estimation result for the ultrasound image P1. Then, the error between the estimation result and the correct answer data F is calculated, and the parameters are updated by the parameter update unit 56 based on the error. Next, the ultrasound image P1 is input to the learning model 30, and the position, type, etc. of the region of interest are estimated and output from the output layer 52C. Then, the error between the estimation result and the correct answer data is calculated, and the parameters are updated by the parameter update unit 56 based on the error. Note that the correct answer data F is data indicating the position and type of the region of interest in the ultrasound image P1, and therefore can be used in common in the machine learning of the composite image C1 and the ultrasound image P1.
[0091] FIG. 15 is a flow diagram illustrating a learning method performed using the machine learning system 10.
[0092] First, the selection unit 22A accepts the selection of an ultrasound image (step S20). Next, the selection unit 22A accepts the selection of a figure (step S21). Thereafter, the figure synthesis unit 22B synthesizes the selected ultrasound image and the selected figure to generate a synthesized image (step S22). Next, the learning unit 22C causes the learning model 30 to perform machine learning using the synthesized image (step S23). Thereafter, the learning unit 22C causes the learning model 30 to perform machine learning using an ultrasound image that has not been synthesized (step S24).
[0093] In this embodiment, machine learning is performed using a composite image and an ultrasound image that has no graphic combined therein. How to use the composite image and the ultrasound image for machine learning will be described below.
[0094] FIG. 16 is a diagram conceptually illustrating an example of training data according to the present invention.
[0095] FIG. 16(A) shows the training data used in the first embodiment. In the first embodiment, synthetic images are used as training data, so mini-batches composed of synthetic images are prepared. In the case shown in FIG. 16(A), training data for batch A 80, batch B 82, batch C 84, batch D 86, and batch E 88 are prepared, and each batch is composed of, for example, 200 synthetic images. Therefore, in one epoch, batch A 80, batch B 82, batch C 84, batch D 86, and batch E 88 are all input to the learning model 30, and machine learning is performed.
[0096] FIG. 16(B) shows the training data used in the second embodiment. In the second embodiment, composite images and ultrasound images without graphics are used as training data, so mini-batches composed of composite images and ultrasound images are prepared. In the case shown in FIG. 16(B), training data for batch A 80, batch a 90, batch B 82, batch b 92, batch C 84, batch c 94, batch D 86, batch d 96, batch E 88, and batch e 98 are prepared. Note that each of batch A 80, batch B 82, batch C 84, batch D 86, and batch E 88 is composed of, for example, 200 composite images. Also, each of batch a 90, batch b 92, batch c 94, batch d 96, and batch e 98 is composed of, for example, 200 ultrasound images. Note that the same ultrasound images are used for batch A 80 and batch a 90. Specifically, the ultrasound images used for the composite images that make up batch A 80 make up batch a 90, and the same is true for the other batches. In one epoch, batch A 80, batch a 90, batch B 82, batch b 92, batch C 84, batch c 94, batch D 86, batch d 96, batch E 88, and batch e 98 are all input into learning model 30, and machine learning is performed.
[0097] Figure 16(C) shows another example of training data used in the second embodiment. In the case shown in Figure 16(C), a first epoch consisting of A batch 80, B batch 82, C batch 84, D batch 86, and E batch 88, and a second epoch consisting of a batch 90, b batch 92, c batch 94, d batch 96, and e batch 98 are performed. Note that, as explained in Figure 16(B), the same ultrasound images are used in A batch 80 and a batch 90.
[0098] As described above, in this embodiment, a composite image and a non-composite ultrasound image are used as training data, so that the learning model 30, which recognizes the position and type of the region of interest regardless of the superimposed figure, can be trained more effectively.
[0099] <Third embodiment> Next, a third embodiment will be described. In this embodiment, a recognizer configured with a trained model for which machine learning has been performed by the above-described machine learning system 10 will be described. This recognizer is installed in an ultrasound endoscope system.
[0100] FIG. 17 is a schematic diagram showing the overall configuration of an ultrasound endoscope system (ultrasound diagnostic apparatus) that includes, as a recognizer 206, a learning model 30 (trained model) for which machine learning has been completed by the above-described machine learning system 10.
[0101] As shown in FIG. 17, the ultrasonic endoscope system 102 includes an ultrasonic scope 110, an ultrasonic processor device 112 that generates ultrasonic images, an endoscopic processor device 114 that generates endoscopic images, a light source device 116 that supplies illumination light to the ultrasonic scope 110 to illuminate the inside of the body cavity, and a monitor 118 that displays ultrasonic images and endoscopic images.
[0102] The ultrasound scope 110 includes an insertion section 120 that is inserted into a body cavity of a subject, a handheld operation section 122 that is connected to the base end of the insertion section 120 and that is operated by the surgeon, and a universal cord 124 that has one end connected to the handheld operation section 122. The other end of the universal cord 124 is provided with an ultrasound connector 126 that is connected to the ultrasound processor device 112, an endoscope connector 128 that is connected to the endoscope processor device 114, and a light source connector 130 that is connected to the light source device 116.
[0103] The ultrasonic scope 110 is detachably connected to the ultrasonic processor 112, the endoscope processor 114, and the light source 116 via these connectors 126, 128, and 130. In addition, a tube 132 for supplying air and water and a tube 134 for suction are connected to the light source connector 130.
[0104] The monitor 118 receives the video signals generated by the ultrasonic processor device 112 and the endoscope processor device 114 and displays the ultrasonic image and the endoscopic image. The display of the ultrasonic image and the endoscopic image can be switched appropriately and displayed on the monitor 118 by switching only one of the images, or both images can be displayed simultaneously.
[0105] The handheld operation section 122 is provided with an air / water supply button 136 and a suction button 138, a pair of angle knobs 142, and a treatment tool insertion port 144.
[0106] The insertion section 120 has a distal end, a proximal end, and a longitudinal axis 120a, and is composed of, in order from the distal end side, a distal end body 150 made of a hard member, a bending section 152 connected to the proximal end side of the distal end body 150, and a thin, long, flexible soft section 154 connecting the proximal end side of the bending section 152 to the distal end side of the hand-operated operation section 122. That is, the distal end body 150 is provided on the distal end side of the longitudinal axis 120a of the insertion section 120. The bending section 152 is remotely operated to be bent by rotating a pair of angle knobs 142 provided on the hand-operated operation section 122. This allows the distal end body 150 to be oriented in a desired direction.
[0107] An ultrasonic probe 162 and a bag-shaped balloon 164 that encases the ultrasonic probe 162 are attached to the tip body 150. The balloon 164 can be inflated or deflated by supplying water from a water tank 170 or by suctioning water from inside the balloon 164 with a suction pump 172. The balloon 164 is inflated until it abuts against the inner wall of the body cavity to prevent attenuation of the ultrasonic waves and ultrasonic echoes (echo signals) during ultrasonic observation.
[0108] An endoscopic observation unit (not shown) having an observation unit and an illumination unit equipped with an objective lens, an imaging element, etc. is attached to the distal end body 150. The endoscopic observation unit is provided behind the ultrasound probe 162 (on the handheld operation unit 122 side).
[0109] FIG. 18 is a block diagram illustrating an embodiment of an ultrasound processor device.
[0110] The ultrasound processor device 112 shown in FIG. 18 recognizes the position and type of a region of interest in an ultrasound image based on sequentially acquired time-series ultrasound images, and notifies the user of information indicating the recognition result.
[0111] 18 is composed of a transmitting / receiving unit 200, an image generating unit 202, a CPU (Central Processing Unit) 204, a recognizer 206, a display control unit 210, and a memory 212, and the processing of each unit is realized by one or more processors. Note that the recognizer 206 is equipped with the learning model 30 that has been trained by the above-mentioned machine learning system 10, with the parameters retained as they are.
[0112] The CPU 204 operates based on various programs including the ultrasound image processing program according to the present invention stored in the memory 212, and controls the transmitter / receiver unit 200, the image generator unit 202, the recognizer 206, and the display controller unit 210, and also functions as a part of each of these units.
[0113] The transmitting / receiving unit 200 and the image generating unit 202 sequentially acquire ultrasound images in time series.
[0114] The transmitting section of the transmitting / receiving section 200 generates a plurality of drive signals to be applied to a plurality of ultrasonic transducers of the ultrasonic probe 162 of the ultrasonic scope 110, and applies the plurality of drive signals to the plurality of ultrasonic transducers by giving each of the drive signals a respective delay time based on a transmission delay pattern selected by a scanning control section (not shown).
[0115] The receiving section of the transmitting / receiving unit 200 amplifies the multiple detection signals output from the multiple ultrasonic transducers of the ultrasonic probe 162, and converts the analog detection signals into digital detection signals (also called RF (Radio Frequency) data). This RF data is input to the image generating unit 202.
[0116] The image generation unit 202 performs reception focusing processing by adding together the detection signals represented by the RF data and applying delay times to the detection signals based on the reception delay pattern selected by the scan control unit. This reception focusing processing forms sound ray data in which the focus of the ultrasonic echo is narrowed.
[0117] The image generation unit 202 further corrects for attenuation due to distance according to the depth of the ultrasonic reflection position using STC (Sensitivity Time Gain Control) for the sound ray data, then generates envelope data by performing envelope detection processing using a low-pass filter or the like, and stores the envelope data for one frame, or more preferably for multiple frames, in a cine memory (not shown).The image generation unit 202 performs preprocessing such as log (logarithmic) compression and gain adjustment on the envelope data stored in the cine memory to generate a B-mode image.
[0118] In this way, the transmitting and receiving unit 200 and the image generating unit 202 sequentially acquire time-series B-mode images (hereinafter referred to as "ultrasound images").
[0119] The recognizer 206 performs a process of recognizing information about the position of a region of interest in the ultrasound image based on the ultrasound image, and a process of classifying the region of interest into one of a plurality of classes (types) based on the ultrasound image. The recognizer 206 is configured with a trained model, and machine learning is performed by the above-mentioned machine learning system 10.
[0120] The regions of interest in this example are various organs in the ultrasound image (tomographic image of a B-mode image), such as the pancreas, main pancreatic duct, spleen, splenic vein, splenic artery, and gallbladder.
[0121] When time-series ultrasound images are input sequentially, the recognizer 206 recognizes the position of the region of interest for each input ultrasound image, outputs information related to the position, and also recognizes to which of multiple classes the region of interest belongs, and outputs information indicating the recognized class (class information).
[0122] The position of the region of interest may be, for example, the center position of a rectangle surrounding the region of interest. In this example, the class information is information indicating the type of organ.
[0123] The display control unit 210 is composed of a first display control unit 210A that displays time-series ultrasound images on the monitor 118, which is a display unit, and a second display control unit 210B that displays information related to the region of interest on the monitor 118.
[0124] The first display control unit 210A causes the monitor 118 to display the ultrasound images sequentially acquired by the transmitting and receiving unit 200 and the image generating unit 202. In this example, the monitor 118 displays a moving image showing an ultrasound tomographic image.
[0125] In addition, the first display control unit 210A performs a reception process to receive a phrase command from the handheld operation unit 122 of the ultrasound scope 110, and when, for example, the freeze button on the handheld operation unit 122 is operated and a freeze instruction is received, the first display control unit 210A performs a process to switch the sequential display of ultrasound images displayed on the monitor 118 to a fixed display of one ultrasound image (the current ultrasound image).
[0126] The second display control unit 210B displays class information indicating the classification of the region of interest recognized by the recognizer 206 in a superimposed manner at the position of the region of interest of the ultrasound image displayed on the monitor 118.
[0127] Furthermore, when second display control section 210B receives a freeze instruction, it also fixes the relative position of the class information with respect to the region of interest during the period in which the ultrasound image displayed on monitor 118 is fixedly displayed as a still image.
[0128] As described above, in this embodiment, the learning model 30 that has undergone machine learning in the machine learning system 10 is used as the recognizer 206. This allows for accurate recognition of the region of interest in the ultrasound image.
[0129] <Other> In the above embodiment, the hardware structure of the processing units (e.g., the selection unit 22A, the graphic synthesis unit 22B, and the learning unit 22C) that execute various processes is the following various processors. The various processors include a CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) and functions as various processing units, a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacture, and a dedicated electrical circuit, such as an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing specific processes.
[0130] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types (for example, multiple FPGAs, or a combination of a CPU and an FPGA). Also, multiple processing units may be configured with a single processor. Examples of multiple processing units configured with a single processor include, first, a configuration in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server, and this processor functions as multiple processing units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a system-on-chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.
[0131] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.
[0132] The above-described configurations and functions can be realized by any hardware, software, or a combination of both. For example, the present invention can be applied to a program that causes a computer to execute the above-described processing steps (processing procedures), a computer-readable recording medium (non-transitory recording medium) on which such a program is recorded, or a computer on which such a program can be installed.
[0133] Although examples of the present invention have been described above, it goes without saying that the present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the present invention. [Explanation of symbols]
[0134] 10: Machine learning systems 12: Communications Department 14: Image database 16: Graphic database 20:Operation unit 22: Processor 22A: Selection section 22B: Graphic synthesis section 22C: Learning Department 24:RAM 26:ROM 28:Display section 30: Learning model 32: Bus
Claims
1. A machine learning system comprising: an image database in which a plurality of medical images are stored; a figure database in which figures are stored that are used as observation aids when observing the medical images to be superimposed on the medical images; a processor; and a learning model, The processor: a selection receiving process for receiving a selection of the medical image including the region of interest stored in the image database and a selection of the figure stored in the figure database; a graphic synthesis process for synthesizing the selected medical image and the graphic to generate a synthesized image; a learning process for causing the learning model to perform learning using the synthetic image to output at least one of the position of the region of interest and the type of the region of interest as an estimation result; A machine learning system that performs
2. The machine learning system of claim 1, wherein the medical image accepted in the selection acceptance process is used at least twice in the learning process, the first time using the medical image that has not undergone the graphic synthesis process, and the second time using the synthetic image that has undergone the graphic synthesis process to train the learning model.
3. the processor performs a size reception process for receiving a size of the graphics to be synthesized in the graphics synthesis process; The machine learning system according to claim 1 or 2, wherein the graphic composition process comprises combining the graphic with the medical image based on the size accepted in the size acceptance process.
4. The machine learning system according to any one of claims 1 to 3, wherein the graphic database stores at least one of the maximum and minimum values of the size of the graphic and area information that can be synthesized into the medical image in association with the graphic.
5. the image database stores information about the region of interest in association with the medical image; The machine learning system according to claim 1 , wherein the graphic synthesis process uses information about the region of interest to synthesize the graphic so that at least a portion of the graphic is superimposed on the region of interest.
6. the image database stores information about the region of interest in association with the medical image; The machine learning system according to claim 1 , wherein the graphic synthesis process uses information about the region of interest to synthesize the graphic while excluding the region of interest.
7. In the graphic composition process, a plurality of the composite images are generated by changing the position and size of the selected graphic, The machine learning system according to claim 1 , wherein the learning process causes the learning model to learn using a plurality of the synthetic images.
8. The machine learning system according to claim 1 , wherein the medical image is an image obtained from an ultrasound diagnostic device.
9. A trained model for recognizing regions of interest from medical images, comprising: The parameters of the trained model are trained by the machine learning system according to any one of claims 1 to 8, A trained model that causes a computer to accept the medical image as input, process the input medical image based on the parameters, recognize the region of interest from the medical image, and output the recognition results.
10. A learning method using a machine learning system including an image database storing a plurality of medical images, a figure database storing figures to be superimposed on the medical images and used as an observation aid when observing the medical images, a processor, and a learning model, The processor: a selection receiving step of receiving a selection of the medical image including the region of interest stored in the image database and a selection of the figure stored in the figure database; a graphic synthesis step of synthesizing the selected medical image and the graphic to generate a synthesized image; a learning process of causing the learning model to perform learning using the synthetic image to output at least one of the position of the region of interest and the type of the region of interest as an estimation result; A learning method in which:
11. A program for executing a learning method using a machine learning system including an image database storing a plurality of medical images, a figure database storing figures used as observation aids when observing the medical images to be superimposed on the medical images, a processor, and a learning model, the processor, a selection receiving step of receiving a selection of the medical image including the region of interest stored in the image database and a selection of the figure stored in the figure database; a graphic synthesis step of synthesizing the selected medical image and the graphic to generate a synthesized image; a learning process of causing the learning model to perform learning using the synthetic image to output at least one of the position of the region of interest and the type of the region of interest as an estimation result; A program that executes the following.
12. A non-transitory computer-readable recording medium having the program according to claim 11 recorded thereon.
Citation Information
Patent Citations
Data generation device, data generation method, and data generation program
JP2019087078A
Diagnosis support apparatus and x-ray CT apparatus
JP2020192006A
Image discrimination model construction method, image discrimination model, and image discrimination method
JP2021005266A
Radiation image processing device and program
JP2021013685A
Image processing apparatus, object detection apparatus, and image processing method
WO2016157499A1