Identification device, learning device, identification method, learning method, identification program, and learning program
The identification device uses multi-task learning and semantic segmentation to accurately identify abnormal regions in dental X-ray images, addressing the challenges of complex image analysis and diverse calcification patterns.
Patent Information
- Application Number
- JP2023190310
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-05-19
AI Technical Summary
Existing methods for identifying calcification in dental panoramic X-ray photographs face challenges due to the complexity of the images, diverse shapes and sizes of calcified areas, and ambiguous boundary shapes, which hinder accurate identification of abnormal regions.
The proposed solution involves an identification device and method that utilize multi-task learning with semantic segmentation to generate a second image from the first image, allowing for accurate identification of abnormal regions. This approach combines a vision transformer and a CNN for global and local feature extraction, and employs a residual neural network for identification.
The method achieves high accuracy in identifying the presence or absence of abnormal regions, improving upon previous methods by effectively handling complex images and diverse calcification patterns without relying on artificial thresholds.
Smart Images

Figure 2025077824000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an identification device, a learning device, an identification method, a learning method, an identification program, and a learning program.
Background Art
[0002] Atherosclerosis is a cause of stroke and heart attack, but it has no symptoms and there are few opportunities to visit a doctor. Calcification of carotid artery atherosclerosis may be found in dental panoramic X-ray photographs taken by dentists. If the presence or absence of calcification can be identified (discriminated) during a dentist's examination, it can create an opportunity to visit a doctor and lead to early treatment.
[0003] However, dental panoramic X-ray photographs are very complex images that include bones and other tissues, and the shapes and sizes of calcified areas are also diverse. In addition, the shape of the boundary is ambiguous, making it even more difficult to identify.
[0004] Patent Document 1 discloses using U-Net and ResNet (Residual Neural Networks) in medical images. Patent Document 2 discloses a technique for detecting calcification by segmentation and a neural network.
[0005] Non-Patent Document 1 discloses applying a Vision Transformer to a U-Net, which is a CNN (Convolutional Neural Network) connected in series. Non-Patent Document 2 discloses connecting a CNN and a Transformer in series in a U-Net.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
[0007] [Non-Patent Document 1] Hu Cao et. al., Swin-Unet: Unet-like Pure Transformer for medical Image Segmentation, arXiv, 12 May 2021 [Non-Patent Document 2] Jieneng Chen. al., TransUNet: Transformers Make Strong Encorders for Medical Image Segmentation, arXiv, 8 Feb 2021 [Non-Patent Document 3] S. Gao et. al., Res2Net: A new multi-scale backbone architecture, IEEE Trans. on Pattern Analysis and Machine Intelligence, vol 43, no.2, pp.652-662, 2021 [Non-Patent Document 4] A. Navon et. al., Multi-task learning as a bargaining game, arViv:2202.01017, 2022 [Non-Patent Document 5] T. Murano et. al., New Method of Detecting Calcification Regions in Dental Panoramic Radiographs Based on U-PraNet, 2021 20th International Symposium on Communication and Information Technologies (ISCIT), Tottori, Japan, Oct.20-22, W1-3, pp.11-14, 2019 [Non-Patent Document 6] Y. Zhang et. al., TransFuse: Fusing Transformers and CNNs for Medicial Image Segmentation, arXiv: 2102.08005, 2021 [Summary of the Invention] [Problems to be Solved by the Invention]
[0008] In recent years, the introduction of AI (Artificial Intelligence) into medical images has been progressing. In image processing in AI, CNN has been widely used.
[0009] In CNN, in order to generate features by convolution, it is good at aggregating nearby features into features, but it is difficult to obtain features considering distant features. That is, since CNN cannot perform global learning, it is difficult to perform learning considering the relative positional relationship of objects.
[0010] Here, in medical images, the positional relationship between abnormal regions such as tumors and bones and other tissues may be important information. That is, it is required to consider global information such as surrounding information.
[0011] However, with the methods disclosed in the prior art documents, global learning may be difficult. Also, with the methods disclosed in the prior art documents, it may be difficult to handle cases where there are large differences between target objects and cases where there are objects that are not similar to the target objects. Further, with the methods disclosed in the prior art documents, parameters such as thresholds may have to be determined artificially. Also, with the methods disclosed in the prior art documents, it may be difficult to determine the presence or absence of calcification.
[0012] One aspect of the present invention aims to accurately identify the presence or absence of an abnormal region.
Means for Solving the Problems
[0013] To solve the above problems, an identification device according to the present invention includes an acquisition unit that acquires a first image, a generation unit that generates a second image, which is an image after segmentation, by performing semantic segmentation on the first image, and an identification unit that identifies the presence or absence of an abnormal region in the second image, and uses multi-task learning with reference to a loss function in the generation model of the generation unit and a loss function in the identification model of the identification unit for the learning of the generation unit and the identification unit.
[0014] To solve the above problems, a learning device according to the present invention includes an acquisition unit that acquires a first image, a generation unit that generates a second image, which is an image after segmentation, by performing semantic segmentation on the first image, an identification unit that identifies the presence or absence of an abnormal region in the second image, and a learning unit that uses multi-task learning with reference to a loss function in the generation model of the generation unit and a loss function in the identification model of the identification unit for the learning of the generation unit and the identification unit.
[0015] In order to solve the above problems, the identification method according to the present invention includes an acquisition step of acquiring a first image, a generation step of generating a second image which is an image after segmentation by performing semantic segmentation on the first image, and an identification step of identifying the presence or absence of an abnormal region in the second image, and multi-task learning with reference to a loss function in the generation model of the generation step and a loss function in the identification model of the identification step is used for the learning of the generation step and the identification step.
[0016] In order to solve the above problems, the learning method according to the present invention includes an acquisition step of acquiring a first image, a generation step of generating a second image which is an image after segmentation by performing semantic segmentation on the first image, an identification step of identifying the presence or absence of an abnormal region in the second image, and a learning step of using multi-task learning with reference to a loss function in the generation model of the generation step and a loss function in the identification model of the identification step for the learning of the generation step and the identification step.
[0017] The learning device or identification device according to each aspect of the present invention may be realized by a computer. In this case, a control program for the learning device or identification device that realizes the learning device or identification device by operating the computer as each part (software element) included in the learning device or identification device, and a computer-readable recording medium on which it is recorded also fall within the scope of the present invention.
Advantages of the Invention
[0018] According to the present invention, the presence or absence of an abnormal region can be identified with high accuracy.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Best Mode for Carrying Out the Invention
[0020] 〔Embodiment 1〕 Hereinafter, an embodiment according to one aspect of the present invention (hereinafter, also referred to as "this embodiment") will be described with reference to the drawings. In the drawings, the same or corresponding parts are denoted by the same reference numerals and their description will not be repeated.
[0021] (Outline of Learning Device 1) FIG. 1 is a block diagram showing the configuration of the main part of the learning device 1 according to this embodiment. The learning device 1 includes a control unit 100. The learning device 1 may include a storage unit 200 and an input / output unit 300. The control unit 100 includes a generation unit 10, a discrimination unit 40, a learning unit 60, and an acquisition unit 70. The control unit 100 may include a pre-training unit 50. The generation unit 10 includes an initial guidance area estimation mechanism 20 and a boundary estimation mechanism 30.
[0022] The control unit 100 is a functional block that performs learning processing and identification processing in the learning device 1. That is, for the image acquired by the acquisition unit 70, it performs learning processing of the model or identification processing using the learned model.
[0023] The storage unit 200 stores the model generated or used by the control unit 100. The storage unit 200 also stores the parameters used for the model. Note that the storage unit 200 may store a program for operating the learning device 1.
[0024] The input / output unit 300 has a function of inputting and outputting inputs from a keyboard, a mouse, a touch panel, etc. and outputs to a display, etc. Note that the input / output unit 300 may output the model learned by the learning device 1 to an external device. Specifically, it is a case where the model stored in the storage unit 200 is saved to a storage medium and moved.
[0025] FIG. 2 is a diagram showing the processing in the stage of performing learning in the learning device 1. The acquisition unit 70 acquires the first image 3 which is an input image and inputs it to the generation unit 10. The generation unit 10 processes the first image 3 based on the generation model 11, derives various feature amounts, and then generates the second image 4 using the feature amounts. The identification unit 40 inputs the second image 4. The identification unit 40 processes the second image 4 based on the identification model 41 and identifies the presence or absence of a calcified region as an abnormal region. Note that the generation unit 10 and the identification unit 40 are connected in series.
[0026] Note that the presence or absence of the abnormal region output by the identification unit 40 may be the probability of the presence or absence of the abnormal region. Also, in the case where there is no abnormal region, a vector of (1,0) may be output, and in the case where there is an abnormal region, a vector of (0,1) may be output, which is a one-hot representation.
[0027] The learning unit 60 compares the result of the presence or absence of the abnormal region output by the identification unit 40 with the teacher data. Based on the comparison result, backpropagation is performed on the generation model 11 and the identification model 41 so as to reduce the loss function. That is, the first loss 12 in the generation model 11 and the second loss 42 in the identification model 41 are optimized in cooperation with each other.
[0028] Here, it is preferable that the generation unit 10 uses more of the first images 3 of the non-healthy subjects. This is because the features of the abnormal region can be learned by learning a large number of images including the abnormal region in the non-healthy subjects.
[0029] Also, in the generation unit 10, if learning is performed using the first images 3 of the healthy subjects, there is a risk of degrading the learning regarding the abnormal region. Therefore, it is preferable not to use the first images 3 of the healthy subjects as much as possible.
[0030] However, in the identification unit 40, not only the first images 3 of the non-healthy subjects but also the first images 3 of the healthy subjects are necessary. Therefore, it is preferable to train the generation unit 10 and the identification unit 40 using both types of images.
[0031] Here, it is preferable to train the generation unit 10 and the identification unit 40 while preventing the degradation of learning in the generation unit 10. For this purpose, multi-task learning is adopted in the present embodiment.
[0032] (Overview of the generation unit 10) The generation unit 10 has a function of generating the second image 4 from the input first image 3. The generation unit 10 makes the feature amounts in the first image 3 manifest by means of the initial guidance region estimation mechanism 20. Then, the probability that each pixel is an abnormal region is estimated using the manifested feature amounts. The boundary estimation mechanism 30 estimates the probability in more detail in the peripheral part of the abnormal region estimated by the initial guidance region estimation mechanism 20. The probability of each pixel is reconstructed into the form of an image to generate the second image 4. That is, the generation unit 10 includes the initial guidance region estimation mechanism 20 and the boundary estimation mechanism 30 connected in series to the initial guidance region estimation mechanism 20. As a result, the second image 4 is an image representing the probability that each pixel is an abnormal region, and is a result of performing semantic segmentation on the abnormal region based on the first image 3.
[0033] Here, semantic segmentation means performing labeling on each pixel in an image. In the present embodiment, the probability of the existence of an abnormal region is attached to each pixel as a label or a pixel value. Specifically, each pixel value in the second image 4 has the probability of the existence of an abnormal region in at least one of the channels.
[0034] (Initial guidance region estimation mechanism 20) FIG. 3 is a network diagram showing the processing of the generation unit 10. The initial guidance region estimation mechanism 20 includes a vision transformer 21 that inputs the first image 3 and outputs a first feature amount 22b, and a CNN 23 that inputs the first image 3 and outputs a second feature amount 24b. In the generation unit 10, the vision transformer 21 and the CNN 23 are connected in parallel.
[0035] The vision transformer 21 divides the first image 3 into a plurality of non-overlapping square patches 21a. The first encoder 22 in the vision transformer 21 reads the plurality of patches 21a and determines the first feature amount 22b by regarding each as a token 22a in the transformer.
[0036] Here, when determining the first feature quantity 22b, the first encoder 22 may generate a plurality of types of feature quantities with different numbers of channels at various resolutions by performing downsampling or upsampling on the patch 21a. Note that in the first encoder 22, a feature quantity considering the global arrangement of each patch 21a is calculated by similarity calculation for the patch 21a, and is used as the first feature quantity 22b.
[0037] The CNN 23 convolves the first image 3 by the second encoder 24 to obtain a second feature quantity 24b. At this time, convolution is performed so that the resolution of the second feature quantity 24b is the same as that of the first feature quantity 22b.
[0038] Note that both the first feature quantity 22b and the second feature quantity 24b are calculated to be feature quantities of four types of resolutions. Note that the types of resolutions are not limited to four types.
[0039] In order to fuse the first feature quantity 22b calculated by the vision transformer 21 and the second feature quantity 24b calculated by the CNN 23 so that they can be processed simultaneously, first feature fusion units 25a to fourth feature fusion units 25d are provided. Each of the first feature fusion units 25a to fourth feature fusion units 25d fuses feature quantities of the same resolution. Note that at this time, they may be fused by multiplying by a predetermined coefficient, or the predetermined parameters may be changed by learning.
[0040] The feature quantities that have passed through the feature fusion units 25a to 25d are integrated by the intermediate decoder 26. Then, they are normalized by the sigmoid function 27 to obtain a third feature quantity 27a. At this stage, an intermediate result before inputting to the boundary estimation mechanism 30 is generated.
[0041] The initial guidance area estimation mechanism 20 has a function of estimating approximately where in the image the abnormal area is located and its approximate shape. When estimating the approximate position and shape, by using the global feature amount by the vision transformer 21, each value in the third feature amount 27a can be emphasized or suppressed according to the arrangement of each part in the image.
[0042] (Boundary Estimation Mechanism 30) The boundary estimation mechanism 30 generates the second image 4 with reference to the first feature amount 22b and the second feature amount 24b. For this purpose, the boundary estimation mechanism 30 repeats the same process three times. Each repetition is composed of a feature combiner 31, a reverse attention 32, an adder 33, and a sigmoid function 34. To distinguish each repetition, 1st to 3rd are added, with 1st adding a, 2nd adding b, and 3rd adding c respectively. Note that the number of repetitions is not limited to three and is one less than the number of types of the first feature amount 22b and the second feature amount 24b.
[0043] The first feature combiner 31a inputs the output of the second feature fuser 25b and changes the resolution using a feature pyramid. As a result, the feature amount and resolution of the previous stage can be unified for processing.
[0044] Reverse attention 32a is performed to obtain a feature amount that emphasizes the boundary portion of the target object by taking the element-by-element product of the feature amount created by the first feature combiner 31a and the feature amount of the previous stage (here, the third feature amount 27a).
[0045] After that, the first adder 33a that adds the feature amount of the previous stage (here, the third feature amount 27a) to the emphasized feature amount is applied.
[0046] By applying the first sigmoid function 34a to the added feature amount, a new feature amount 35a is obtained.
[0047] For the new feature quantity 35a, the second feature combiner 31b, the second reverse attention 32b, the second adder 33b, and the second sigmoid function 34b are applied to obtain a new feature quantity 35b. Note that for each process, the new feature quantity 35a is used as the feature quantity of the previous stage. Also, the second feature combiner 31b inputs the output of the third feature fuser 25c.
[0048] For the new feature quantity 35b, the third feature combiner 31c, the third reverse attention 32c, the third adder 33c, and the third sigmoid function 34c are applied to obtain a new feature quantity 35c. Note that for each process, the new feature quantity 35b is used as the feature quantity of the previous stage. Also, the third feature combiner 31c inputs the output of the fourth feature fuser 25d.
[0049] For the new feature quantity 35c, the sigmoid function 36 is applied to generate the normalized second image 4. Here, by using the sigmoid function, since each pixel value falls within the range of 0 to 1, the second image 4 becomes an image with the existence probability of the abnormal region as the pixel value. In the boundary estimation mechanism 30, with each additional stage, the resolution of the feature quantity is upsampled, and it can be enlarged to the same size as the resolution of the original image (the first image 3).
[0050] Note that the boundary estimation mechanism 30 has a function of clearly separating the boundary between the abnormal region and the normal region regarding the approximate position of the abnormal region estimated by the initial guidance region estimation mechanism 20. The boundary estimation mechanism 30 upsamples the feature quantity over multiple stages and gradually drops it as a detailed feature quantity. That is, due to the boundary estimation mechanism 30 outputting the shape of the clear abnormal region, it can be identified by the identification unit 40.
[0051] (Pre-training by MAE) Here, for the first encoder 22 in the vision transformer 21, pre-training is performed in advance by the pre-training unit 50. Fig. 4 is a network diagram showing the pre-training of the vision transformer 21.
[0052] Among the plurality of patches 21a created based on the first image 3, a masked patch 21b obtained by masking some of them and an unmasked patch 21a are combined and referred to as a third image 5. Note that the third image 5 is preferably an image including a plurality of non-masked regions (unmasked patches 21a) that are not adjacent to each other. In pre-training, among the third image 5, the unmasked patch 21a is converted into a feature amount 22a by the first encoder 22, and the feature amount 22a corresponding to the masked patch 21b is returned to a patch 22d by the first decoder 22c, and the first image 3 is reconstructed by rearranging the patches 22d in the original array. That is, in pre-training, the first image 3 without defects is learned from the third image 5 in which the first image 3 has been damaged.
[0053] Through pre-training, the first encoder 22 can calculate a feature amount capable of reconstructing the first image 3. Pre-training the first encoder 22 in this way is called MAE (Masked Auto Encoder).
[0054] By using MAE, even when there are few first images 3, images can be increased by changing the arrangement of the masked patches 21b, and the first encoder 22 can be trained. In addition, since the image can be reconstructed for an image with masked defects (third image 5), the global relationship in an image without abnormal regions can be learned. As a result, even for the first image 3 including an abnormal region, it becomes possible to infer where the abnormal region exists, and the abnormal region can be more accurately identified.
[0055] As described above, the first encoder 22 in the vision transformer 21 is learned using the third image 5 including a plurality of non-masked regions that are not adjacent to each other.
[0056] (Identification unit 40) The identification unit 40 inputs the second image 4 and identifies the presence or absence of an abnormal region. As the identification unit 40, the architecture based on ResNet (Residual Neural Networks) shown in Non-Patent Document 3 was used.
[0057] Since the identification unit 40 has a network structure using residual blocks, it can adopt a deep CNN network structure, enabling highly accurate identification.
[0058] Regarding the presence or absence of an abnormal region, one-hot representation or the probability regarding the presence or absence of an abnormal region is output.
[0059] Here, the identification unit 40 identifies the second image 4 only through learning without having parameters such as artificially determined thresholds. Therefore, since there is no intervention of artificial judgment, there is an advantage of having less risk of making incorrect parameter decisions.
[0060] (Multi-task learning) The learning unit 60 causes the generation unit 10 and the identification unit 40 to be learned using multi-task learning (MTL: Multi Task Learning). In multi-task learning, even for models with mutually conflicting tendencies in a plurality of tasks, each model can be learned without degrading the performance. That is, the performance of the generation model 11 in the generation unit 10 and the identification model 41 in the identification unit 40 are simultaneously learned without degrading their performance.
[0061] Here, first, the reason for performing multi-task learning simultaneously will be explained. If single-task learning (STL: Single Task Learning) is performed only by the identification unit 40, depending on the data, when proceeding with the learning to identify the presence or absence of an abnormal region, it may converge to output that there is an abnormal region in any image. Therefore, it is difficult to identify the first image 3 only with the identification unit 40.
[0062] FIG. 5 is a diagram for explaining the process according to the comparative example, and shows the process when single-task learning is performed on the generation unit 10, the weights are fixed, and then single-task learning is performed on the discrimination unit 40. In the case of FIG. 5, although the abnormal region can be identified with the performance as described later, the accuracy is insufficient.
[0063] Therefore, in single-task learning, there is a trade-off relationship depending on the states of the mutual models, and there may be a conflict in the gradients of the loss functions of each model. Therefore, in the present embodiment, multi-task learning is performed on the generation unit 10 and the discrimination unit 40 so as not to damage the mutual models.
[0064] In the present embodiment, multi-task learning is performed. In this multi-task learning, a gradient matrix G is formed with the gradients of the first loss 12 in the generation model 11 and the gradients of the second loss 42 in the discrimination model 41 as column elements, respectively. Also, a coefficient vector α is derived from the following equation. Note that “(t)” in the gradient matrix G and the coefficient vector α represents an index for iteration.
[0065]
Equation
[0066]
Equation
[0067] By using the loss function, it is possible to improve the performance of the generation model 11 using the data of abnormal subjects while not degrading the performance of the discrimination unit 40. Conversely, it is possible to improve the performance of the discrimination unit 40 so as not to degrade the performance of the generation model 11 using the data of normal subjects. That is, optimization becomes possible by using the loss function.
[0068] (Comparative Example 1) Hereinafter, a comparative example for showing the superiority of the generation unit 10 and the discrimination unit 40 in the present embodiment will be described. There are two types of comparative examples.
[0069] As Comparative Example 1, the method described in Non-Patent Document 5 is cited. FIG. 6 is a network diagram of the method shown in Non-Patent Document 5. In Comparative Example 1, the point that the feature amount in the first image 3 is derived by the initial guidance area estimation mechanism 20a and the second image 4a with a clear boundary is generated from the feature amount by the boundary estimation mechanism 30a is the same as in the present embodiment.
[0070] However, the initial guidance area estimation mechanism 20a employs a CNN-based encoder-decoder structure. Therefore, it extracts local feature information and cannot consider the global feature information between pixels at distant positions. As a result, information on the position of a site having a feature amount close to the abnormal area in the image is missing, and it is considered that detection cannot be performed with sufficient accuracy.
[0071] In addition, as a process in the boundary estimation mechanism 30a, a process for clarifying the boundary of the abnormal area is also performed, but the details are different from those of the boundary estimation mechanism 30.
[0072] In Comparative Example 1, in the image output through the boundary estimation mechanism 30a, the presence or absence of the abnormal area is identified by the size of the area of the abnormal area. That is, it can be said that the matter finally determining the identification result is an artificially determined threshold value. Therefore, since there is no guarantee that the threshold value can be set appropriately, it is considered to have low reliability. On the other hand, in the present embodiment, since there is no intervention of artificial judgment, there is an advantage that the risk of making an incorrect threshold / parameter determination is small.
[0073] FIG. 7 shows the results of generating the second image 4a from the first image 3 in Embodiment 1 and Comparative Example 1. In addition, for comparison with the present embodiment, the second image 4 and the correct answer image 6 judged by the doctor are also shown.
[0074] As shown in FIG. 7, it can be seen that the abnormal area that could not be identified in Comparative Example 1 can be identified in the present embodiment.
[0075] (Comparative Example 2) As Comparative Example 2, the method described in Non-Patent Document 6 is cited. FIG. 8 is a network diagram in the method shown in Non-Patent Document 6. In the method of Non-Patent Document 6, in order to extract the feature amount in the first image 3, a network 21c using a transformer and a network 23c using a CNN are connected in parallel. Although this network structure is similar to that connected in parallel in the initial guidance region estimation mechanism 20 in the present embodiment, in Comparative Example 2, it does not have a boundary estimation mechanism 30 and an identification unit 40.
[0076] FIG. 9 shows the results of generating the second image 4b from the first image 3 in Embodiment 1 and Comparative Example 2. In addition, for comparison with the present embodiment, the second image 4 and the correct answer image 6 judged by the doctor are also shown together.
[0077] As shown in FIG. 9, it can be seen that the abnormal region that was over-detected in Comparative Example 2 can be normally identified in the present embodiment.
[0078] (Effect of the proposed method according to Embodiment 1) The effect of multi-task learning is shown by comparing the case of performing multi-task learning in FIG. 2 with the case of not performing multi-task learning in FIG. 5.
[0079] Table 1 shows the test results when multi-task learning is used in Comparative Example 1, Comparative Example 2, and the present embodiment. Note that the accuracy, F-value, and AUC (Area Under the Curve) were evaluated as the test results in Table 1.
[0080]
Table 1
[0081] As shown in Table 1, in the method of this embodiment, higher accuracy is achieved in all evaluation indicators than in Comparative Example 1 and Comparative Example 2. This is considered to be because in this embodiment, in the initial guidance area estimation mechanism 20, by using both the vision transformer 21 and the CNN 23, global learning and local learning can be achieved simultaneously, and the third feature amount 27a having the feature amounts of both can be generated.
[0082] (Effect of multi-task learning) Next, the effect of multi-task learning is shown.
[0083] Table 2 shows the test results when multi-task learning was not used in Comparative Example 1, Comparative Example 2, and this embodiment. Note that accuracy, F value, and AUC were evaluated as the test results in Table 2.
[0084]
Table 2
[0085] (Effect of MAE) Table 3 shows the results of comparing mDice and mIoU when MAE was used and when it was not used in Comparative Example 1 and this embodiment.
[0086]
Table 3
[0087] As shown in Table 3, it can be seen that in both Comparative Example 1 and this embodiment, using MAE results in higher accuracy than not using MAE.
[0088] (Effect according to the type of teacher data used in MAE) In this embodiment, in MAE, learning was performed using the first images 3 of non - healthy and healthy individuals as teacher images. Here, a comparative example in which learning was performed using only the first images 3 of non - healthy individuals as teacher images was compared with the learning results in this embodiment in which learning was performed using the first images 3 including healthy individuals in addition to non - healthy individuals as teacher images.
[0089]
Table 4
[0090] (Parentheses) Therefore, it can be seen that by using multi - task learning and MAE, the abnormal region can be identified with high accuracy. In particular, regarding multi - task learning, by connecting the generation unit 10 in parallel with the vision transformer 21 and the CNN 23, a feature amount having both global feature amounts and local feature amounts can be created, and the accuracy can be improved.
[0091] 〔Embodiment 2〕 Other embodiments of the present invention will be described below. For convenience of explanation, members having the same functions as the members described in the above embodiment are denoted by the same reference numerals, and the description thereof will not be repeated.
[0092] (Identification device 2) FIG. 10 is a block diagram showing the configuration of the main part of the identification device 2 according to the present embodiment. The identification device 2 includes a control unit 100a. The identification device 2 may include a storage unit 200 and an input / output unit 300. The control unit 100a includes a generation unit 10, an identification unit 40, and an acquisition unit 70. That is, the identification device 2 according to the present embodiment is different from the learning device 1 according to Embodiment 1 in that the control unit 100a does not include a pre-learning unit 50 and a learning unit 60 with respect to the control unit 100.
[0093] The identification device 2 stores the generation model 11 and the identification model 41 learned by the learning device 1 in the storage unit 200. Therefore, the identification device 2 can identify an abnormal region from the first image 3 based on the generation model 11 and the identification model 41.
[0094] 〔Modification〕 In Embodiments 1 and 2, the identification of calcification is targeted as the abnormal region, but it is not limited thereto. For example, examples of the abnormal region include tumors, polyps, and hemangiomas.
[0095] 〔Summary〕 The identification device according to Aspect 1 of the present invention includes an acquisition unit that acquires a first image, a generation unit that generates a second image, which is an image after segmentation, by performing semantic segmentation on the first image, and an identification unit that identifies the presence or absence of an abnormal region in the second image, and uses multi-task learning referring to the loss function in the generation model of the generation unit and the loss function in the identification model of the identification unit for the learning of the generation unit and the identification unit.
[0096] According to the above configuration, the presence or absence of an abnormal region can be identified from the second image generated using the first image.
[0097] The identification device according to aspect 2 of the present invention is, in the above aspect 1, wherein the generation unit includes an initial guidance area estimation mechanism and a boundary estimation mechanism connected in series to the initial guidance area estimation mechanism, and the initial guidance area estimation mechanism includes a vision transformer that inputs the first image and outputs a first feature amount, and a convolutional neural network that inputs the first image and outputs a second feature amount, and the boundary estimation mechanism may be configured to generate the second image with reference to the first feature amount and the second feature amount.
[0098] According to the above configuration, in the initial guidance area estimation mechanism in the generation unit, by using both a vision transformer and a convolutional neural network, the boundary estimation mechanism can generate a highly accurate second image.
[0099] The identification device according to aspect 3 of the present invention is, in the above aspect 2, wherein the vision transformer may be configured to be learned using a third image including a plurality of non-mask areas that are not adjacent to each other.
[0100] According to the above configuration, in the vision transformer, by using a pre-learned one using a third image including a plurality of non-mask areas that are not adjacent to each other, even when using a small number of teacher images for generating the third image, a learning effect with sufficient generalization ability can be obtained.
[0101] The identification device according to aspect 4 of the present invention is, in the above aspect 2 or 3, wherein in the generation unit, the vision transformer and the convolutional neural network are connected in parallel, and the generation unit and the identification unit are connected in series.
[0102] According to the above configuration, the vision transformer generates a global feature amount, the convolutional neural network generates a local feature amount, and the generation unit uses the integrated feature amount as the second image, enabling the identification unit to perform highly accurate identification.
[0103] In the identification device according to aspect 5 of the present invention, in any of aspects 1 to 4 above, the identification unit may be configured as a residual neural network architecture.
[0104] According to the above configuration, by using a residual neural network, a deep CNN network structure can be adopted in the identification unit, enabling highly accurate identification.
[0105] In the identification device according to aspect 6 of the present invention, in any of aspects 1 to 5 above, the abnormal region may be configured as a calcified region.
[0106] According to the above configuration, a calcified region can be identified as an abnormal region.
[0107] In the identification device according to aspect 7 of the present invention, in any of aspects 1 to 6 above, a learning unit that updates the parameters in the generation unit and the identification unit by referring to the gradients of the loss functions of the generation unit and the identification unit may be further provided.
[0108] According to the above configuration, the learning unit can perform multi-task learning considering the gradients of the loss functions of the generation unit and the identification unit.
[0109] The learning device according to aspect 8 of the present invention includes an acquisition unit that acquires a first image, a generation unit that generates a second image, which is an image after segmentation, by performing semantic segmentation on the first image, an identification unit that identifies the presence or absence of an abnormal region in the second image, and a learning unit that uses multi-task learning with reference to the loss function in the generation model of the generation unit and the loss function in the identification model of the identification unit for learning the generation unit and the identification unit.
[0110] According to the above configuration, a model that can identify the presence or absence of an abnormal region from the second image generated using the first image can be learned.
[0111] The learning device according to aspect 9 of the present invention may be configured such that the first image includes an image including the abnormal region and an image not including the abnormal region.
[0112] According to the above configuration, by performing learning using an image having an abnormal region and an image not having an abnormal region, global learning can be performed even with a small number of first images.
[0113] The identification method according to aspect 10 of the present invention includes an acquisition step of acquiring a first image, a generation step of generating a second image, which is an image after segmentation, by performing semantic segmentation on the first image, and an identification step of identifying the presence or absence of an abnormal region in the second image, and uses multi-task learning with reference to a loss function in the generation model in the generation step and a loss function in the identification model in the identification step for learning in the generation step and the identification step.
[0114] According to the above configuration, the presence or absence of an abnormal region can be identified from the second image generated using the first image.
[0115] The learning method according to aspect 11 of the present invention includes an acquisition step of acquiring a first image, a generation step of generating a second image, which is an image after segmentation, by performing semantic segmentation on the first image, an identification step of identifying the presence or absence of an abnormal region in the second image, and a learning step of using multi-task learning with reference to a loss function in the generation model in the generation step and a loss function in the identification model in the identification step for learning in the generation step and the identification step.
[0116] According to the above configuration, it is possible to learn a model that can identify the presence or absence of an abnormal region from the second image generated using the first image.
[0117] The identification program according to aspect 12 of the present invention is an identification program for causing a computer to function as the identification device in the above aspect 1, and may also be configured to cause a computer to function as the acquisition unit, the generation unit, and the identification unit.
[0118] The learning program according to aspect 13 of the present invention is a learning program for causing a computer to function as the learning device in the above aspect 8, and may also be configured to cause a computer to function as the acquisition unit, the generation unit, the identification unit, and the learning unit.
[0119] 〔Example of implementation by software〕 The functions of the learning device 1 or the identification device 2 (hereinafter referred to as "device") are programs for causing a computer to function as the device, and can be realized by programs for causing a computer to function as each control block of the device (especially each part included in the control unit 100 or the control unit 100a).
[0120] In this case, the above device includes a computer having at least one control device (for example, a processor) and at least one storage device (for example, a memory) as hardware for executing the above program. By executing the above program with this control device and storage device, each function described in the above embodiments is realized.
[0121] The above program may be recorded on one or more computer-readable recording media, not temporarily. This recording medium may or may not be provided in the above device. In the latter case, the above program may be supplied to the above device via any wired or wireless transmission medium.
[0122] In addition, part or all of the functions of each of the above control blocks can also be realized by a logic circuit. For example, an integrated circuit in which a logic circuit functioning as each of the above control blocks is formed is also included in the scope of the present invention. In addition to this, for example, it is also possible to realize the functions of each of the above control blocks by a quantum computer.
[0123] 〔Supplementary Notes〕 The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope indicated in the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
Explanation of Reference Numerals
[0124] 1 Learning device 2 Identification device 3 First image 4 Second image 5 Third image 10 Generation unit 11 Generation model 12 First loss 20 Initial guidance area estimation mechanism 21 Vision transformer 22b First feature amount 23 CNN (Convolutional Neural Network) 24b Second feature amount 30 Boundary estimation mechanism 40 Identification unit 41 Identification model 42 Second loss 50 Pre-learning unit 60 Learning unit 70 Acquisition unit 100 Control unit 200 Storage unit 300 Input / output unit
Claims
1. An acquisition unit that acquires a first image; A generation unit that performs semantic segmentation on the first image to generate a second image that is a segmented image; an identification unit that identifies the presence or absence of an abnormal area in the second image, A classification device that utilizes multi-task learning that refers to a loss function in a generative model of the generation unit and a loss function in a discriminative model of the classification unit for training the generation unit and the classification unit.
2. the generation unit includes an initial guidance area estimation mechanism and a boundary estimation mechanism connected in series to the initial guidance area estimation mechanism, the initial guidance area estimation mechanism includes a vision transformer that inputs the first image and outputs a first feature amount, and a convolutional neural network that inputs the first image and outputs a second feature amount; The identification device according to claim 1 , wherein the boundary estimation mechanism generates the second image by referring to the first feature amount and the second feature amount.
3. The apparatus of claim 2 , wherein the vision transformer is trained using a third image that includes a plurality of non-adjacent unmasked regions.
4. In the generation unit, the vision transformer and the convolutional neural network are connected in parallel, The identification device according to claim 2 , wherein the generation unit and the identification unit are connected in series.
5. The device according to claim 1 , wherein the classifier is a residual neural network architecture.
6. The identification device according to claim 1 , wherein the abnormal region is a calcified region.
7. The classification device according to claim 1 , further comprising a learning unit that updates parameters in the generation unit and the classification unit by referring to gradients of the loss functions of the generation unit and the classification unit, respectively.
8. An acquisition unit that acquires a first image; A generation unit that performs semantic segmentation on the first image to generate a second image that is a segmented image; An identification unit that identifies the presence or absence of an abnormal area in the second image; A learning device comprising: a learning unit that uses multi-task learning with reference to a loss function in a generative model of the generation unit and a loss function in a discriminative model of the discrimination unit for training the generation unit and the discrimination unit.
9. The learning device according to claim 8 , wherein the first images include an image including the abnormal region and an image not including the abnormal region.
10. acquiring a first image; A generating step of generating a second image, which is a segmented image, by performing semantic segmentation on the first image; and a discrimination step of discriminating the presence or absence of an abnormal area in the second image, A classification method in which multi-task learning that refers to a loss function in a generative model in the generation step and a loss function in a discriminative model in the classification step is utilized for learning in the generation step and the classification step.
11. acquiring a first image; A generating step of generating a second image, which is a segmented image, by performing semantic segmentation on the first image; An identification step of identifying the presence or absence of an abnormal area in the second image; A learning method comprising: a learning step in which multi-task learning referring to a loss function in a generative model of the generation step and a loss function in a discriminative model of the discrimination step is utilized for learning the generation step and the discrimination step.
12. An identification program for causing a computer to function as the identification device according to claim 1 , the identification program causing a computer to function as the acquisition unit, the generation unit, and the identification unit.
13. A learning program for causing a computer to function as the learning device according to claim 8 , the learning program causing a computer to function as the acquisition unit, the generation unit, the identification unit, and the learning unit.
Citation Information
Patent Citations
Intravascular ultrasound imaging and calcium detection methods
JP2022549877A
Automated tumor detection based on image processing
JP2023517058A
Cited By
Deep learning model generation device and deep learning model generation method
JP2026020262A
Method for product inspection using artificial neural network model based on unsupervised and supervised learning
KR103001161B1