Training data production system, training data production method, and recording medium
By performing image correspondence and assigning training information based on the similarity of the photographed objects in a medical image generation device, the problem of low similarity in time series images is solved, and the efficiency of assigning image label information and the recognition accuracy of machine learning are improved.
Patent Information
- Application Number
- CN202080098221.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-09
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2040-03-09
AI Technical Summary
In the prior art, medical image generation devices have a small number of images with low or high similarity in time series images, resulting in an inability to effectively assign label information, which affects the recognition accuracy of machine learning.
A plurality of medical images are acquired by the image acquisition unit, the image correspondence unit corresponds them based on the similarity of the photographed objects to form a corresponding image group, and the assigned object image is displayed by the image output unit, the input acceptance unit accepts the training information, and the training information assignment unit automatically assigns the training information based on the representative training information.
It improves the efficiency of assigning image label information, reduces the amount of user operations, increases the quantity and quality of training images, and improves the recognition accuracy of machine learning.
Smart Images

Figure CN115243601B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a training data production system, a training data production method, a recording medium, and the like. Background Art
[0002] As image recognition technology, machine learning such as deep learning is widely used. In machine learning, a large number of training images suitable for learning need to be input, and one of the methods for preparing such a large number of training images is data augmentation. Data augmentation is a method of increasing the number of training images by processing the original images that were actually taken to generate new images. However, from the perspective of improving the recognition accuracy of machine learning, it is preferable to prepare a large number of original images and assign training labels to the original images. Patent document 1 discloses a medical image generation device that generates images for machine learning of an ultrasonic diagnostic device. In Patent document 1, the medical image generation device sets a medical image displayed when the user performs an image save operation or a freeze operation as an image of interest, and mechanically extracts multiple images similar to the image of interest from a time series image. The user assigns label information to the image of interest. The medical image generation device assigns label information to multiple images similar to the image of interest based on the label information attached to the image of interest.
[0003] Prior art literature
[0004] Patent Literature
[0005] Patent Document 1: Japanese Patent Application Publication No. 2019-118694 Summary of the Invention
[0006] Problems to be solved by the invention
[0007] In Patent Document 1, the image that becomes the focus image when the user performs an image save or freeze operation. Therefore, an image with low similarity to the time series images, or an image with a small number of highly similar images within the time series images, may become the focus image. In this case, there is a problem in that the medical image generation device cannot use the label information assigned to the focus image to assign label information to the time series images, or the number of images that can be labeled by the medical image generation device is small. For example, if an image captured by temporarily changing the camera angle significantly in an endoscope becomes the focus image, it may be impossible to assign label information based on similarity to images captured at other angles.
[0008] Means for solving problems
[0009] One embodiment of the present disclosure relates to a training data production system, which includes: an image acquisition unit that acquires multiple medical images; an image correspondence unit that performs correspondence between the medical images included in the multiple medical images based on the similarity of the captured objects, and produces a corresponding image group obtained by forming a group of medical images that can correspond to each other; an image output unit that outputs an image that will become an object of assignment representing training information, that is, an assignment object image, to a display unit based on the corresponding image group; an input acceptance unit that accepts input of the representative training information; and a training information assignment unit that assigns training information to the medical images included in the corresponding image group based on the representative training information input for the assignment object image.
[0010] Other aspects of the present disclosure relate to a training data production method, which includes the following steps: obtaining multiple medical images; performing correspondence between the medical images included in the multiple medical images based on the similarity of the captured objects, and producing a corresponding image group obtained by forming a group of medical images that can correspond to each other; based on the corresponding image group, outputting an image that will become an assignment object representing training information, namely, an assignment object image, to a display unit; accepting input of the representative training information; and assigning training information to the medical images included in the corresponding image group based on the representative training information input for the assignment object image.
[0011] Yet another embodiment of the present disclosure relates to a training data production program that causes a computer to perform the following processing: obtaining a plurality of medical images; performing correspondence between the medical images included in the plurality of medical images based on the similarity of the captured objects, and producing a corresponding image group obtained by forming a group of medical images that can correspond to each other; based on the corresponding image group, outputting an image that will become an object of assignment representing training information, namely, an assignment object image, to a display unit; accepting input of the representative training information; and assigning training information to the medical images included in the corresponding image group based on the representative training information input for the assignment object image. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is the first configuration example of the training data creation system.
[0013] Figure 2 This is a flowchart of the processing in the first structural example.
[0014] Figure 3 It is a diagram for explaining the operation of the image corresponding unit.
[0015] Figure 4 This is a diagram illustrating a method for determining a corresponding image group.
[0016] Figure 5It is a diagram for explaining the operations of the image output unit, the input acceptance unit, and the training information provision unit.
[0017] Figure 6 This is a flowchart of the processing in the first modified example of the first structural example.
[0018] Figure 7 This is an example of an image displayed on the display unit in the second modification of the first structural example.
[0019] Figure 8 This is a configuration example of a training data creation system in a third modified example of the first configuration example.
[0020] Figure 9 This is a flowchart of the processing in the third modification example of the first structural example.
[0021] Figure 10 This is a configuration example of a training data creation system in a fourth modified example of the first configuration example.
[0022] Figure 11 It is a diagram for explaining a first operation example of the input adjustment unit.
[0023] Figure 12 It is a diagram for explaining a second operation example of the input adjustment unit.
[0024] Figure 13 This is a flowchart of the processing in the second structural example.
[0025] Figure 14 It is a diagram for explaining the operation of the image corresponding unit.
[0026] Figure 15 It is a diagram for explaining the operations of the image output unit, the input acceptance unit, and the training information provision unit.
[0027] Figure 16 It is a diagram illustrating a second modified example of the second configuration example.
[0028] Figure 17 This is a flowchart of the processing in the third modification example of the second configuration example.
[0029] Figure 18 It is a diagram illustrating a third modified example of the second configuration example. DETAILED DESCRIPTION
[0030] The present embodiment is described below. In addition, the present embodiment described below does not unduly limit the contents described in the claims. In addition, all the structures described in this embodiment are not necessarily essential structural elements of the present disclosure.
[0031] 1. First Configuration Example
[0032] Figure 1This is a first configuration example of the training data production system 10. The training data production system 10 includes a processing unit 100, a storage unit 200, a display unit 300, and an operation unit 400. Figure 1 , the endoscope 600 and the learning device 500 are also shown in the figure. However, when the training data creation system 10 creates training data, the endoscope 600 and the learning device 500 do not need to be connected to the training data creation system 10.
[0033] The processing unit 100 includes an image acquisition unit 110, an image correspondence unit 120, an image output unit 130, an input acceptance unit 140, and a training information assignment unit 150. The image acquisition unit 110 acquires a plurality of medical images. The image correspondence unit 120 performs correspondence between the medical images included in the plurality of medical images based on the similarity of the photographed objects, and creates a corresponding image group. The corresponding image group is a group composed of medical images that can correspond to each other in the above correspondence. Based on the corresponding image group, the image output unit 130 outputs an image that will become an assignment object representing training information, that is, an assignment object image, to the display unit 300. The input acceptance unit 140 accepts input representing training information from the operation unit 400. The training information assignment unit 150 assigns training information to the medical images included in the corresponding image group based on the representative training information input to the assignment object image.
[0034] In this way, the image to be assigned, which is displayed on the display unit 300 in order for the user to assign training information, is displayed based on a corresponding image group that is associated based on the similarity of the photographed object. That is, since the similarity is determined before being presented to the user, the image to be assigned is an image that has similarity with each medical image in the corresponding image group. Thus, representative training information is added to medical images that correspond more to images with high similarity, so that the training information assigning unit 150 can automatically assign training information to more medical images using one representative training information. As a result, the number of images to be assigned that the user should input as representative training information can be reduced, and the user can create training images with less working hours.
[0035] The following, Figure 1 The details of the training data production system 10 are described below.
[0036] The training data creation system 10 is, for example, an information processing device such as a PC (Personal Computer). Alternatively, the training data creation system 10 may be a system in which a terminal device and an information processing device are connected via a network. For example, the terminal device includes a display unit 300, an operating unit 400, and a storage unit 200, and the information processing device includes a processing unit 100. Alternatively, the training data creation system 10 may be a cloud system in which multiple information processing devices are connected via a network.
[0037] The storage unit 200 stores a plurality of medical images captured by the endoscope 600. Medical images are images of the body captured by a medical endoscope. A medical endoscope is a video endoscope such as a digestive tract endoscope, or a rigid endoscope used for surgery. The plurality of medical images are time-series images. For example, the endoscope 600 captures a dynamic image of the body, and each frame of the dynamic image corresponds to a respective medical image. The storage unit 200 is a storage device such as a semiconductor memory or a hard disk drive. The semiconductor memory is, for example, a volatile memory such as a RAM, or a non-volatile memory such as an EEPROM.
[0038] The processing unit 100 is a processor. The processor may also be an integrated circuit device such as a CPU, a microcomputer, a DSP, an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array). The processing unit 100 may include one or more processors. In addition, the processing unit 100 as a processor may be, for example, a processing circuit or processing device composed of one or more circuit components, or a circuit device in which one or more circuit components are mounted on a substrate.
[0039] The operation of the processing unit 100 can also be realized by software processing. That is, a program that describes the operation of all or part of the image acquisition unit 110, image matching unit 120, image output unit 130, input receiving unit 140, and training information imparting unit 150 included in the processing unit 100 is stored in the storage unit 200. The processor executes the program stored in the storage unit 200, thereby realizing the operation of the processing unit 100. In addition, the program can also be further described in Figure 8 The program of the detection unit 160 described later in the description can be stored in an information storage medium that is a computer-readable medium. The information storage medium can be implemented, for example, by an optical disc, a memory card, an HDD, or a semiconductor memory. A computer is a device that includes an input device, a processing unit, a storage unit, and an output unit.
[0040] Figure 2 1 is a flowchart of the processing in the first structural example. In step S1, the image acquisition unit 110 acquires multiple medical images. Specifically, the image acquisition unit 110 is an access control unit that reads data from the storage unit 200, acquires multiple medical images from the storage unit 200, and outputs the multiple medical images to the image corresponding unit 120.
[0041] In step S2, the image matching unit 120 extracts feature points from multiple medical images and matches each feature point between the images. In step S3, the image matching unit 120 selects a medical image with the most matching feature points, and the image output unit 130 outputs the selected medical image as the assigned image to the display unit 300. The display unit 300 displays the assigned image. The display unit 300 is, for example, a liquid crystal display device or an EL (electroluminescence) display device.
[0042] In step S4, the input accepting unit 140 accepts input of training information for the displayed assignment target image. Specifically, the user uses the operating unit 400 to assign training information to the assignment target image displayed on the display unit 300, and the assigned training information is input to the input accepting unit 140. The input accepting unit 140 is, for example, a communication interface between the operating unit 400 and the processing unit 100. The operating unit 400 is a device for the user to perform operational input or input information into the training data creation system 10, and may be, for example, a pointing device, a keyboard, or a touch panel.
[0043] In step S5, the training information assigning unit 150 assigns training information to each medical image associated by the image association unit 120 based on the training information assigned to the assignment target image, and stores the training information in the storage unit 200. The medical image to which the training information is assigned is called a training image.
[0044] Training images stored in the storage unit 200 are input to the learning device 500. The learning device 500 includes a storage unit 520 and a processor 510. The training images input from the storage unit 200 are stored in the storage unit 520 as training data 521. A machine-learned learning model 522 is stored in the storage unit 520. The processor 510 performs machine learning on the learning model 522 using the training data 521. The learning model 522 obtained through machine learning is transmitted as a learned model to the endoscope for use in image recognition within the endoscope.
[0045] Figure 3 1 is a diagram for explaining the operation of the image matching unit 120. Figure 3 In , IA1 to IAm are a plurality of medical images acquired by the image acquisition unit 110. m is an integer greater than or equal to 3. Here, m=5, that is, IAm=IA5.
[0046] exist Figure 3In the figure, no feature point is detected in IA1, the same feature point FPa is detected from IA2 to IA5, the same feature point FPb is detected from IA2 to IA4, and the same feature point FPc is detected from IA4 and IA5. Feature points are points that represent the characteristics of an image, and are points that are characterized by feature quantities detected from the image. Various feature quantities can be conceived as feature quantities, such as edges, gradients, or statistics of the image. The image correspondence unit 120 extracts feature points using methods such as SIFT (Scale Invariant Feature Transform), SURF (Speeded-Up Robust Features), and HOG (Histograms of Oriented Gradients).
[0047] The image correspondence unit 120 uses the feature points FPa to FPc to match a group of images with high similarity of the captured objects as a corresponding image group GIM. "Similarity of the captured objects" refers to the degree of similarity of the objects projected in the medical images. The higher the similarity, the higher the possibility that the objects are the same. When feature points are used, the greater the number of identical feature points, the higher the similarity of the captured objects can be judged. For example, the image correspondence unit 120 matches medical images with a number of identical feature points above a threshold. Figure 3 In the example, the image correspondence unit 120 corresponds medical images with two or more identical feature points. IA4 has two or more identical feature points relative to IA2, IA3, and IAm, respectively. Therefore, IA2 to IA4 and IAm are set as the corresponding image group GIM.
[0048] Figure 4 This figure illustrates a method for determining a corresponding image group. Here, it is assumed that multiple medical images IX1 to IX5 are obtained. The image correspondence unit 120 determines the similarity of each medical image IX1 to IX5 with other medical images. A zero indicates a high similarity, and an x indicates a low similarity. IX1, IX2, IX3, IX4, and IX5 are determined to have high similarity with two, three, three, two, and two medical images, respectively.
[0049] The image correspondence unit 120 selects a medical image with more similarities as an assigned target image, and sets the medical image with high similarity to the assigned target image as a corresponding image group. Figure 4In the example, IX2 is selected as the assigned image, and IX1 to IX3 and IX5 are set as the corresponding image group. Alternatively, IX3 is selected as the assigned image, and IX1 to IX4 are set as the corresponding image group. The image assignment unit 120 may also select, for example, the image with the smaller number among IX2 and IX3, that is, the image that is earlier in the time series, as the assigned image. Alternatively, if IX4 is included in other corresponding image groups, the image assignment unit 120 may select IX1 to IX3 and IX5 as the corresponding image group, or select the corresponding image group taking into account other corresponding image groups.
[0050] Figure 5 1 is a diagram for explaining the operations of the image output unit 130, the input receiving unit 140, and the training information providing unit 150. Figure 5 In the above table, IB1 to IBn are equivalent to Figure 3 IA2~IAm. Figure 4 According to the method described in , IB3 is selected from the medical images IB1 to IBn as the corresponding image group GIM as the assigned target image RIM. n is an integer greater than or equal to 2.
[0051] The image output unit 130 causes the display unit 300 to display the assignment target image RIM selected by the image matching unit 120. The user assigns representative training information ATIN to the assignment target image RIM via the operation unit 400. The representative training information ATIN is training information assigned to the assignment target image RIM. Figure 5 The figure shows an example in which the training information is the contour information surrounding the detection object in machine learning, but the training information can also be text, rectangle, contour or area. Text is information used as training data for AI that performs classification processing. Text is information that represents the content or status of the photographed object, such as the type of lesion or the malignancy of the lesion. Rectangle is information used as training data for AI that performs detection. A rectangle is a rectangle that is circumscribed to the photographed object such as a lesion, indicating the existence and location of the lesion. Contours or regions are information used as training data for AI that performs segmentation. Contours are information that represent the boundary between the photographed object such as a lesion and the area outside the photographed object such as normal mucosa. Regions are information that represent the area occupied by the photographed object such as a lesion, and include contour information.
[0052] The training information imparting unit 150 imparts training information AT1 to ATn to the medical images IB1 to IBn included in the corresponding image group GIM based on the representative training information ATIN accepted by the input accepting unit 140. Figure 5In the example, representative training information ATIN is directly assigned to IB3, which is the assigned target image RIM, as training information AT3. The training information assigning unit 150 geometrically transforms the representative training information ATIN based on the feature points for medical images other than IB3, thereby assigning training information. Specifically, between two images with two or more identical feature points, the position and correspondence of these feature points can reveal the amount of parallel translation, rotation, and scaling between the images. The training information assigning unit 150 transforms the representative training information ATIN into training information AT1, AT2, AT4, and ATn through geometric transformation using these amounts of parallel translation, rotation, and scaling.
[0053] According to the present embodiment, the image output unit 130 outputs the medical image selected from the corresponding image group GIM to the display unit 300 as the assigned target image RIM.
[0054] In this way, any medical image among the medical images included in the corresponding image group GIM is displayed as an assigned object image RIM on the display unit 300. The assigned object image RIM is associated with other medical images in the corresponding image group GIM, so by using this correspondence, training information can be assigned to each medical image based on the representative training information ATIN.
[0055] In addition, in the present embodiment, in the correspondence based on the similarity of the photographed objects, the number of medical images corresponding to each medical image is set to the number of corresponding images. Figure 4 IX1 to IX5 have corresponding image numbers of 2, 3, 3, 2, and 2, respectively. The image output unit 130 outputs the medical image selected based on the corresponding image number to the display unit 300 as the assignment target image RIM.
[0056] Specifically, the image correspondence unit 120 selects an assigned target image RIM from a plurality of medical images based on the number of corresponding images, and sets the medical image corresponding to the assigned target image RIM among the plurality of medical images as the corresponding image group GIM. More specifically, the image correspondence unit 120 selects the medical image with the largest number of corresponding images from the plurality of medical images for which the number of corresponding images has been calculated as the assigned target image RIM.
[0057] In this way, the number of images corresponding to one assigned image RIM increases, thereby reducing the number of assigned images RIM that the user should assign to represent the training information ATIN. This can reduce the burden on the user in creating training images. In addition, in the above-mentioned structural example, the assigned image RIM is selected by the image matching unit 120, but the assigned image RIM can also be selected by the image output unit 130. For example, Figure 4As shown in IX2 and IX3, when there are multiple candidates with the same number of highly similar images, the image output unit 130 may select any candidate from the multiple candidates as the assigned target image. For example, the image output unit 130 may select a candidate image that contains a large number of feature points as the assigned target image.
[0058] In this embodiment, the input accepting unit 140 accepts representative contour information indicating the contour of a specific region assigned to the target image RIM as representative training information ATIN. The training information assigning unit 150 assigns contour information to the medical images included in the corresponding image group GIM as training information AT1 to ATn.
[0059] A specific region is a region targeted for detection by AI, which is obtained through machine learning based on training images generated by the training data creation system 10. According to this embodiment, contour information representing the contours of the specific region is assigned to the medical image as training information AT1 to ATn. Therefore, by performing machine learning based on this training information, segmented specific region detection is possible.
[0060] In addition, in this embodiment, the image matching unit 120 extracts feature points FPa to FPc representing the features of each of the multiple medical images IA1 to IAm, and matches the medical images IB1 to IBn having the same feature points in the multiple medical images IA1 to IAm, thereby creating a corresponding image group GIM. The training information assigning unit 150 assigns contour information representing the contour of a specific area to the medical images IB1 to IBn included in the corresponding image group GIM using the same feature points FPa to FPc as training information AT1 to ATn.
[0061] In this way, the contour information of each medical image can be assigned using the information of the feature points matched when creating the corresponding image group GIM. As described above, between images having the same feature points, the representative contour information can be transformed into the contour information of each medical image through geometric transformation.
[0062] 2. Modification of the First Structural Example
[0063] In the first variant, the similarity of the photographed objects is determined based on the image similarity. Figure 1 same. Figure 6 is a flowchart of the processing in the first variant. Steps S11, S15 and Figure 2 Steps S1 and S5 are the same, so their description is omitted.
[0064] In step S12, the image correspondence unit 120 calculates the image similarity between the medical images included in the plurality of medical images. Image similarity is an indicator that indicates the degree to which two images are similar, that is, an indicator that indicates the similarity of the images themselves. When the images are similar, it can be determined that the similarity of the photographed objects reflected in the image is also high. Examples of image similarity include SSD (Sum of Squared Difference), SAD (Sum of Absolute Difference), NCC (Normalized Cross Correlation), ZNCC (Zero-mean Normalized Cross Correlation), and the like.
[0065] In step S13, the image matching unit 120 selects images with high image similarity from a plurality of medical images as assigned target images. For example, the image matching unit 120 determines the similarity by comparing the image similarity with a threshold value. Figure 4 The image output unit 130 outputs the selected image to be assigned to the display unit 300 .
[0066] In step S14, the input accepting unit 140 accepts input of training information for the displayed assignment target image. When image similarity is used, the training information may be, for example, text, but is not limited thereto.
[0067] In the second modification, the corresponding status is displayed. Figure 7 This is an example of an image displayed on the display unit 300 in the second modification.
[0068] The image output unit 130 outputs the correspondence information 303 between the assigned target image RIM and the medical images 302 other than the assigned target image RIM among the plurality of medical images to the display unit 300. Figure 7 In FIG, 301 is an arrow indicating the flow of time. That is, medical images 302 are displayed in a time series along arrow 301. Correspondence information 303 is a line connecting the assigned target image RIM and the medical image 302. Images connected by a line are images that are determined to have a high similarity to the imaged objects and a correspondence is established. Images not connected by a line are images that are determined to have a low similarity to the imaged objects and a correspondence is not established.
[0069] In this way, the user can know the correspondence between the assigned target image RIM and each medical image. For example, the user can determine whether an appropriate medical image has been selected as the assigned target image RIM.
[0070] The input accepting unit 140 accepts an input designating any medical image in the medical images 302 other than the assigned target image RIM as a new assigned target image. The input accepting unit 140 changes the assigned target image based on the accepted input. That is, when the user observes the corresponding information 303 displayed on the display unit 300 and determines that there is a more appropriate assigned target image than the current assigned target image RIM, the user selects the medical image. The input accepting unit 140 uses the medical image selected by the user as the new assigned target image.
[0071] In this way, the user can reselect a medical image that is determined to be suitable as a medical image for assigning representative training information as an assignment target image, and the user can assign representative training information to the assignment target image. In this way, a training image that can achieve higher-precision machine learning can be produced.
[0072] Alternatively, after the user selects a new target image, the image association unit 120 may reset the corresponding image group, and the image output unit 130 may display the corresponding information on the display unit 300. Specifically, the image association unit 120 may set medical images that are highly similar to the newly selected target image to the new corresponding image group, and output the corresponding information to the image output unit 130.
[0073] In the third modified example, display adjustment of a specific area is performed. Figure 8 This is a configuration example of the training data creation system 10 in the third modification. Figure 8 In the embodiment, the processing unit 100 further includes a detection unit 160. In addition, the image output unit 130 includes a display adjustment unit 131. In addition, the description of the components that have already been described will be omitted as appropriate.
[0074] Figure 9 is a flowchart of the processing in the third modification example. Steps S21, S26, and S27 are Figure 2 Steps S1, S4, and S5 are the same, so their description is omitted.
[0075] In step S22, the detection unit 160 detects a specific region from each of the multiple medical images. The specific region is the region to which training information is assigned. In step S23, the image matching unit 120 extracts feature points from each of the multiple medical images and matches the feature points between the images. At this time, the image matching unit 120 may extract feature points based on the specific region detected by the detection unit 160. For example, the image matching unit 120 may extract feature points that characterize the specific region.
[0076] In step S24, the display adjustment unit 131 scales the assigned target image based on the status of the corresponding image and displays it on the display unit 300. For example, if the size of the specific area projected in the corresponding image group is large or small, the display adjustment unit 131 scales the assigned target image to adjust the size of the specific area. In step S25, the display adjustment unit 131 adjusts the detected specific area so that it is easily visible and displays the assigned target image on the display unit 300. For example, the display adjustment unit 131 moves the specific area parallel to the center of the screen, or rotates or performs a projection transformation so that the specific area is easily visible. A projection transformation is a transformation that changes the inclination of the surface where the specific area is located, that is, a transformation that changes the angle between the surface where the specific area is located and the camera's line of sight. In addition, both steps S25 and S24 can be implemented, or either one can be implemented.
[0077] According to the present embodiment, the display adjustment unit 131 applies a geometric transformation to the assigned image and then outputs the result to the display unit 300. The geometric transformation may be, for example, parallel translation, rotation, scaling, projection transformation, or a combination thereof.
[0078] In this way, geometric transformation can be performed to make the specific area assigned representative training information by the user easier to see. This allows the user to be presented with the specific area in an easily visible manner, reducing the user's workload or allowing the user to accurately assign representative training information.
[0079] In addition, according to the present embodiment, the detection unit 160 detects information of a specific area from a plurality of medical images through image processing. The display adjustment unit 131 performs display adjustment of the specific area by performing geometric transformation based on the information of the detected specific area. Image processing is an image recognition process that identifies the position or area of a specific area from an image. As image recognition processing, various technologies can be used, as an example, an image recognition technology using machine learning, or an image recognition technology using feature extraction. The information of the detected specific area is the position, size, outline, area, or a combination thereof of the specific area. By using this information, parallel movement, rotation, or scaling, etc. can be performed in a manner that appropriately displays the specific area.
[0080] In this way, the specific region to which the training information should be assigned is detected in advance by the detection unit 160 , and the display adjustment unit 131 can geometrically transform the assignment target image so as to appropriately display the specific region.
[0081] In the fourth modification, the input representative training information is adjusted. Figure 10 This is a configuration example of the training data creation system 10 in the fourth modification. Figure 10In FIG, the input accepting unit 140 includes an input adjusting unit 141. In addition, description of components already described will be omitted as appropriate.
[0082] Figure 11 This figure illustrates a first operational example of the input adjustment unit 141. The input adjustment unit 141 detects estimated outline information ATES for a specific region from a target image through image processing. The input acceptance unit 140 accepts representative outline information ATINb representing the outline of the specific region in the target image as representative training information. If the estimated outline information ATES has an accuracy level exceeding a predetermined value, the input adjustment unit 141 corrects the representative outline information ATINb based on the estimated outline information ATES.
[0083] The input adjustment unit 141 performs AI-based image recognition processing and outputs estimated contour information ATES and its estimated accuracy. For example, the accuracy is output for each position along the contour. Based on the estimated contour information ATES, the input adjustment unit 141 corrects the representative contour information ATINb for portions where the accuracy exceeds a specified value. For example, the input adjustment unit 141 replaces the representative contour information ATINb with the estimated contour information ATES for portions where the accuracy exceeds a specified value. The input adjustment unit 141 outputs the corrected representative contour information as final representative training information ATF to the training information imparting unit 150.
[0084] According to this embodiment, if the estimated contour information based on image processing is accurate, the representative contour information input by the user is corrected based on this estimated contour information. For example, even if the contour information input by the user is incorrect, correction to the correct contour information is possible. In this case, since correction is only performed when the estimated contour information is accurate, highly accurate correction can be achieved.
[0085] Figure 12 This figure illustrates a second example of the operation of the input adjustment unit 141. The input adjustment unit 141 detects estimated outline information ATES for a specific area from a given target image through image processing. The input acceptance unit 140 accepts representative outline information ATINb representing the outline of the specific area in the given target image as representative training information. If the estimated outline information ATES has an accuracy greater than a specified value, the input adjustment unit 141 outputs representative training information based on the estimated outline information ATES. If the estimated outline information ATES has an accuracy less than the specified value, the input adjustment unit 141 outputs representative training information based on the representative outline information ATINb.
[0086] The input adjustment unit 141 outputs the estimated contour information ATES and its estimated accuracy by performing image recognition processing based on AI processing. For example, the accuracy of the entire contour is output. When the accuracy is greater than a specified value, the input adjustment unit 141 outputs the estimated contour information ATES as the final representative training information ATF to the training information giving unit 150. In addition, when the accuracy is less than a specified value, the input adjustment unit 141 outputs the representative contour information ATINb as the final representative training information ATF to the training information giving unit 150. Figure 12 The middle diagram shows a case where the estimated outline information ATES is selected.
[0087] In the fifth modification, feature points are extracted again based on the input representative training information. Figure 1 same.
[0088] The image matching unit 120 extracts feature points representing image characteristics from each of the multiple medical images and matches the medical images with the same feature points, thereby creating a corresponding image group. The input receiving unit 140 receives representative contour information representing the contour of a specific area assigned to the target image as representative training information. When the contour of the representative contour information deviates from the feature point by more than a specified value, the image matching unit 120 re-extracts and matches the feature points based on the position of the contour of the representative contour information.
[0089] Specifically, the image correspondence unit 120 calculates the distance between the contour representing the contour information and the feature point assigned to the object image. The distance is, for example, the distance between the point in the contour closest to the feature point and the feature point. When there are multiple feature points in the assigned object image, the distance is calculated for each feature point, and the shortest distance or average distance is calculated as the final distance. When the calculated distance is greater than the specified value, the image correspondence unit 120 re-extracts the feature points of each medical image. At this time, the image correspondence unit 120 extracts the feature points that exist near the contour representing the contour information input by the user, and uses the feature points to reconstruct the corresponding image group.
[0090] According to this embodiment, feature points are extracted again based on the contour representing the contour information input by the user. This re-extracts feature points associated with the specific region identified by the user. By reconstructing a corresponding image set using these feature points, a corresponding image set appropriately associated with the feature points of the specific region can be obtained. This allows the training information assignment unit 150 to accurately assign contour information to each medical image based on the representative contour information.
[0091] 3. Second Configuration Example
[0092] The second configuration example of the training data production system 10 is described. Figure 1 same. Figure 13 This is a flowchart of the processing in the second structural example. Steps S31, S32 and Figure 2 Steps S1 and S2 are the same, so their description is omitted.
[0093] In step S33, the image matching unit 120 constructs a 3D model by synthesizing multiple medical images based on the matching results in step S32. In step S34, the image output unit 130 adjusts the pose of the 3D model so that the corresponding points in the corresponding medical images appear more frequently, and outputs a 2D image of the pose as the assigned target image to the display unit 300. In step S35, the input receiving unit 140 receives input of representative training information for the assigned target image. Examples of training information include, but are not limited to, rectangles, contours, and regions.
[0094] In step S36, the input receiving unit 140 automatically adjusts the input representative training information according to the nearby contours in the 3D model, or automatically adjusts the input representative training information according to the depth of the 3D model. In step S37, the training information assigning unit 150 assigns training information to the medical image or region corresponding to the 3D model based on the representative training information assigned to the 3D model.
[0095] Figure 14 This is a diagram for explaining the operation of the image matching unit 120. Figure 14 Omitted, but with Figure 3 Similarly, feature points are extracted from the medical images IB1 to IBn, and medical images with high similarity to the captured objects are matched as a corresponding image group GIM.
[0096] The image correspondence unit 120 uses the correspondence of the medical images IB1 to IBn included in the corresponding image group GIM to synthesize the 3D model MDL using the medical images IB1 to IBn. Specifically, the image correspondence unit 120 estimates the camera position and line of sight direction at which each medical image was captured based on the correspondence between the feature points in each medical image and the feature points between the medical images, and estimates the 3D model MDL based on the estimation result. As a synthesis method of the 3D model MDL, for example, SfM (Structure from Motion) can be used.
[0097] Figure 15This figure illustrates the operations of the image output unit 130, the input receiving unit 140, and the training information assigning unit 150. The image output unit 130 generates assigned target images RIM1 and RIM2 based on the 3D model MDL. These assigned target images RIM1 and RIM2 are two-dimensional images of the 3D model MDL as viewed from lines of sight VL1 and VL2. For example, line of sight VL2 is viewed from the opposite direction to line of sight VL1. Specifically, if the appearance of the 3D model MDL in the assigned target image RIM1 is called the front surface, the appearance of the 3D model MDL in the assigned target image RIM2 is called the back surface.
[0098] The user assigns representative training information ATIN1 and ATIN2 to the assignment target images RIM1 and RIM2, respectively. Figure 15 The figure illustrates a case where the training information is contour information. Based on the representative training information ATIN1 and ATIN2, the input accepting unit 140 assigns the representative training information to the 3D model MDL. Specifically, the input accepting unit 140 generates the contour of a specific region in the 3D model MDL based on the contours of the representative training information ATIN1 and ATIN2. For example, if the specific region is a convex portion, the boundary between the convex portion and the surface other than the convex portion forms the contour of the specific region in the 3D model MDL.
[0099] The training information assigning unit 150 assigns training information AT1 to ATn to the medical images IB1 to IBn of the corresponding image group GIM based on the representative training information assigned to the 3D model MDL. When the 3D model MDL is generated, the camera position and line of sight direction corresponding to each medical image are estimated. In addition, if a 3D model is generated, the distance between each point on the model and the camera is also estimated. The training information assigning unit 150 uses the camera position and line of sight direction to transform the representative training information in the 3D model MDL into training information AT1 to ATn in the medical images IB1 to IBn.
[0100] According to the present embodiment, the image output unit 130 outputs an image generated using the corresponding image group GIM to the display unit 300 as an assignment target image.
[0101] The corresponding image group GIM is associated based on the similarity of the photographed subject, so the image generated using the corresponding image group GIM is an image related to the photographed subject. By displaying such an image as the assigned target image on the display unit 300, the photographed subject is appropriately presented to the user.
[0102] In this embodiment, the image output unit 130 generates a 3D model MDL of the captured object based on the corresponding image group GIM, and outputs images RIM1 and RIM2 based on the 3D model MDL to the display unit 300 as assignment target images.
[0103] In this way, the 3D model MDL of the imaged object is presented to the user. By using the 3D model MDL, it is expected that training information can be appropriately assigned to lesions, etc., which are important to the 3D structure. Furthermore, regardless of the imaging method of the imaged object in each medical image, the 3D model MDL observed from any camera viewpoint can be presented to the user, allowing for the display of an image suitable for assigning representative training information.
[0104] Specifically, the image output unit 130 displays the assigned object image when the 3D model MDL is observed with a camera line of sight that shows more corresponding points. The corresponding points are the same feature points between the medical images used to construct the 3D model MDL. Since the correspondence of the feature points is used when constructing the 3D model MDL, the positions of the feature points on the 3D model MDL can be known. Therefore, it is possible to know which camera line of sight shows more corresponding points when observing the 3D model MDL, and display the assigned object image observed with that camera line of sight. Figure 15 , two assignment target images RIM1 and RIM2 are shown in the figure. However, for example, an image observed with a camera sight line having the largest number of corresponding points is set as either RIM1 or RIM2.
[0105] Displaying an image showing a large number of corresponding points means displaying a portion with high correspondence accuracy when constructing the 3D model MDL based on the corresponding image group GIM. Such an assigned target image is believed to correspond with each medical image in the corresponding image group GIM with high accuracy. Therefore, assigning training information to each medical image based on the representative training information assigned to the assigned target image can generate highly accurate training information.
[0106] In addition, in this embodiment, the image output unit 130 outputs a plurality of two-dimensional images RIM1 and RIM2 having different viewing directions relative to the three-dimensional model MDL as a plurality of assigned target images to the display unit 300. The input receiving unit 140 receives representative contour information representing the contour of a specific region of the assigned target image as representative training information ATIN1 and ATIN2 for each of the plurality of two-dimensional images. The training information assigning unit 150 assigns contour information as training information AT1 to ATn to each of the medical images IB1 to IBn included in the corresponding image group GIM.
[0107] When observing the 3D model MDL from one viewing direction, there may be parts that are not visible due to the back surface of a convex part, etc. According to this embodiment, by displaying multiple 2D images with different viewing directions, 2D images obtained by observing the 3D model MDL from various directions are displayed, and representative contour information is assigned to each of the 2D images. In this way, contour information is assigned to the 3D model MDL without omission. In addition, Figure 15In the example, two two-dimensional images are displayed as the assignment target images. However, the number of two-dimensional images to be displayed is not limited to two, and a plurality of two-dimensional images may be displayed.
[0108] In addition, the above description uses the case where "the image generated using the corresponding image group" is an image based on the 3D model MDL as an example, but is not limited to this. As in the modified example described later, it can also be an expanded view of the digestive tract, etc. generated using the corresponding image group GIM.
[0109] 4. Modification of the Second Structural Example
[0110] In the first modification, input adjustment of representative profile information is performed. The hardware structure of the training data production system 10 is the same as Figure 10 same.
[0111] The input accepting unit 140 accepts representative contour information indicating contours of specific regions assigned to the target images RIM1 and RIM2 as representative training information ATIN1 and ATIN2. The input adjusting unit 141 adjusts the input representative contour information based on the contour and depth of the three-dimensional model MDL.
[0112] Depth refers to the depth of the 3D model MDL when observed in a 2D image displayed as an assigned object image, and is equivalent to the line of sight of the camera when generating the 2D image. For example, when a contour is added to a convex portion of the 3D model MDL, when the contour assigned by the user is observed in the depth direction, there may be multiple surfaces of the 3D model MDL in the depth direction. In such a case, the input adjustment unit 141 adjusts the representative contour information so that the contour is added to the foremost surface. In addition, when the representative contour information assigned by the user is away from the contour of the 3D model MDL in the assigned object image, the representative contour information is adjusted in a manner consistent with the contour of the 3D model MDL close to the representative contour information.
[0113] According to this embodiment, representative contour information can be appropriately assigned to the 3D model MDL based on the representative contour information assigned to the target image, which is a 2D image. Specifically, by taking into account the depth of the 3D model MDL, representative contour information can be assigned to the contour desired by the user. Furthermore, by taking into account the contour of the 3D model MDL, even if the contour information entered by the user is incorrect, it can be corrected to the correct contour information.
[0114] In the second modification, input errors of representative contour information are detected. The hardware structure of the training data production system 10 is similar to Figure 1 same. Figure 16 : is a diagram illustrating a second modified example. Figure 16 Although one image to be assigned is shown in FIG, a plurality of images to be assigned may be displayed as described above.
[0115] The image output unit 130 generates a two-dimensional image based on the three-dimensional model MDL and outputs the two-dimensional image as the assigned target image RIM. The input receiving unit 140 receives representative contour information ATINc representing the contour of a specific region assigned to the target image RIM as representative training information. If the three-dimensional model MDL is not present at the depth of the contour of the representative contour information ATINc in the viewing direction when generating the two-dimensional image using the three-dimensional model MDL, the input receiving unit 140 outputs an error message indicating a change in the viewing direction.
[0116] exist Figure 16 , the 3D model MDL displayed in the assigned object image RIM does not overlap with the representative contour information ATINc. That is, in the line of sight direction, there is no 3D model MDL in the depth of the contour representing the contour information ATINc. In this case, the input acceptance unit 140 outputs an error message. For example, the error message is input to the image output unit 130, and the image output unit 130 causes the display unit 300 to display a display content urging the user to change the line of sight based on the error message. Alternatively, the error message is input to the image corresponding unit 120, and the image corresponding unit 120 performs the correspondence of the medical images again based on the error message, thereby regenerating the corresponding image group and the 3D model, and the image output unit 130 displays the assigned object image again.
[0117] According to this embodiment, when inappropriate representative outline information ATINc is input to the 3D model MDL, an error message is output, thereby preventing the use of the inappropriate representative outline information. For example, consider a situation where an inappropriate 3D model is not generated, or a situation where the camera's line of sight is inappropriate when presenting the 3D model to the user. According to this embodiment, even in these situations, the use of inappropriate representative outline information can be prevented.
[0118] In the third modification, a development view is generated using the corresponding image group, and an assignment target image based on the development view is displayed. That is, in the third modification, the "image generated using the corresponding image group" is the development view. Figure 17 This is a flowchart of the processing in the third variant. Steps S41 and S42 are Figure 2 Steps S1 and S2 are the same, so their description is omitted. Figure 18 It is a diagram illustrating a third modified example.
[0119] In step S43, the image correspondence unit 120 synthesizes the medical images IB1 to IBn according to the correspondence of the medical images IB1 to IBn in the corresponding image group GIM, thereby producing the expanded view TKZ. Specifically, based on the correspondence between the feature points of the medical images, the image correspondence unit 120 fits the medical images IB1 to IBn in a manner that the positions of the corresponding feature points are consistent, thereby producing the expanded view TKZ. Figure 18, although the bonding of IB1 to IB3 among IB1 to IBn is shown in the figure, IB4 to IBn are also bonded in the same manner.
[0120] The image output unit 130 outputs the assignment target image based on the development view TKZ to the display unit 300. The image output unit 130 may display the entire development view TKZ on the display unit 300, or may display a portion of the development view TKZ showing a specific area on the display unit 300.
[0121] In step S44, the input accepting unit 140 accepts the representative training information ATIN input by the user. Figure 18 The case where the training information is outline information is shown in FIG. The input accepting unit 140 automatically adjusts the input representative training information according to the nearby outline in the developed view.
[0122] In step S46, the training information assigning unit 150 assigns training information to the medical images or regions corresponding to the expanded view based on the representative training information assigned to the expanded view. Since the expanded view TKZ was created by affixing medical images IB1 to Ibn in step S43, the positions of the medical images within the expanded view TKZ are known. The training information assigning unit 150 uses the positions of the medical images within the expanded view TKZ to convert the representative training information ATIN into training information for each medical image.
[0123] The present embodiment and its variations have been described above, but the present disclosure is not directly limited to the various embodiments and their variations. During the implementation stage, the constituent elements can be modified and concretized without departing from the main purpose. In addition, the multiple constituent elements disclosed in the above-mentioned various embodiments and variations can be appropriately combined. For example, several constituent elements can be deleted from all the constituent elements described in the various embodiments and variations. Furthermore, the constituent elements described in different embodiments and variations can also be appropriately combined. In this way, various modifications and applications can be made without departing from the main purpose of the present disclosure. In addition, in the specification or the drawings, a term that is recorded at least once together with a different term in a broader sense or with the same meaning can be replaced with the different term in any part of the specification or the drawings.
[0124] Explanation of symbols
[0125] 10 training data production system, 100 processing unit, 110 image acquisition unit, 120 image corresponding unit, 130 image output unit, 131 display adjustment unit, 140 input acceptance unit, 141 input adjustment unit, 150 training information assignment unit, 160 detection unit, 200 storage unit, 300 display unit, 301 arrow, 302 medical image, 303 corresponding information, 400 operation unit, 500 learning device, 510 processor, 520 storage unit, 521 training data, 522 learning model, 600 endoscope, AT1~ATn training information, ATES estimated contour information, ATIN represents training information, ATINb represents contour information, FPa~FPc feature points, GIM corresponding image group, IA1~IAm medical images, IB1~IBn medical images, MDL three-dimensional model, RIM assigned object image, TKZ expansion diagram.
Claims
1. A training data production system, characterized in that: Include: an image acquisition unit that acquires a plurality of medical images; An image correspondence unit that performs correspondence between the medical images included in the plurality of medical images based on similarity of the photographed objects, and creates a corresponding image group formed by groups of medical images that can correspond to each other; an image output unit that outputs an assignment target image, which is an image representing an assignment target of the training information, to a display unit based on the corresponding image group; an input receiving unit that receives representative contour information indicating a contour of a specific region of the target image as the representative training information; as well as A training information imparting unit imparts contour information as training information to the medical images included in the corresponding image group based on the representative training information input for the imparting target image.
2. The training data production system according to claim 1, wherein: The image output unit outputs the medical image selected from the corresponding image group to the display unit as the assigned target image.
3. The training data production system according to claim 2, wherein: When the number of medical images corresponding to each medical image in the correspondence based on the similarity is set as the number of corresponding images, The image output unit outputs the medical image selected according to the corresponding image number to the display unit as the assigned object image.
4. The training data production system according to claim 2, wherein: The image correspondence unit selects the assigned object image from the multiple medical images based on the number of medical images corresponding to each medical image in the correspondence, and uses the medical image corresponding to the assigned object image among the multiple medical images as the corresponding image group.
5. The training data production system according to claim 1, wherein: The image correspondence unit extracts feature points representing features of the image from each of the multiple medical images, and creates the corresponding image group by performing the correspondence on the medical images having the same feature points among the multiple medical images. The training information imparting unit imparts contour information indicating a contour of a specific region to the medical images included in the corresponding image group as the training information using the same feature points.
6. The training data production system according to claim 2, wherein: The image output unit includes a display adjustment unit that performs geometric transformation on the target image and outputs the resultant image to the display unit.
7. The training data production system according to claim 6, characterized in that: The training data production system includes a detection unit that detects information of a specific area from the plurality of medical images through image processing. The display adjustment unit adjusts the display of the specific area by performing the geometric transformation based on the information of the specific area.
8. The training data production system according to claim 2, wherein: The input accepting unit includes an input adjusting unit. The input adjustment unit detects estimated contour information of a specific area from the assignment target image through image processing, and corrects the representative contour information based on the estimated contour information when the estimated contour information has an accuracy level greater than a predetermined value.
9. The training data production system according to claim 2, wherein: The input accepting unit includes an input adjusting unit. The input adjustment unit detects estimated contour information of a specific area from the assigned object image through image processing, and outputs the representative training information based on the estimated contour information when the estimated contour information has an accuracy greater than a specified value; and outputs the representative training information based on the representative contour information when the estimated contour information has an accuracy less than a specified value.
10. The training data production system according to claim 1, wherein: The image correspondence unit extracts feature points representing features of the image from each of the multiple medical images, and creates the corresponding image group by performing the correspondence on the medical images having the same feature points among the multiple medical images. When the contour of the representative contour information is separated from the feature point by more than a predetermined value, the image matching unit extracts and matches the feature points again based on the position of the contour of the representative contour information.
11. The training data production system according to claim 1, wherein: The image output unit outputs correspondence information between the assignment target image and medical images other than the assignment target image among the plurality of medical images to the display unit.
12. The training data production system according to claim 11, wherein: The input accepting unit accepts input for designating any medical image among the multiple medical images other than the assignment target image as the new assignment target image, and changes the assignment target image based on the accepted input.
13. The training data production system according to claim 1, wherein: The image output unit outputs an image generated using the corresponding image group to the display unit as the assignment target image.
14. The training data production system according to claim 13, wherein: The image output unit generates a three-dimensional model of the photographed object based on the corresponding image group, and outputs an image based on the three-dimensional model to the display unit as the assigned target image.
15. The training data production system according to claim 14, wherein: The image output unit outputs a plurality of two-dimensional images having different sight lines with respect to the three-dimensional model to the display unit as the plurality of assigned target images. The input receiving unit receives the representative contour information as the representative training information for each of the plurality of two-dimensional images. The training information imparting unit imparts contour information as the training information to the medical images included in the corresponding image group.
16. The training data production system according to claim 14, wherein: The input accepting unit includes an input adjusting unit. The input adjustment unit adjusts the input representative outline information based on the outline and depth of the three-dimensional model.
17. The training data production system according to claim 14, wherein: The image output unit generates a two-dimensional image using the three-dimensional model and outputs the two-dimensional image as the assigned target image. The input accepting unit outputs error information instructing to change the viewing direction when the three-dimensional model is not present at the depth of the outline representing the outline information in the viewing direction when the two-dimensional image is generated using the three-dimensional model.
18. A method for producing training data, characterized in that: The following steps are involved: Acquire multiple medical images; Performing correspondence between the medical images included in the plurality of medical images based on similarity of the photographed objects, and producing a corresponding image group formed by groups of medical images that can correspond to each other; outputting, based on the corresponding image group, an assignment target image, which is an image to be assigned representing the training information, to a display unit; receiving representative contour information indicating a contour of a specific region of the target image as the representative training information; as well as Based on the representative training information input for the assignment target image, contour information is assigned as training information to the medical images included in the corresponding image group.
19. A computer-readable recording medium storing a program for causing a computer to execute the following processing: Acquire multiple medical images; Performing correspondence between the medical images included in the plurality of medical images based on similarity of the photographed objects, and producing a corresponding image group formed by groups of medical images that can correspond to each other; outputting, based on the corresponding image group, an assignment target image, which is an image to be assigned representing the training information, to a display unit; receiving representative contour information indicating a contour of a specific region of the target image as the representative training information; as well as Based on the representative training information input for the assignment target image, contour information is assigned as training information to the medical images included in the corresponding image group.
Citation Information
Patent Citations
Medical image generation apparatus
JP2019118694A
Presentation method of defect image
JP2013254286A
Apparatus and Method for Displaying Capsule Endoscope Image, and Record Media Storing Program for Carrying out that Method
US20100165088A1