Apparatus and method for generating training data, and apparatus and method for generating training models
The learning data generation apparatus efficiently synthesizes image data by combining regions of interest from multiple images, addressing the time-consuming nature of large data training and reducing image seam reflection, thus accelerating the learning process.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FUJIFILM CORP
- Filing Date
- 2022-10-26
- Publication Date
- 2026-05-07
AI Technical Summary
Training using large amounts of training data requires a considerable amount of time.
A learning data generation apparatus and method that synthesizes image data by combining regions of interest from multiple images, ensuring the regions are separated by a threshold distance from the boundary line to minimize the impact of image transitions during convolutional processing.
Enables efficient learning by reducing the amount of training data required and minimizing the reflection of image seams in the learning process, thereby accelerating the training process.
Smart Images

Figure 0007855010000001 
Figure 0007855010000002 
Figure 0007855010000003
Abstract
Description
Technical Field
[0001] The present invention relates to a learning data generation device and method, and a learning model generation device and method, and particularly relates to a learning data generation device and method for a learning model that performs image recognition, and a learning model generation device and method.
Background Art
[0002] In recent years, learning models for performing image recognition have enabled the generation of models with high recognition accuracy by deep learning (see Non-Patent Document 1, etc.) if there is a large amount of learning data.
[0003] Patent Document 1 describes a technique for increasing learning data by synthesizing an image of a recognition target with an image used as an input image during learning.
[0004] Patent Document 2 describes a technique for increasing the variation of learning data by extracting an image of a specific part from an image of a recognition target, performing image conversion processing on the extracted part image, and synthesizing it with the image of the recognition target.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Non-Patent Documents
[0006]
Non-Patent Document 1
Summary of the Invention
[0007] However, it has been pointed out that training using large amounts of training data requires a considerable amount of time.
[0008] One embodiment of the technology described herein provides a learning data generation apparatus and method that enables efficient learning, as well as a learning model generation apparatus and method. [Means for solving the problem]
[0009] (1) A learning data generation device that generates learning data, comprising a processor, wherein the processor acquires first image data and second image data, each having a region of interest, and when the positional relationship between the region of interest of the first image data and the region of interest of the second image data satisfies predetermined conditions, the device synthesizes the image of the region containing the region of interest of the first image data and the image of the region containing the region of interest of the second image data to generate third image data.
[0010] (2) The learning data generation apparatus of (1), wherein the predetermined conditions include that the region of interest of the first image data is located within a first region in the image, and the region of interest of the second image data is located within a second region different from the first region in the image.
[0011] (3) The learning data generation apparatus of (2), wherein the predetermined conditions include that the region of interest of the first image data is located within the first region at a distance of a threshold or more from the boundary line separating the first region and the second region, and the region of interest of the second image data is located within the second region at a distance of a threshold or more from the boundary line.
[0012] (4) The learning data generation device according to (2) or (3), wherein the predetermined conditions include that a plurality of regions of interest in the first image data are located within the first region at a distance of a threshold or more from the boundary line separating the first region and the second region, and a plurality of regions of interest in the second image data are located within the second region at a distance of a threshold or more from the boundary line.
[0013] (5) A training data generator according to (3) or (4), wherein, when training data is used to train a neural network using convolutional processing, the threshold is set based on the size of the receptive field of the first convolutional layer.
[0014] (6) A training data generation device according to any one of (2) to (5), wherein the processor synthesizes the image of the first region of the first image data with the image of the region of the second image data other than the first region to generate the third image data.
[0015] (7) The processor generates the third image data by overwriting the image of the region of the first image data other than the first region with the image of the region of the second image data other than the first region, as described in (6).
[0016] (8) A learning data generation device according to any one of (1) to (7), wherein the predetermined condition includes the fact that the region of interest of the first image data and the region of interest of the second image data are separated by a threshold or more.
[0017] (9) The learning data generation device of (8), wherein the processor sets a boundary line between the region of interest of the first image data and the region of interest of the second image data to divide the image into multiple regions, and synthesizes the image of the first image data of the region of the first image data that contains the region of interest among the multiple regions of the first image data divided by the boundary line and the image of the second image data of the region of the second image data that contains the region of interest among the multiple regions of the second image data that contains the boundary line to generate the third image data.
[0018] (10) The processor generates a third image data by overwriting the image of the region of interest in the first image data with the image of the region of interest in the second image data. (9)
[0019] (11) A training data generator, one of (8) to (10), wherein when training data is used to train a neural network using convolutional operations, the threshold is set based on the size of the receptive field of the first convolutional layer.
[0020] (12) The processor acquires first correct data indicating the correct answer of the first image data and second correct data indicating the correct answer of the second image data, and generates third correct data indicating the correct answer of the third image data from the first correct data and the second correct data. The learning data generation device according to any one of (1) to (11).
[0021] (13) The processor generates third correct data indicating the correct answer of the third image data from the first correct data and the second correct data according to the conditions when generating the third image data from the first image data and the second image data. The learning data generation device according to (12).
[0022] (14) The first correct data and the second correct data are mask data for the target area. The learning data generation device according to (12) or (13).
[0023] (15) A learning model generation device that generates a learning model, comprising a processor. The processor acquires the third image data generated by any one of the learning data generation devices according to (1) to (14), and learns the learning model using the third image data. The learning model generation device.
[0024] (16) The processor further uses at least one of the first image data and the second image data used for generating the third image data to learn the learning model. The learning model generation device according to (15).
[0025] (17) The processor performs learning using the third image data and learning using at least one of the first image data and the second image data. The learning model generation device according to (16).
[0026] (18) The processor excludes the boundary region of the image synthesis of the third image data and learns the learning model. The learning model generation device according to any one of (15) to (17).
[0027] (19) A method for generating training data, comprising the steps of: acquiring first image data and second image data, each having a region of interest; determining whether the region of interest of the first image data and the region of interest of the second image data are in a specific positional relationship; and, if the positional relationship between the region of interest of the first image data and the region of interest of the second image data satisfies predetermined conditions, synthesizing an image of the region containing the region of interest of the first image data and an image of the region containing the region of interest of the second image data to generate a third image data.
[0028] (20) A method for generating a learning model, comprising the steps of: acquiring first image data and second image data, each having a region of interest; generating third image data by combining an image of the region containing the region of interest of the first image data and an image of the region containing the region of interest of the second image data, when the positional relationship between the region of interest of the first image data and the region of interest of the second image data satisfies a predetermined condition; and training a learning model using the third image data. [Effects of the Invention]
[0029] According to the present invention, learning can be done efficiently. [Brief explanation of the drawing]
[0030] [Figure 1] A diagram showing an example of training data. [Figure 2] Conceptual diagram for generating training data [Figure 3] A diagram showing an example of image segmentation. [Figure 4] A diagram showing an example where the location of the lesion cannot be identified. [Figure 5] Figure showing an example of new image data. [Figure 6] A diagram showing an example of new correct answer data. [Figure 7] Block diagram showing an example of the hardware configuration of a training data generation device. [Figure 8] Block diagram of the main functions of the learning data generation device [Figure 9] A flowchart illustrating an example of the procedure for generating new training data. [Figure 10] This diagram shows an example of combining four image data. [Figure 11] This diagram shows an example of setting a boundary line by dynamically changing it. [Figure 12] A diagram showing an example of the newly generated image data. [Figure 13] Conceptual diagram for determining whether synthesis is possible or not [Figure 14] Block diagram of the main functions of the learning data generation device [Figure 15] A flowchart illustrating an example of the procedure for generating new training data. [Figure 16] A diagram showing another example of boundary line setting. [Figure 17] This diagram shows an example of dynamically switching the boundary settings for each training data set being synthesized. [Figure 18] This figure shows an example of setting boundaries when there are multiple areas of interest. [Figure 19] Block diagram of the main functions of the learning model generation device. [Modes for carrying out the invention]
[0031] Preferred embodiments of the present invention will be described below with reference to the attached drawings.
[0032] [Training Data Generation Device (Training Data Generation Method)] [First Embodiment] This section explains the process using the example of generating a learning model that recognizes lesions from images (endoscopic images) of tubular organs such as the stomach and large intestine. In particular, it explains the process of generating a learning model that recognizes the area occupied by the lesion within the image, that is, a learning model that performs image segmentation (especially semantic segmentation). In this case, examples of learning models that can be used include U-net, FCN (Fully Convolutional Network), SegNet, PSPNet (Pyramid Scene Parsing Network), and Deeplabv3+. These are types of neural networks that use convolutional processing, i.e., convolutional neural networks (CNN or ConvNet).
[0033] Figure 1 shows an example of training data.
[0034] As shown in the figure, the training data consists of pairs of image data and ground truth data.
[0035] The image data is training image data. The training image data consists of image data that includes the object to be recognized. As described above, in this embodiment, a learning model that recognizes lesions is generated from images taken with an endoscope. Therefore, the training image data consists of image data taken with an endoscope and includes image data that includes lesions. In particular, it consists of image data of the target organ for image recognition taken with an endoscope. For example, when recognizing a lesion in the stomach, it consists of image data of the stomach taken with an endoscope.
[0036] Ground truth data is data that shows the correct answers for the training image data. In this embodiment, the training image data consists of image data in which the lesion is distinguished from the rest of the image. Figure 1 shows an example in which ground truth data is constructed using so-called mask images. In this case, the ground truth data is constructed using image data of images with the lesion masked (images with the lesion filled in). Image data of images with the lesion masked is an example of mask data.
[0037] Thus, training data consists of pairs of image data and ground truth data (image pairs). A large number of these image pairs are prepared to build a dataset, and the training model is then trained using this constructed dataset.
[0038] [Overview of training data generation] Figure 2 is a conceptual diagram of the generation of training data.
[0039] As shown in the figure, in this embodiment, two training data sets are combined to generate new training data.
[0040] The newly generated training data will be referred to as "new training data." The image data and ground truth data that make up this new training data will be referred to as "new image data" and "new ground truth data," respectively.
[0041] Furthermore, the two sets of training data used to generate new training data will be referred to as "First Training Data" and "Second Training Data." The image data and ground truth data that make up the First Training Data will be referred to as "First Image Data" and "First Ground Truth Data," respectively. The image data and ground truth data that make up the Second Training Data will be referred to as "Second Image Data" and "Second Ground Truth Data," respectively.
[0042] The new training data is generated as follows:
[0043] First, training data containing the recognition target is acquired within the image data. The acquired training data is designated as the first training data. In this embodiment, the image data is image data taken with an endoscope, and the recognition target is the lesion. The lesion is an example of a region of interest.
[0044] Next, the system determines which region of the image data (first image data) that constitutes the first training data contains the lesion. In this embodiment, the image data is divided into two regions, and the system determines which region contains the lesion.
[0045] Figure 3 shows an example of image segmentation.
[0046] As shown in the figure, in this embodiment, the image is divided into two equal parts vertically. The straight line separating the regions is called the boundary line BL. The region above the boundary line BL is called the upper region UA, and the region below it is called the lower region LA. Figure 3 shows an example where the lesion X is located in the lower region LA. Therefore, in the example in Figure 3, it is determined that the lesion X is located in the lower region LA.
[0047] Figure 4 shows an example where the location of the lesion cannot be identified.
[0048] As shown in the figure, if lesion X spans two regions, it is impossible to determine which region the lesion is located in. Therefore, in this case, it is determined that the location of the lesion cannot be identified.
[0049] Here, the case where lesion X spans two regions is when lesion X lies on the boundary line BL. Therefore, the condition for identifying the location of lesion X is that lesion X does not lie on the boundary line BL.
[0050] Furthermore, in this embodiment, the following is required in order to determine that the lesion X is located in the upper region UA or the lower region LA: namely, the lesion X is required to be separated from the boundary line BL by a threshold Th or more.
[0051] As described above, in this embodiment, new image data is generated by combining two image data (first image data and second image data). As will be described later, the new image data is generated by combining the upper region of the first image data and the lower region of the second image data. Alternatively, it is generated by combining the lower region of the first image data and the upper region of the second image data. In other words, the images of the opposite regions of the first and second image data are joined together by the boundary line BL to generate new image data. In the new image data generated in this way, the image switches at the seam (see Figure 5). Therefore, if a lesion exists near the seam, the part where the image switches may be reflected in the learning process. In other words, there is a risk that an image that does not actually exist may be reflected in the learning process.
[0052] Therefore, in this embodiment, it is required that the lesion X is not located near the boundary line BL, that is, that it is located at a distance of Th or more from the boundary line BL. This requirement is set from the perspective of the impact on learning. Therefore, the threshold Th is set from the perspective of the impact on learning. Accordingly, when the generated training data is used for training a neural network using convolutional processing, it is preferable to set it based on the size of the receptive field. In particular, it is preferable to set it based on the size of the receptive field of the first convolutional layer. For example, as shown in Figure 3, suppose that the size (length × width) of the receptive field RF of the first convolutional layer is m × n. In this embodiment, since the boundary line BL is set horizontally, the threshold Th is at least m Set the value to a value greater than / 2. This prevents the lesion region, including the image transition area, from being convolved in at least the first convolutional layer, thereby suppressing the reflection of the image transition area in the learning process.
[0053] Furthermore, when we say that the lesion X is located more than a threshold Th away from the boundary line BL, it means that the distance between the pixel closest to the boundary line BL among the pixels constituting the lesion X and the boundary line BL is greater than or equal to the threshold Th.
[0054] If the location of the lesion cannot be identified from the image data in the first training data obtained, the next training data is obtained. In other words, the above process is repeated until training data is obtained that allows the location of the lesion to be identified.
[0055] Once the location of lesion X is identified in the first image data, training data to be used for synthesis is acquired. The acquired training data, like the first training data, is training data that includes the target to be recognized within the image data. The image data is image data taken with an endoscope. The acquired training data is designated as the second training data.
[0056] Next, in the image data that constitutes the second training data (second image data), it is determined whether a lesion is located in a specific region within the image. Here, a specific region is a region in the first image data to be synthesized in which no lesion is located. Therefore, the specific region changes depending on the region in the first image data to be synthesized in which the lesion is located. If the lesion is located in the upper region UA in the first image data to be synthesized, the lower region LA becomes the specific region. In this case, the upper region UA is an example of the first region, and the lower region LA is an example of the second region. On the other hand, if the lesion is located in the lower region LA in the first image data to be synthesized, the upper region UA becomes the specific region. In this case, the lower region LA is an example of the first region, and the upper region UA is an example of the second region.
[0057] In the second image data, in order to determine that a lesion is located in a specific region, it is required that the lesion is located in the specific region at a distance of a threshold Th or more from the boundary line BL.
[0058] If the lesion is located in a specific region in the second image data, and the positional relationship between the lesion in the first image data and the lesion in the second image data satisfies predetermined conditions, then synthesis is performed and new image data is generated. Synthesis is performed as follows: That is, the lesion in the first image data is includeNew image data is generated by combining the image of the region and the image of the region containing the lesion in the second image data. Therefore, for example, if the lesion is located in the upper region of the first image data, the image of the upper region of the first image data and the image of the lower region of the second image data are combined to generate new image data. On the other hand, if the lesion is located in the lower region of the first image data, the image of the lower region of the first image data and the image of the upper region of the second image data are combined to generate new image data.
[0059] Figure 5 shows an example of the new image data.
[0060] As shown in the figure, new image data is generated in which the lesion X is included in both the upper region UA and the lower region LA of the image. In this embodiment, the new image data is an example of the third image data.
[0061] The method of synthesis is not particularly limited. For example, a method of synthesis by overwriting can be employed. That is, a method can be employed in which the image of a part of one image data (the area other than the area of interest) is overwritten with the image of the corresponding area (the area of interest) of the other image data, and then the images are synthesized. For example, if the area of interest is located in the upper area of the first image data, an image of the lower area (the area of interest) is extracted from the second image data, and the image of the lower area (the area other than the area of interest) of the first image data is overwritten with the extracted image. Alternatively, an image of the upper area (the area of interest) is extracted from the first image data, and the image of the upper area (the area other than the area of interest) of the second image data is overwritten with the extracted image. In addition, a method can be employed in which images of the area to be synthesized are extracted from each image data and then synthesized. For example, if the area of interest is located in the upper area of the first image data, an image of the upper area is extracted from the first image data, and an image of the lower area is extracted from the second image data. The images extracted from each image data are then joined together to generate new image data.
[0062] The same process is used to synthesize the ground truth data to generate new ground truth data. That is, the first and second ground truth data are synthesized under the same conditions as the new image data to generate new ground truth data. For example, if the new image data is generated by synthesizing the upper region of the first image data and the lower region of the second image data, then the upper region of the first ground truth data and the lower region of the second ground truth data are synthesized to generate new ground truth data. On the other hand, if the new image data is generated by synthesizing the lower region of the first image data and the upper region of the second image data, then the lower region of the first ground truth data and the second ground truth data above The image of the side region is combined with the new ground truth data to generate new ground truth data. In this embodiment, the new ground truth data is an example of third ground truth data.
[0063] Figure 6 shows an example of the new correct answer data. This figure shows the correct answer data for the new image data shown in Figure 5.
[0064] As shown in the figure, corresponding to the new image data (see Figure 5), images (mask images) containing the lesion X in the upper region UA and lower region LA of the image, respectively, are generated as new ground truth data.
[0065] Furthermore, if the lesion is not located in a specific region of the image data in the training data acquired as the second training data, the next training data is acquired. In other words, the above process is repeated until training data in which the lesion is located in a specific region is acquired.
[0066] As described above, in this embodiment, new image data is generated by combining image data containing the lesion (area of interest) in one region (first region) obtained by dividing the image vertically into two equal parts, and image data containing the lesion (area of interest) in the other region (second region). Then, new ground truth data is generated by combining images under the same conditions as the new image data. This makes it possible to generate training data containing two lesions (areas of interest) in a single training data set. Furthermore, this reduces the amount of training data required.
[0067] [Hardware configuration] Figure 7 is a block diagram showing an example of the hardware configuration of a training data generation device.
[0068] The learning data generation device 1 is, for example, composed of a computer and includes a processor 2, main memory 3, auxiliary storage device 4, input device 5 and output device 6, etc. In other words, the learning data generation device 1 of this embodiment functions as a learning data generation device when the processor 2 executes a predetermined program (learning data generation program). The auxiliary storage device 4 stores the program executed by the processor 2 and various data necessary for processing, etc. Learning data necessary for generating new learning data, and the generated new learning data are also stored in this auxiliary storage device 4. The input device 5 is an operation unit that includes a keyboard, mouse, and other devices. study It includes an input interface for taking in the training data necessary for generating data. The output device 6 includes a display and an output interface for outputting the newly generated training data, etc.
[0069] Figure 8 is a block diagram of the main functions of the learning data generation device.
[0070] As shown in the figure, the learning data generation device 1 mainly has functions such as a first learning data acquisition unit 11, a location identification unit 12, a second learning data acquisition unit 13, a synthesis feasibility determination unit 14, a new learning data generation unit 15, and a new learning data recording unit 16. The functions of each unit are realized by the processor 2 executing a predetermined program.
[0071] The first training data acquisition unit 11 acquires training data to be used as the first training data. In this embodiment, the training data to be used as the first training data is acquired from the auxiliary storage device 4. Therefore, it is assumed that training data is already stored in the auxiliary storage device 4. This training data is training data used to generate new training data. Therefore, it is training data that includes a region of interest within the image. This training data is also used as the second training data.
[0072] The position identification unit 12 performs a process to identify the location of the lesion, which is the region of interest, in the image data (first image data) that constitutes the first learning data. In this embodiment, it performs a process to determine whether the lesion is located in the upper region UA or the lower region LA. As described above, in order to determine that the lesion is located in the upper region UA or the lower region LA, it is required that the lesion is located in the upper region UA or the lower region LA at a distance of a threshold Th or more from the boundary line BL.
[0073] The second training data acquisition unit 13 acquires training data to be used as the second training data. As described above, it acquires training data to be used as the second training data from the auxiliary storage device 4.
[0074] The synthesis feasibility determination unit 14 performs a process to determine whether the acquired second training data can be synthesized. Specifically, it determines whether a lesion is located in a specific region in the image data (second image data) that constitutes the second training data. As described above, a specific region is a region in the first image data to be synthesized in which no lesion is located. If the lesion is located in the upper region UA in the first image data to be synthesized, the lower region LA becomes the specific region. On the other hand, if the lesion is located in the lower region LA in the first image data to be synthesized, the upper region UA becomes the specific region. If the synthesis feasibility determination unit 14 determines that the lesion is located in the specific region in the acquired second training data, it determines that synthesis is possible. In order to determine that a lesion is located in the specific region, it is required that the lesion is located in the specific region at a distance of a threshold Th or more from the boundary line BL.
[0075] The new training data generation unit 15 performs the process of generating new training data. Specifically, it generates new training data by combining the first training data with the second training data which has been determined to be synthesizable with the first training data. In this case, if the lesion is located in the upper region UA of the first image data, the image of the upper region UA of the first image data and the image of the lower region LA of the second image data are combined to generate new image data. On the other hand, if the lesion is located in the lower region LA of the first image data, the image of the lower region LA of the first image data and the image of the upper region UA of the second image data are combined to generate new image data. In addition, new ground truth data is generated in accordance with the generation of new image data. That is, new ground truth data is generated under the same conditions as the conditions for generating new image data. Therefore, for example, if the lesion is located in the upper region UA of the first image data, the image of the upper region UA of the first ground truth data and the image of the lower region LA of the second ground truth data are combined to generate new ground truth data. On the other hand, if the lesion is located in the lower region LA of the first image data, the image of the lower region LA of the first ground truth data and the image of the upper region UA of the second ground truth data are combined to form a new correct answer Data is generated.
[0076] The new learning data recording unit 16 performs the process of recording the new learning data generated by the new learning data generation unit 15. As an example, in this embodiment, the generated new learning data is recorded in the auxiliary storage device 4.
[0077] [Generating new training data] Figure 9 is a flowchart showing an example of the procedure for generating new training data.
[0078] First, the first training data is obtained (step S1). Specifically, one of the multiple training data stored in the auxiliary storage device 4 is read to obtain the first training data.
[0079] Next, the location of the lesion in the acquired first training data is identified (step S2). Specifically, it is determined whether the lesion is located in the upper region or the lower region of the image data constituting the first training data (first image data). Then, based on the result of this determination process, it is determined whether or not the location of the lesion has been identified (step S3).
[0080] If the location of the lesion could not be identified in step S2 (if step S3 is No), it is determined whether there is any unprocessed first training data (step S4). That is, it is determined whether there is any training data that has not yet been used as first training data. If there is no unprocessed first training data, the process ends. On the other hand, if there is unprocessed first training data, the process returns to step S1, the unprocessed first training data is obtained, and the processing from step S2 onwards is performed. That is, the first training data to be processed is switched.
[0081] If the location of the lesion can be identified in step S2 (if step S3 is Yes), then the second training data is acquired (step S5). Similar to the first training data, one of the multiple training data stored in the auxiliary storage device 4 is read to acquire the second training data.
[0082] Next, it is determined whether the acquired second training data can be combined (step S6). Specifically, it is determined whether the lesion is located in a specific region in the image data that constitutes the second training data (second image data). As described above, the specific region is determined by the first training data to be combined. If the lesion is located in the upper region of the first image data in the first training data to be combined, the lower region is set as the specific region. On the other hand, if the lesion is located in the lower region of the first image data in the first training data to be combined, the upper region is set as the specific region.
[0083] If it is determined that synthesis is not possible, the system checks whether there is any unprocessed second training data (step S7). That is, it checks whether there is any training data that has not yet been used as second training data. If there is no unprocessed second training data, the process ends. On the other hand, if there is unprocessed second training data, the system returns to step S5, obtains that unprocessed second training data, and determines whether synthesis is possible (step S6). That is, it switches the second training data to be processed.
[0084] On the other hand, if it is determined that the data can be combined, the process of generating new training data is performed (step S8). Specifically, the first image data of the first training data and the second image data of the second training data are combined to generate new image data for the new training data. Also, the first ground truth data of the first training data and the second ground truth data of the second training data are combined to generate new ground truth data for the new training data.
[0085] Here, the new image data is generated by combining the image of the region containing the lesion in the first image data and the image of the region containing the lesion in the second image data. Therefore, for example, if the lesion is contained in the upper region of the first image data, the new image data is generated by combining the image of the upper region of the first image data and the image of the lower region of the second image data. Similarly, if the lesion is contained in the lower region of the first image data, the new image data is generated by combining the image of the lower region of the first image data and the image of the upper region of the second image data. In the same way, the first ground truth data and the second ground truth data are combined to generate new ground truth data. The generated new training data is stored in the auxiliary storage device 4.
[0086] After generating new training data, determine whether there is any unprocessed first training data (Step S9). .vinegar In other words, the process determines whether there is any training data that has not yet been used as the first training data. If there is no unprocessed first training data, the process ends. On the other hand, if there is unprocessed first training data, the process returns to step S1 and starts generating new training data using the unprocessed training data.
[0087] The training data used to generate new training data will be considered processed training data and will not be used to generate new training data in the future. Similarly, training data in which the location of the lesion could not be identified as the first training data will also be considered processed training data. Therefore, training data in which the location of the lesion could not be identified as the first training data will not be used to generate new training data in the future. On the other hand, training data that was determined to be unsuitable for synthesis as the second training data will not be considered processed training data. This is because this training data may be able to be synthesized with other training data as the first training data.
[0088] As described above, the learning data generation device 1 of this embodiment can extract only the region containing the lesion from two learning data sets to generate new learning data. This reduces the amount of learning data and thus the time required for learning. In other words, it enables efficient learning.
[0089] [Differentiation] [When the training data to be synthesized has multiple regions of interest] In the above embodiment, the case in which the number of regions of interest (lesions) included in the first training data and the second training data is one was described, but the application of the present invention is not limited to this. The same can be applied when the training data used as the target for synthesis has multiple regions of interest. In this case, it is preferable that all regions of interest satisfy the synthesis conditions (predetermined conditions). For example, when synthesizing an image by dividing it into two equal parts vertically as in the above embodiment, it is preferable that for the first training data, all regions of interest included in its image data (first image data) are located in the upper or lower region. Similarly, for the second training data, it is preferable that all regions of interest included in its image data (second image data) are located in a specific region. This makes it possible to generate new image data that makes full use of the information of the regions of interest included in the training data.
[0090] Furthermore, in order to determine that all areas of interest included in the first image data are located in the upper or lower region, it is preferable to require the following additional conditions: that all areas of interest included in the first image data are located in the upper or lower region, separated from the boundary line by a threshold or more. Similarly, in order to determine that all areas of interest included in the second image data are located in a specific region, it is preferable to require that all areas of interest included in the second image data are located in the specific region, separated from the boundary line by a threshold or more. This helps to suppress the reflection of image seams in the learning process.
[0091] [Division of area] In the above embodiment, the example described was the case where an image is divided into two regions, upper and lower, and then combined. However, the manner of division is not limited to this. In addition, for example, a method of dividing the image horizontally into two equal parts and combining them can also be employed. Alternatively, a method of dividing the image diagonally into two equal parts and combining them can also be employed.
[0092] Furthermore, although the above embodiment described the case of combining two training data, the number of training data to be combined is not limited to this. It is also possible to combine three or more training data to generate new training data. In this case, the image is divided according to the number of training data to be combined. For example, when combining three training data to generate new training data, the image is divided into three regions. Similarly, when combining four training data to generate new training data, the image is divided into four regions. The manner of division is not particularly limited. For example, when combining three training data, the image is divided into three vertically or horizontally. Or, it is divided into three circumferentially. Also, for example, when combining four training data, the image is divided into four vertically or horizontally. Or, it is divided into four circumferentially. For each divided region, the images of the corresponding regions of each training data are combined to generate new training data. Figure 10 is a diagram showing an example of combining four image data. The same figure shows an example of combining four image data by dividing the image into four equal parts in the circumferential direction. The new image data is generated by placing the image of the first region of the first image data in the first region (upper left region). The image of the second region of the second image data is placed in the second region (upper right region). The image of the third region of the third image data is placed in the third region (lower left region). The image of the fourth region of the fourth image data is placed in the fourth region (lower right region). Here, the image data selected as the first image data is the image data that has the lesion (region of interest) X in the first region (upper left region). The image data selected as the first image data is the image data that has the lesion X in the second region (upper right region). The image data selected as the third image data is the image data that has the lesion X in the third region (lower left region). The image data selected as the fourth image data is the image data that has the lesion X in the fourth region (lower right region).
[0093] [Setting Boundaries] In the above embodiment, the boundary line is fixed and the images of predetermined regions are combined, but the image data of the first training data (first imageThe configuration may dynamically change the position of the boundary line depending on the position of the region of interest contained in the data. In this case, the area in which the image is synthesized changes depending on the position of the region of interest contained in the first image data.
[0094] Figure 11 shows an example of setting the boundary line by dynamically changing it. The figure shows an example of setting the boundary line BL that divides the image into two halves vertically by dynamically changing it.
[0095] First, the location of the lesion (area of interest) X is identified within the first image data. Next, the distance from the upper edge of lesion X to the top edge of the image is calculated. Note that the upper edge of lesion X is equivalent to the uppermost pixel among the pixels that make up lesion X. Similarly, the distance from the lower edge of lesion X to the bottom edge of the image is calculated. Note that the lower edge of lesion X is equivalent to the lowermost pixel among the pixels that make up lesion X. The calculated distances are compared, and the area with the longer distance is selected as the boundary line BL setting area. Figure 11 shows an example where the upper area of lesion X is selected as the boundary line BL setting area. The boundary line BL is set in the selected setting area. In this case, the boundary line BL is set at a distance D from the upper edge of lesion X.
[0096] Here, the distance D is set in the same way as the threshold Th in the above embodiment, from the perspective of its impact on learning. Therefore, when the generated training data is used to train a neural network using convolutional processing, it is set based on the size of the receptive field, in particular the size of the receptive field of the first convolutional layer.
[0097] Thus, the boundary line can also be set for each training data set, depending on the location of the region of interest contained in the image data of the first training data set.
[0098] In the example shown in Figure 11, the second image data to be synthesized is selected from images that include the lesion in the region above the boundary line BL.
[0099] Figure 12 shows an example of the new image data.
[0100] As shown in the figure, the first image data is placed below the defined boundary line BL, and the second image data is placed above it. Image The image data into which the elements are placed is generated as new image data.
[0101] [Boundary Structure] In the above embodiment, the boundary line is composed of a horizontal straight line, but it can also be composed of a diagonal straight line. Furthermore, it can be composed of a curve instead of a straight line. In addition, it can be composed of a straight line that is partially bent (a so-called polyline).
[0102] [Second Embodiment] [overview] In this embodiment, when generating new training data by combining two training data sets, the feasibility of combining the two training data sets is determined based on the distance between the regions of interest contained in each training data set.
[0103] The following outlines the method for generating training data in this embodiment. Here, we will explain using the example of splitting an image vertically into two halves and then combining them. Furthermore, similar to the first embodiment described above, we will explain using the example of generating a training model that recognizes lesions (areas of interest) from endoscopic images.
[0104] Figure 13 is a conceptual diagram of the determination of whether synthesis is possible or not.
[0105] The lesion included in the first image data is designated as the first lesion X1, and the lesion included in the second image data is designated as the second lesion X2.
[0106] The distance between the first lesion X1 and the second lesion X2 is calculated, and based on the calculated distance, it is determined whether or not the lesions can be combined.
[0107] Here, the distance between the first lesion X1 and the second lesion X2 is the distance between them within the superimposed image data of the first and second image data. In other words, it is the distance between them when the first and second image data are superimposed. In this embodiment, the image is divided vertically and then combined, so the distance V in the vertical direction of the image is calculated.
[0108] If the calculated distance V is greater than or equal to the threshold ThV, it is determined that synthesis is possible. In other words, synthesis is determined to be possible if the first lesion X1 and the second lesion X2 are separated by a distance of ThV or more. Here, the threshold ThV is set from the perspective of its impact on learning, similar to the threshold Th in the first embodiment described above. Therefore, when the generated training data is used to train a neural network using convolutional processing, it is set based on the size of the receptive field, in particular, the size of the receptive field of the first convolutional layer. For example, if the size (length × width) of the receptive field of the first convolutional layer is m × n, the threshold ThV is set to a value at least greater than m.
[0109] When two image data can be combined, a boundary line BL is set between the two lesion areas X1 and X2. In this embodiment, the image is divided vertically into two halves and combined, so a horizontal boundary line BL is set. The boundary line BL is set at the midpoint between the two lesion areas X1 and X2.
[0110] After setting the boundary line BL, the image is divided along the set boundary line BL, and the images of the regions containing the lesion are combined to generate new image data. In the example shown in Figure 13, the image of the lower region of the first image data and the image of the upper region of the second image data are combined to generate new image data.
[0111] In this embodiment, the distance V between the first lesion X1 and the second lesion X2 is an example of a positional relationship. Furthermore, the condition for determining that synthesis is possible, namely the condition that the distance V is greater than or equal to the threshold ThV, is an example of a predetermined condition.
[0112] [Hardware configuration] Figure 14 is a block diagram of the main functions of the learning data generation device.
[0113] As shown in the figure, the learning data generation device mainly has functions such as a first learning data acquisition unit 21, a second learning data acquisition unit 22, a distance calculation unit 23, a synthesis feasibility determination unit 24, a boundary line setting unit 25, a new learning data generation unit 26, and a new learning data recording unit 27. The functions of each unit are realized by the processor executing a predetermined program.
[0114] The first learning data acquisition unit 21 performs the process of acquiring learning data to be used as the first learning data. In this embodiment, the learning data to be used as the first learning data is acquired from the auxiliary storage device 4.
[0115] The second training data acquisition unit 22 performs the process of acquiring training data to be used as the second training data. Similar to the first training data, it acquires training data to be used as the second training data from the auxiliary storage device 4.
[0116] The distance calculation unit 23 performs a process to calculate the distance between lesions contained in the first training data and the second training data. That is, it calculates the distance between a lesion (first lesion) contained in the image data of the first training data (first image data) and a lesion (second lesion) contained in the image data of the second training data (second image data). In this embodiment, the distance V in the vertical direction of the image is calculated.
[0117] The synthesis feasibility determination unit 24 performs a process to determine whether or not the two training data can be synthesized based on the distance calculated by the distance calculation unit 23. Specifically, it determines whether or not synthesis is possible based on whether the distance V calculated by the distance calculation unit 23 is greater than or equal to the threshold ThV. If the distance V is greater than or equal to the threshold ThV, it is determined that synthesis is possible.
[0118] The boundary line setting unit 25 performs the process of setting a boundary line when the two training data can be combined. In this embodiment, a horizontal boundary line is set at the midpoint between the two lesion areas (midpoint in the vertical direction) (see Figure 13).
[0119] The new training data generation unit 26 performs the process of generating new training data by combining the first training data and the second training data. Specifically, it divides the image based on the set boundary line and combines the images of the regions containing the lesion to generate new training data. For example, if the lesion is located in the region below the set boundary line in the first training data, the image of the region below the boundary line of the first image data and the image of the region above the boundary line of the second image data are combined to generate new image data. Similarly, for the ground truth data, the image of the region below the boundary line of the first ground truth data and the image of the region above the boundary line of the second ground truth data are combined to generate new ground truth data. Also, for example, if the lesion is located in the region above the set boundary line in the first training data, the image of the region above the boundary line of the first image data and the image of the region below the boundary line of the second image data are combined to generate new image data. Similarly, for the ground truth data, the image of the region above the boundary line of the first ground truth data and the image of the region below the boundary line of the second ground truth data are combined to generate new ground truth data. Similar to the first embodiment described above, the synthesis method is not particularly limited. Methods such as synthesis by overwriting, or synthesis by extracting images of the region to be synthesized from each image data and then synthesizing them can be employed.
[0120] [Generating new training data] Figure 15 is a flowchart showing an example of the procedure for generating new training data.
[0121] First, the first training data is acquired (step S11). Specifically, one of the multiple training data stored in the auxiliary storage device 4 is read to acquire the first training data.
[0122] Next, the second training data is acquired (step S12). Similar to the first training data, one of the multiple training data stored in the auxiliary storage device 4 is read to acquire the second training data.
[0123] Next, the distance between lesions (areas of interest) contained in the acquired first training data and second training data is calculated (step S13). That is, the distance V (distance in the vertical direction of the images) between the lesion (first lesion) contained in the image data of the first training data (first image data) and the lesion (second lesion) contained in the image data of the second training data (second image data) is calculated. The distance here is the distance between the two when the images of each image data are superimposed (see Figure 13).
[0124] Next, based on the calculated distance, it is determined whether the two training data can be combined (step S14). Here, it is determined whether the calculated distance V is greater than or equal to the threshold ThV, and the ability to combine them is determined. If the calculated distance V is greater than or equal to the threshold ThV, it is determined that the data can be combined. On the other hand, if the calculated distance V is less than the threshold ThV, it is determined that the data cannot be combined.
[0125] If it is determined that the data cannot be synthesized, the presence or absence of unprocessed second training data is determined (step S15). That is, the presence or absence of training data that has not yet been used as second training data is determined.
[0126] If there is unprocessed second training data, the process returns to step S12, one of the unprocessed second training data is selected, and the distance between the lesions is calculated between it and the newly selected second training data (step S13). In other words, the second training data is modified, and the feasibility of synthesis is determined again.
[0127] On the other hand, if there is no unprocessed second training data, the presence or absence of unprocessed first training data is determined (step S16). That is, the presence or absence of training data that has not yet been used as first training data is determined.
[0128] If there is no unprocessed first training data, the process ends. On the other hand, if there is unprocessed first training data, the process returns to step S11, one of the unprocessed first training data is selected, and processing begins anew. In other words, the first training data is modified, and the process of generating new training data begins.
[0129] In step S14, if it is determined that the images can be combined, a boundary line is set (step S17). In this embodiment, a boundary line BL is set that divides the image vertically (see Figure 13). The boundary line BL is set at the midpoint between the first lesion X1 and the second lesion X2 (the midpoint in the vertical direction of the image).
[0130] After setting the boundary line BL, new training data is generated (step S18). That is, new image data and new ground truth data are generated.
[0131] New image data is generated by combining the image of the region containing the lesion from the first image data with the image of the region containing the lesion from the second image data. Therefore, for example, if the lesion is located in the region above the boundary line BL in the first image data, the new image data is generated by combining the image of the region above the boundary line BL in the first image data with the image of the region below the boundary line BL in the second image data. Similarly, if the lesion is located in the region below the boundary line BL in the first image data, the new image data is generated by combining the image of the region below the boundary line BL in the first image data with the image of the region above the boundary line in the second image data. In the same manner, the first ground truth data and the second ground truth data are combined to generate new ground truth data. The generated new training data is stored in the auxiliary storage device 4.
[0132] After generating new training data, it is determined whether there is any unprocessed first training data (step S19). If there is no unprocessed first training data, the process ends. On the other hand, if there is unprocessed first training data, the process ends. 11 Returning to the previous step, we retrieve one data point from the unprocessed first training data and begin the process of generating new training data.
[0133] The training data used to generate new training data is considered processed training data and will not be used to generate new training data in the future. Similarly, first training data that is determined to be uncomposite (first training data for which there is no second training data that can be composed) is also considered processed training data. On the other hand, even if second training data is determined to be uncomposite, if the first training data is switched, it will not be considered processed training data. This is because there is a possibility that it can be composed with other first training data.
[0134] As described above, according to this embodiment, similar to the first embodiment, new training data can be generated by extracting only the region containing the lesion from two training data sets. This reduces the amount of training data and thus the time required for training. In other words, training can be performed efficiently.
[0135] [Differentiation] [Image division method] In the above embodiment, the case of dividing an image into two halves vertically and combining them was described as an example, but the manner in which the image is divided is not limited to this. The boundary line is set according to the manner in which the image is divided.
[0136] Figure 16 shows another example of how to set a boundary line.
[0137] The figure shows an example of splitting an image horizontally into two halves and combining them. In this case, the boundary line BL is set vertically.
[0138] In this case, the feasibility of combining the images is determined based on the distance between the lesions in the lateral direction of the images. Specifically, it is determined based on the lateral distance H between the lesion (first lesion) X1 in the first image data and the lesion (second lesion) X2 in the second image data. If the distance H is greater than or equal to the threshold ThH, the two training data sets are deemed capable of being combined. On the other hand, if the distance H is less than the threshold ThH, the data sets are deemed incapable of being combined.
[0139] New training data is generated by combining regions that contain lesions. For example, if a lesion is located in the region to the left of the boundary of the first image data, the image of the region to the left of the boundary of the first image data and the image of the region to the right of the boundary of the second image data are combined to generate new image data. On the other hand, if a lesion is located in the region to the right of the boundary of the first image data, the image of the region to the right of the boundary of the first image data and the image of the region to the left of the boundary of the second image data are combined to generate new image data. New ground truth data is generated using a similar method.
[0140] [A method for dynamically changing the boundary settings] In the above embodiment, the method of dividing the image is fixed, but it may also be configured to switch for each training data to be combined. In other words, it may be configured to dynamically change the boundary line setting for each training data to be combined.
[0141] Figure 17 shows an example of dynamically switching the boundary settings for each training data set being synthesized.
[0142] First, calculate the distance V between the first lesion X1 and the second lesion X2 in the vertical direction of the image. Determine whether the calculated distance V is greater than or equal to the threshold ThV.
[0143] If the calculated distance V is greater than or equal to the threshold ThV, the image is split vertically to generate new training data. In this case, a horizontal boundary line is set between the first lesion X1 and the second lesion X2. The images of the upper and lower regions of the set boundary line are combined to generate new training data.
[0144] On the other hand, if the calculated distance V is less than the threshold ThV, the horizontal distance is calculated. That is, the horizontal distance of the image direction In the following steps, the distance H between the first lesion X1 and the second lesion X2 is calculated. It is then determined whether the calculated distance H is greater than or equal to the threshold ThH.
[0145] If the calculated distance H is greater than or equal to the threshold ThH, the image is divided horizontally to generate new training data. In this case, a vertical boundary line (a boundary line extending vertically in the image) is set between the first lesion X1 and the second lesion X2. The image of the region to the right of the set boundary line and the image of the region to the left are combined to generate new training data.
[0146] On the other hand, if the calculated distance H is less than the threshold ThH, it is determined that synthesis is not possible.
[0147] In this way, by setting boundaries according to the training data to be synthesized, the number of combinations of training data that can be synthesized can be increased.
[0148] In the above example, we explained the case where the image is divided by a horizontal or vertical boundary line, but it is also possible to divide the image by setting the boundary line diagonally. In other words, as long as one region separated by the boundary line contains the region of interest of one training data set, and the other region contains the region of interest of the other training data set, the method of setting the boundary line is not particularly limited. Therefore, the boundary line may be set with a polyline, or it may be set with a curve.
[0149] Furthermore, the method for setting the optimal boundary line is not limited to the examples above, and various methods can be employed. Therefore, it is also possible to configure the system to directly determine the optimal boundary line from the location information of the lesions contained in the first training data and the location information of the lesions contained in the second training data.
[0150] [When there are multiple areas of interest] Figure 18 shows an example of setting boundaries when there are multiple regions of interest.
[0151] As shown in the figure, when the training data used to generate new training data (training data used for synthesis) has multiple regions of interest, it is preferable to set the boundary line such that all regions of interest of one training data set are included in one region separated by the boundary line, and all regions of interest of the other training data set are included in the other region. Here, "all regions of interest of one training data set are included in one region separated by the boundary line" means that all regions of interest of one training data set are included in one region, separated from the boundary line by a predetermined threshold or more. Similarly, "all regions of interest of the other training data set are included in the other region separated by the boundary line" means that all regions of interest of the other training data set are included in the other region, separated from the boundary line by a predetermined threshold or more.
[0152] The example shown in Figure 18 illustrates a case where the first training data has two lesions (first lesions) X1a and X1b within its image data (first image data), and the second training data has two lesions (second lesions) X2a and X2b within its image data (second image data). In this case, the boundary line BL is set such that all lesions (first lesions X1a and X1b) within the first image data are located in one region (the region to the left of the boundary line BL in Figure 18), and all lesions (second lesions X2a and X2b) within the second image data are located in the other region (the region to the right of the boundary line BL in Figure 18).
[0153] Furthermore, a prerequisite for image synthesis is that the distance between all lesions in the first image data and all lesions in the second image data must be greater than or equal to a threshold. This condition is met between the nearest lesion and the other lesions if it is met between them. Therefore, if the distance between the nearest lesion and the other lesions is greater than or equal to a threshold, it can be determined that image synthesis is possible.
[0154] [Generating a Learning Model] Next, we will explain how to generate a learning model using the generated training data. Here, we will explain using the example of generating a learning model that recognizes lesions from images taken with an endoscope, in particular a learning model that recognizes the area occupied by the lesion within the image (a learning model that performs image segmentation).
[0155] [Learning Model Generation Device (Learning Model Generation Method)] The generation of the learning model is performed using a learning model generation device. This device consists of a computer. The same computer used to generate the training data can be used. Therefore, a description of its hardware configuration is omitted.
[0156] Figure 19 is a block diagram of the main functions of the learning model generation device.
[0157] As shown in the figure, the learning model generation device 100 has functions such as a learning data acquisition unit 111 that acquires learning data, a learning unit 112 that trains the learning model 200 using the acquired learning data, and a learning control unit 113 that controls the learning. The functions of each unit are realized by a processor in the computer executing a predetermined program (learning model generation program). The program executed by the processor, and the data necessary for processing, etc., are stored in an auxiliary storage device in the computer.
[0158] The learning data acquisition unit 111 acquires the learning data to be used for learning. This learning data is the new learning data (third learning data) generated by the learning data generation device 1. The learning data is stored in advance as a dataset in the auxiliary storage device. Therefore, the learning data acquisition unit 111 sequentially reads and acquires the learning data from the auxiliary storage device.
[0159] The learning unit 112 trains the learning model 200 using the learning data acquired by the learning data acquisition unit 111. As described above, the learning model used for image segmentation can be, for example, U-net, FCN, SegNet, PSPNet, Deeplabv3+, etc. Since the training of these models is a well-known technique, a detailed explanation will be omitted.
[0160] The learning control unit 113 controls the acquisition of learning data by the learning data acquisition unit 111 and the learning process performed by the learning unit 112.
[0161] The learning model generation device 100, configured as described above, uses the learning data acquired by the learning data acquisition unit 111 to train the learning model 200 and generate a learning model that performs desired image recognition. In this embodiment, a learning model that recognizes the region of a lesion from an endoscopic image is generated. Here, the learning data acquired by the learning data acquisition unit 111 is learning data generated by synthesizing multiple learning data. Therefore, compared to training using the original learning data (learning data before synthesis), the same learning effect can be obtained with a smaller number of data. In addition, this can shorten the training time.
[0162] Generally, in deep learning, a single dataset is repeatedly trained to generate a learning model with the desired accuracy. Therefore, in this embodiment as well, the learning model is repeatedly trained using a dataset composed of new training data.
[0163] The generated learning model is applied to an image recognition device or system. In this embodiment, it is applied to an endoscope device or system. For example, it is incorporated into an endoscopic image processing device that processes images taken with an endoscope (endoscopic images) and used for automatic recognition of lesions.
[0164] [Differentiation] [Learning using the first training data and / or the second training data] During training, the system can be configured to use not only the new training data but also the training data used to generate the new training data.
[0165] For example, if two sets of training data (first training data and second training data) are combined to generate new training data, the system can be configured to perform training with the first training data and / or the second training data in addition to training with the new training data. In this case, the first training data and / or the second training data may be combined to form a dataset, or some of the multiple training iterations may be replaced with training using the first training data and / or the second training data. As described above, deep learning generates a learning model with the desired accuracy by repeatedly training on a single dataset multiple times. Therefore, the system can be configured to replace at least one of the multiple training iterations with training using the first training data and / or the second training data. For example, a dataset consisting of the new training data and a dataset consisting of the first training data and / or the second training data can be prepared, and training on each dataset can be performed alternately. For example, the first training run uses a dataset consisting of the first and / or second training data, the second training run uses a dataset consisting of new training data, the third training run uses a dataset consisting of the first and / or second training data, the fourth training run uses a dataset consisting of new training data, and so on, alternating between training with each dataset.
[0166] Furthermore, for example, it is possible to prepare a dataset consisting of new training data, a dataset consisting of the first training data, and a dataset consisting of the second training data, and combine training using each dataset. For example, the first training session might be using the dataset consisting of the first training data, the second training session using the dataset consisting of new training data, the third training session using the dataset consisting of the second training data, the fourth training session using the dataset consisting of new training data, and so on, combining training using each dataset.
[0167] Furthermore, it is not necessary to use all of the training data that makes up the dataset in a single training run; training can be performed using only a portion of the training data.
[0168] In this way, by using the training data used to generate the new training data, in addition to the new training data, the impact of synthesis on learning can be reduced. In other words, the impact of image transitions on learning can be reduced.
[0169] [Learning excluding boundary regions] When training a model using new training data, it is also possible to employ a method that excludes the boundary region of image synthesis. In this case, for example, a region to be excluded is set within a certain range on both sides of the boundary line and excluded from the training target. When new training data is generated based on a fixed boundary line, the region to be excluded can be fixed and the model can be trained. The size of the region to be excluded is set considering its impact on training. Therefore, when used for training a neural network using convolutional processing, it is preferable to set it based on the size of the receptive field. Furthermore, it is preferable to set the region to be excluded to at least 1 pixel on both sides of the boundary line.
[0170] [Other embodiments] [Learning Model] In the above embodiment, we described the case of generating a learning model that recognizes lesions from endoscopic images as an example, but the learning model to be generated is not limited to this. The same can be applied to generating learning models for other purposes.
[0171] Furthermore, while the above embodiment described the case of generating a learning model that performs image segmentation, particularly semantic segmentation, the learning models to which the present invention is applied are not limited to this. For example, it can also be applied when generating a learning model that performs instance segmentation as a learning model that performs image segmentation. Examples of learning models that perform instance segmentation include Mask R-CNN and Masklab. In addition, it can also be applied when generating learning models that perform image classification, learning models that perform object detection, and the like.
[0172] [Correct Data] The ground truth data is set according to the model being trained. Therefore, for example, when generating a training model for object detection, ground truth data is generated that shows the location of the region of interest using bounding boxes or similar methods. In this case, the ground truth data can consist of, for example, coordinate information.
[0173] Furthermore, for learning models that perform image classification, ground truth data in the form of image data is not required; they can be constructed using only so-called label information.
[0174] [Hardware configuration] The functions of a learning data generation device and a learning model generation device can be implemented using various types of processors. These include general-purpose processors such as CPUs (Central Processing Units) and / or GPUs (Graphics Processing Units) that execute programs and function as various processing units; programmable logic devices (PLDs) such as FPGAs (Field Programmable Gate Arrays) whose circuit configurations can be changed after manufacturing; and dedicated electrical circuits such as ASICs (Application Specific Integrated Circuits) that have circuit configurations specifically designed to perform particular processing. A program is synonymous with software.
[0175] A single processing unit may be composed of one of these various processors, or it may be composed of two or more processors of the same or different type. For example, a single processing unit may be composed of multiple FPGAs, or a combination of a CPU and an FPGA. Alternatively, multiple processing units may be composed of a single processor. Examples of composing multiple processing units with a single processor include, firstly, a configuration where one or more CPUs and software are combined to form a single processor, and this processor functions as multiple processing units, as is typical of computers used as clients or servers. Secondly, a configuration where a processor is used that realizes the functions of the entire system, including multiple processing units, on a single IC (Integrated Circuit) chip, as is typical of a System on Chip (SoC). Thus, various processing units are configured, in terms of hardware structure, using one or more of the above-mentioned various processors. [Explanation of symbols]
[0176] 1. Training data generation device 2 processors 4 Auxiliary storage 5 Input devices 6. Output device 11. First Training Data Acquisition Unit 12 Location identification part 13. Second Learning Data Acquisition Unit 14 Synthesis possibility judgment unit 15 New training data generation unit 16 New learning data recording unit 21. First Learning Data Acquisition Unit 22 Second Learning Data Acquisition Unit 23 Distance Calculation Unit 24 Synthesis possibility judgment unit 25 Boundary line setting section 26 New learning data generation unit 27 New learning data recording unit 100 Learning Model Generator 111 Training Data Acquisition Unit 112 Learning Department 113 Learning Control Unit 200 Learning Models BL boundary line UA upper area LA lower area RF receptive field X Lesion X1 Lesion (First Lesion) X1a Lesion (First Lesion) X2 Lesion (Second Lesion) X2a Lesion (Second Lesion) S1-S9 Procedure for generating new training data S11-S19 Procedure for generating new training data
Claims
1. A learning data generation device that generates learning data, Equipped with a processor, The aforementioned processor, First image data and second image data, each having a region of interest, are acquired. If the positional relationship between the region of interest in the first image data and the region of interest in the second image data satisfies a predetermined condition, the image of the region including the region of interest in the first image data and the image of the region including the region of interest in the second image data are combined to generate a third image data. It is a training data generation device, The predetermined conditions include the fact that the region of interest in the first image data is located within a first region in the image, and the region of interest in the second image data is located within a second region different from the first region in the image. A device for generating training data.
2. The predetermined conditions include the fact that the region of interest in the first image data is located within the first region at a distance of a threshold or more from the boundary line separating the first region and the second region, and the region of interest in the second image data is located within the second region at a distance of a threshold or more from the boundary line. The learning data generation device according to claim 1.
3. The predetermined conditions include the fact that a plurality of the regions of interest in the first image data are located within the first region at a distance of a threshold or more from the boundary line separating the first region and the second region, and that a plurality of the regions of interest in the second image data are located within the second region at a distance of a threshold or more from the boundary line. The learning data generation device according to claim 1.
4. When the aforementioned training data is used to train a neural network using convolution, The threshold is set based on the size of the receptive field of the first convolutional layer. The learning data generation device according to claim 2 or 3.
5. The processor generates the third image data by combining the image of the first region of the first image data with the image of the region of the second image data other than the first region. The learning data generation device according to claim 1.
6. The processor generates the third image data by overwriting the image of the region of the first image data other than the first region with the image of the region of the second image data other than the first region. The learning data generation device according to claim 5.
7. The predetermined conditions include the fact that the region of interest in the first image data and the region of interest in the second image data are separated by a threshold or more. A learning data generation device according to any one of claims 1 to 3.
8. The aforementioned processor, A boundary line is set between the area of interest in the first image data and the area of interest in the second image data to divide the image into multiple regions. The third image data is generated by combining the image of the first image data in the region containing the region of interest among a plurality of regions of the first image data divided by the boundary line, and the image of the second image data in the region containing the region of interest among a plurality of regions of the second image data divided by the boundary line. The learning data generation device according to claim 7.
9. The processor generates the third image data by overwriting the image of the region of the first image data other than the region of interest with the image of the region of interest of the second image data. The learning data generation device according to claim 8.
10. When the aforementioned training data is used to train a neural network using convolution, The threshold is set based on the size of the receptive field of the first convolutional layer. The learning data generation device according to claim 7.
11. A learning data generation device that generates learning data, Equipped with a processor, The aforementioned processor, First image data and second image data, each having a region of interest, are acquired. If the positional relationship between the region of interest in the first image data and the region of interest in the second image data satisfies a predetermined condition, the image of the region including the region of interest in the first image data and the image of the region including the region of interest in the second image data are combined to generate a third image data. It is a training data generation device, The predetermined conditions include the fact that the region of interest in the first image data and the region of interest in the second image data are separated by a threshold or more. A device for generating training data.
12. The aforementioned processor, First correct answer data indicating the correct answer for the first image data and second correct answer data indicating the correct answer for the second image data are obtained. A third set of correct data indicating the correct image data is generated from the first set of correct data and the second set of correct data. A learning data generation device according to claim 1, 2, 3, or 11.
13. The processor generates a third correct answer data indicating the correct answer of the third image data from the first correct answer data and the second correct answer data, according to the conditions for generating the third image data from the first image data and the second image data. The learning data generation device according to claim 12.
14. The first and second correct data are mask data for the region of interest. The learning data generation device according to claim 12.
15. A learning model generation device that generates a learning model, Equipped with a processor, The aforementioned processor, A third image data generated by the learning data generation device according to claim 1, 2, 3, or 11 is acquired, The learning model is trained using the third image data. A learning model generation device.
16. The processor further uses at least one of the first image data and the second image data used to generate the third image data to train the learning model. A learning model generation device according to claim 15.
17. The processor performs learning using the third image data and learning using at least one of the first image data and the second image data. A learning model generation device according to claim 16.
18. The processor trains the learning model by excluding the boundary region of the image synthesis of the third image data. A learning model generation device according to claim 15.
19. A method for generating training data, which generates training data, The steps include acquiring first image data and second image data, each having a region of interest, The steps include determining whether the region of interest in the first image data and the region of interest in the second image data are in a specific positional relationship, If the positional relationship between the region of interest in the first image data and the region of interest in the second image data satisfies a predetermined condition, the third image data is generated by combining the image of the region of interest in the first image data and the image of the region of interest in the second image data. Includes, The predetermined conditions include the fact that the region of interest in the first image data is located within a first region in the image, and the region of interest in the second image data is located within a second region different from the first region in the image. Method for generating training data.
20. A method for generating training data, which generates training data, The steps include acquiring first image data and second image data, each having a region of interest, The steps include determining whether the region of interest in the first image data and the region of interest in the second image data are in a specific positional relationship, If the positional relationship between the region of interest in the first image data and the region of interest in the second image data satisfies a predetermined condition, the third image data is generated by combining the image of the region of interest in the first image data and the image of the region of interest in the second image data. Includes, The predetermined conditions include the fact that the region of interest in the first image data and the region of interest in the second image data are separated by a threshold or more. Method for generating training data.
21. A method for generating a learning model, The steps include acquiring first image data and second image data, each having a region of interest, If the positional relationship between the region of interest in the first image data and the region of interest in the second image data satisfies a predetermined condition, the third image data is generated by combining the image of the region of interest in the first image data and the image of the region of interest in the second image data. The steps include training the learning model using the third image data, Includes, The predetermined conditions include the fact that the region of interest in the first image data is located within a first region in the image, and the region of interest in the second image data is located within a second region different from the first region in the image. Method for generating a learning model.
22. A method for generating a learning model, The steps include acquiring first image data and second image data, each having a region of interest, If the positional relationship between the region of interest in the first image data and the region of interest in the second image data satisfies a predetermined condition, the third image data is generated by combining the image of the region of interest in the first image data and the image of the region of interest in the second image data. The steps include training the learning model using the third image data, Includes, The predetermined conditions include the fact that the region of interest in the first image data and the region of interest in the second image data are separated by a threshold or more. Method for generating a learning model.
Citation Information
Patent Citations
Medical image processing device, image formation method and image formation program
JP2020018705A
Information processing apparatus, information processing method and program
JP2020060883A
Teacher image generation program, teacher image generation method, and teacher image generation system
JP2021019677A
Image processing method, teacher data generation method, learned model generation method, disease onset prediction method, image processing device, image processing program, and recording medium that records the program
JP2021065606A
Medical image processing apparatus, medical image processing method, and program
JP2021086560A