Image processing device, image processing method, and program

The image processing device with diverse recognizers and teacher label generation enhances AI performance in lesion detection by selecting optimal training data, addressing inconsistent outputs from multiple AIs.

JP7788822B2Active Publication Date: 2025-12-19FUJIFILM CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021148846
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-13
Publication Date
2025-12-19
Estimated Expiration
2041-09-13

AI Technical Summary

Technical Problem

Existing image processing systems for lesion detection in medical imaging face variability in AI performance due to inconsistent training data, leading to differing output results from multiple AIs when processing the same images, which hinders effective machine learning.

Method used

An image processing device that utilizes a plurality of recognizers with diverse training data, including different facilities, devices, and conditions, to determine the suitability of image frames as learning data based on recognition results, and generates teacher labels to enhance machine learning efficiency.

Benefits of technology

This approach allows for the efficient acquisition of training data that can improve AI performance by leveraging diverse recognition results, ensuring effective machine learning and consistent output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007788822000001
    Figure 0007788822000001
  • Figure 0007788822000002
    Figure 0007788822000002
  • Figure 0007788822000003
    Figure 0007788822000003
Patent Text Reader

Abstract

To provide an image processing device, an image processing method, and a program that can efficiently obtain learning data with which effective machine learning can be expected.SOLUTION: An image processing device 10 includes a processor 1 and a plurality of recognizers, and the processor 1 acquires a video acquired by a medical apparatus, causes the plurality of recognizers to perform processing for recognizing a lesion in image frames forming the video to acquire a recognition result of each of the plurality of recognizers, and determines whether to use the image frame as learning data to be used for machine learning on the basis of the recognition result of each of the plurality of recognizers.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device, an image processing method, and a program, and more particularly to an image processing device, an image processing method, and a program for determining learning data to be used in machine learning. [Background technology]

[0002] 2. Description of the Related Art In recent years, in the medical field, images of an object to be examined are used to detect lesions and assist doctors and others in making diagnoses.

[0003] For example, Patent Document 1 describes a technology that receives a plurality of medical data (image data and clinical data) as input and outputs a diagnosis based on this data. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Special Publication No. 2010-504129 Summary of the Invention [Problem to be solved by the invention]

[0005] When detecting lesions from images, AI (Artificial Intelligence: learning model) is trained using training data and teacher data to complete a trained AI (trained model), which is then used to detect lesions. The training data used to train AI is one of the factors that determine the performance of the AI. By using training data that allows for effective machine learning, we can expect to see an improvement in AI performance that is effective relative to the amount of learning.

[0006] On the other hand, even when the same image is input to multiple AIs, the output results of each AI may vary. Such images are difficult for AI to judge, detect, etc., and are excellent as training data. Using such excellent training data to train AI through machine learning can effectively improve AI performance.

[0007] The present invention has been made in consideration of the above circumstances, and its purpose is to provide an image processing device, an image processing method, and a program that can efficiently obtain learning data that can be expected to enable effective machine learning. [Means for solving the problem]

[0008] An image processing device, which is one aspect of the present invention for achieving the above-mentioned object, is an image processing device that includes a processor and a plurality of recognizers, wherein the processor acquires a video captured by a medical device, causes the plurality of recognizers to perform processing to recognize lesions on the image frames that make up the video, acquires the recognition results of each of the plurality of recognizers, and determines whether or not to use the image frames as learning data to be used for machine learning based on the recognition results of each of the plurality of recognizers.

[0009] According to this aspect, an image frame is input to a plurality of recognizers, and whether or not to use the image frame as training data for machine learning is determined based on the recognition results of the plurality of recognizers. This allows for efficient acquisition of training data that can be used for effective machine learning.

[0010] Preferably, the plurality of recognizers differ in at least one of structure, type, and parameters of the recognizers.

[0011] Preferably, each of the plurality of recognizers is trained using different training data.

[0012] Preferably, each of the plurality of recognizers is trained by machine learning using different training data obtained from a different medical device.

[0013] Preferably, the plurality of recognizers are each trained by machine learning using different training data obtained at facilities in different countries or regions.

[0014] Preferably, each of the plurality of recognizers is trained by machine learning using different training data captured under different imaging conditions.

[0015] Preferably, when the processor determines that the image frame to which the diagnosis result is assigned is training data, the processor generates a teacher label for the training data based on the diagnosis result.

[0016] Preferably, the training data determined by the processor is used to train a learning model that performs machine learning.

[0017] Preferably, the processor causes the learning model to learn the training data with sample weights determined based on the distribution of the recognition results of each of the plurality of recognizers.

[0018] Preferably, the processor generates a training label for machine learning based on the distribution of the recognition results.

[0019] Preferably, the processor changes sample weights in the machine learning in accordance with the degree of variability in the recognition results.

[0020] Preferably, the processor causes multiple recognizers to perform a process to recognize lesions on chronologically consecutive image frames, obtains the recognition results of each of the multiple recognizers, and determines whether or not to use the image frames for machine learning based on the recognition results of each of the multiple chronologically consecutive recognizers.

[0021] Preferably, at least one of the plurality of recognizers outputs a recognition result while the video is being captured, and the other recognizers output a recognition result after a first time has elapsed after the video is captured.

[0022] Another aspect of the present invention is an image processing method for an image processing device having a processor and a plurality of recognizers, in which the processor performs the steps of acquiring a video captured by a medical device, causing the plurality of recognizers to perform a process of recognizing lesions on image frames that make up the video and acquiring the recognition results of each of the plurality of recognizers, and determining whether or not to use the image frames as learning data to be used in machine learning based on the recognition results of each of the plurality of recognizers.

[0023] Another aspect of the present invention is a program that causes the processor to execute an image processing method of an image processing device that has a processor and multiple recognizers, and causes the processor to perform the following steps: acquiring a video captured by a medical device; causing the multiple recognizers to perform a process of recognizing lesions on image frames that make up the video and acquiring the recognition results of each of the multiple recognizers; and determining whether or not to use the image frames as learning data to be used in machine learning based on the recognition results of each of the multiple recognizers. [Effects of the Invention]

[0024] According to the present invention, an image frame is input to a plurality of recognizers, and whether or not the image frame will be used as training data for machine learning is determined based on the recognition results of the plurality of recognizers, thereby making it possible to efficiently obtain training data that can be used for effective machine learning. [Brief explanation of the drawings]

[0025] [Figure 1] FIG. 1 is a block diagram showing the main configuration of an image processing device. [Figure 2] FIG. 2 is a diagram conceptually showing an inspection video. [Figure 3] FIG. 3 is a diagram illustrating an example of the recognition unit. [Figure 4] FIG. 4 is a diagram for explaining the determination of whether or not to use the learning data used in machine learning in the learning use determination unit. [Figure 5] FIG. 5 is a flowchart showing an image processing method performed using the image processing device. [Figure 6] FIG. 6 is a block diagram showing the main configuration of the image processing device. [Figure 7] FIG. 7 is a diagram illustrating the learning availability determining unit and the first teacher label generating unit. [Figure 8] FIG. 8 is a diagram illustrating a case where the first truth label generating unit generates truth labels. [Figure 9] FIG. 9 is a functional block diagram showing the main functions of the learning control unit and the learning model. [Figure 10] FIG. 10 is a block diagram showing the main configuration of the image processing device. [Figure 11] FIG. 11 is a diagram illustrating the learning availability determining unit and the second teacher label generating unit. [Figure 12] FIG. 12 shows the case where an image frame is input to the recognition unit. [Figure 13] FIG. 13 is a diagram showing a modified example of the recognition unit. [Figure 14] FIG. 14 is a diagram illustrating a modified example of the learning availability determining unit. [Figure 15] FIG. 15 is a diagram showing the overall configuration of the endoscope device. [Figure 16] FIG. 16 is a functional block diagram of the endoscope device. DETAILED DESCRIPTION OF THE INVENTION

[0026] Preferred embodiments of an image processing apparatus, an image processing method, and a program according to the present invention will be described below with reference to the accompanying drawings.

[0027] First Embodiment FIG. 1 is a block diagram showing the main configuration of an image processing apparatus 10 according to this embodiment.

[0028] The image processing device 10 is mounted on, for example, a computer. The image processing device 10 mainly includes a first processor (processor) 1 and a storage unit 11. The first processor 1 is configured as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit) mounted on the computer. The storage unit 11 is configured as a ROM (Read Only Memory) and RAM (Random Access Memory) mounted on the computer.

[0029] The first processor 1 realizes various functions by executing programs stored in the storage unit 11. The first processor 1 functions as a video acquisition unit 12, a recognition unit 14, and a learning / usability determination unit 16.

[0030] The video acquisition unit 12 acquires an examination video (video) M captured by an endoscope device 500 (see Figures 15 and 16) from a database DB. The endoscope device 500 is an example of a medical device, and the examination video M is an example of a video. The video acquisition unit 12 can acquire videos captured by medical devices other than the above-mentioned examination video M. The examination video M is input via a data input unit of a computer constituting the image processing device 10, and the video acquisition unit 12 acquires the input examination video M.

[0031] 2 is a diagram conceptually showing the examination moving image M acquired by the moving image acquiring unit 12. The examination moving image M is an examination moving image of a large intestine examined by a lower bowel endoscope device.

[0032] As shown in FIG. 2, the examination video M is a video relating to an examination performed between time t1 and time t2. The examination video M is composed of multiple image frames N that are consecutive in time series, and each image frame N has information regarding the time at which it was captured. The image frame N has an image of the large intestine, which is the subject of the examination, captured during a lower endoscopy. Note that, although the examination video M captured during a lower endoscopy has been described in this example, the examination video is not limited to this. For example, the technology of the present disclosure can also be applied to an examination video captured during an upper endoscopy.

[0033] The recognition unit 14 (FIG. 1) performs a process to recognize lesions in image frames N constituting the examination video M acquired by the video acquisition unit 12. The recognition unit 14 is composed of multiple recognizers, and causes the multiple recognizers to perform a process to recognize lesions for each input image frame and output a recognition result. The recognition unit 14 then acquires the recognition results from each of the multiple recognizers. Each recognizer is a trained model that has undergone machine learning in advance. It is preferable that the multiple recognizers have diversity. Here, "diversity" means that the recognizers have different strengths or weaknesses in lesion recognition, or that the output entropy is large when the same image frame N is input. For example, the multiple recognizers may each have undergone machine learning using different training data. Furthermore, for example, the multiple recognizers may each have undergone machine learning using different training data obtained with different medical devices. Note that different training data refers to training data obtained with the same type of medical device (different facilities) or different types of medical devices (different endoscope models, etc.). Furthermore, for example, the multiple recognizers may each have undergone machine learning using different training data obtained with facilities in different countries or regions. Furthermore, for example, each of the multiple recognizers may perform machine learning using different learning data captured under different shooting conditions. Here, the shooting information includes the resolution, exposure time, white balance, frame rate, etc. As described above, the multiple recognizers constituting the recognition unit 14 are provided with the above-described diversity. This prevents the recognition results obtained from the multiple recognizers from always being uniform.

[0034] FIG. 3 is a diagram illustrating an example of the recognition unit 14. As shown in FIG.

[0035] 3, the recognition unit 14 is configured with a first recognizer (recognizer) 14A, a second recognizer (recognizer) 14B, a third recognizer (recognizer) 14C, and a fourth recognizer (recognizer) 14D. The first recognizer 14A to the fourth recognizer 14D are configured with trained models that have been subjected to machine learning in advance.

[0036] For example, the first recognizer 14A to the fourth recognizer 14D perform machine learning using learning data acquired at different facilities or hospitals. Specifically, the first recognizer 14A performs machine learning using learning data acquired at Hospital A, the second recognizer 14B performs machine learning using learning data acquired at Hospital B, the third recognizer 14C performs machine learning using learning data acquired at Hospital C, and the fourth recognizer 14D performs machine learning using learning data acquired at Hospital D.

[0037] Generally, the tendency of examination videos, such as the preferred image quality when capturing examination videos, may differ depending on the facility or hospital. Therefore, as described above, the first recognizer 14A to the fourth recognizer 14D perform machine learning using learning data acquired at different facilities or hospitals, thereby making it possible to configure the recognition unit 14 that has diversity in terms of the tendency of examination videos (such as the image quality of the examination videos).

[0038] The first to fourth recognizers 14A to 14D may perform machine learning on training data in which the distribution of facilities or hospitals constituting the training data is biased. For example, the training data generated by the first recognizer 14A is composed of 50% data from Hospital A, 25% data from Hospital B, 20% data from Hospital C, and 5% data from Hospital D. The training data generated by the second recognizer 14B is composed of 5% data from Hospital A, 50% data from Hospital B, 25% data from Hospital C, and 20% data from Hospital D. The training data generated by the third recognizer 14C is composed of 20% data from Hospital A, 5% data from Hospital B, 50% data from Hospital C, and 25% data from Hospital D. The training data generated by the fourth recognizer 14D is composed of 25% data from Hospital A, 20% data from Hospital B, 5% data from Hospital C, and 50% data from Hospital D.

[0039] Furthermore, for example, the first to fourth recognizers 14A to 14D may perform machine learning using data acquired in different countries or regions. Specifically, the first recognizer 14A performs machine learning using training data acquired in the United States, the second recognizer 14B performs machine learning using training data acquired in the Federal Republic of Germany, the third recognizer 14C performs machine learning using training data acquired in the People's Republic of China, and the fourth recognizer 14D performs machine learning using training data acquired in Japan.

[0040] The techniques (procedures) for endoscopic examinations may differ depending on the country or region. For example, the techniques for endoscopic examinations in Europe are often different from those in Japan due to factors such as a large amount of residue. Therefore, as described above, the first recognizer 14A to the fourth recognizer 14D perform machine learning using learning data acquired in different countries or regions, thereby enabling the recognition unit 14 to be configured with diversity for the techniques (procedures) for endoscopic examinations.

[0041] The first to fourth recognizers 14A to 14D may perform machine learning on training data in which the distribution of countries or regions constituting the training data is biased. For example, the training data generated by the first recognizer 14A is composed of 50% data from the United States, 25% data from the Federal Republic of Germany, 20% data from the People's Republic of China, and 5% data from Japan. The training data generated by the second recognizer 14B is composed of 5% data from the United States, 50% data from the Federal Republic of Germany, 25% data from the People's Republic of China, and 20% data from Japan. The training data generated by the third recognizer 14C is composed of 20% data from the United States, 5% data from the Federal Republic of Germany, 50% data from the People's Republic of China, and 25% data from Japan. The training data generated by the fourth recognizer 14D is composed of 25% data from the United States, 20% data from the Federal Republic of Germany, 5% data from the People's Republic of China, and 50% data from Japan.

[0042] Furthermore, for example, the first recognizer 14A to the fourth recognizer 14D may be configured to have different sizes. For example, the first recognizer 14A is configured as a recognizer that can operate while the endoscope device 500 is acquiring a moving image (immediately after acquiring the moving image: real time). Specifically, the first recognizer 14A receives successive image frames N that constitute the inspection moving image M and outputs a recognition result immediately after the image frame N is input. The second recognizer 14B is configured as a recognizer with a processing capability of 3 FPS (films per second), the third recognizer 14C is configured as a recognizer with a processing capability of 5 FPS, and the fourth recognizer 14D is configured as a recognizer with a processing capability of 10 FPS. The second recognizer 14B, the third recognizer 14C, and the fourth recognizer 14D output recognition results a first time after the moving image is acquired. Here, the first time period is a time period determined by the processing capabilities of the second recognizer 14B, the third recognizer 14C, and the fourth recognizer 14D. As described above, by making the sizes of the first recognizer 14A to the fourth recognizer 14D different, it is possible to use, as learning data, an image frame N that could not be successfully recognized by a recognizer that can operate during video acquisition (a recognizer that the user actually uses).

[0043] The learning use determination unit 16 (Figure 1) determines whether or not to use the image frame N input to the recognition unit 14 as learning data to be used for machine learning, based on the recognition results of each of the multiple recognitions obtained by the recognition unit 14.

[0044] The learning availability determining unit 16 determines whether to use an image frame N as learning data to be used for machine learning using various methods. For example, if the recognition results of the recognizers constituting the recognition unit 14 do not all match, the learning availability determining unit 16 determines the image frame N as learning data to be used for machine learning, and if the recognition results are all match, the learning availability determining unit 16 determines the image frame N as learning data not to be used for machine learning. An image frame N for which the recognition results of multiple recognizers all match is so-called easy learning data, and even if machine learning is performed using this learning data, a high level of machine learning effectiveness cannot be expected. Therefore, the learning availability determining unit 16 determines not to use an image frame N for which the recognition results of multiple recognizers all match as learning data. On the other hand, an image frame N for which the recognition results of multiple recognizers all do not match is learning data that is difficult to recognize, and effective performance improvement can be expected when machine learning is performed. Therefore, the learning availability determining unit 16 determines that an image frame N for which the recognition results of multiple recognizers all do not match is to be used as learning data.

[0045] FIG. 4 is a diagram for explaining the decision made by the learning use decision unit 16 as to whether or not to use the learning data used in machine learning.

[0046] Image frames N1 to N4, which are part of the inspection moving image M and are consecutive in time series, are input to the recognition unit 14 in sequence.

[0047] The first to fourth recognizers 14A to 14D constituting the recognition unit 14 output recognition results 1 to 4 for the input image frames N1 to N4.

[0048] When image frame N1 is input, the first recognizer 14A to the fourth recognizer 14D output recognition results 1 to 4, respectively. Of the output recognition results 1 to 4, only recognition result 1 is different from the other recognition results (recognition results 2 to 4). Therefore, since the recognition results did not all match, the learning use determination unit 16 determines that image frame N1 will be used as learning data for machine learning (image frame N1 is marked with a circle in the figure).

[0049] When image frame N2 is input, the first to fourth recognizers 14A to 14D output recognition results 1 to 4, respectively. The output recognition results 1 to 4 are all consistent. Therefore, the learning use determination unit 16 determines that image frame N3 will not be used as learning data for machine learning because the recognition results are all consistent (image frame N3 is marked with an "x" in the figure).

[0050] Similarly to image frame N1, recognition results 1 to 4 of image frames N3 and N4 were different from the other recognition results (recognition results 2 to 4) only in that recognition result 1 was different. Therefore, since the recognition results did not match in all respects, the learning use determination unit 16 decided to use image frames N3 and N4 as learning data for machine learning (image frame N1 is marked with a circle in the figure).

[0051] As described above, the learning use determination unit 16 determines to use image frame N as learning data if recognition results 1 to 4 all match, and determines to use image frame N as learning data if recognition results 1 to 4 do not all match.

[0052] 5 is a flowchart showing an image processing method performed using the image processing device 10 of this embodiment. The image processing method is performed by the first processor 1 of the image processing device 10 executing a program stored in the storage unit 11.

[0053] First, the video acquisition unit 12 acquires the inspection video M (step S10: video acquisition step). Then, the recognition unit 14 acquires the recognition results of the first recognizer 14A, the second recognizer 14B, the third recognizer 14C, and the fourth recognizer 14D (step S11: result acquisition step). Then, the learning use / non-use determination unit 16 determines whether or not the recognition results 1 to 4 of the first recognizer 14A, the second recognizer 14B, the third recognizer 14C, and the fourth recognizer 14D all match (step S12: learning use / non-use determination step). If the recognition results 1 to 4 all match, the learning use / non-use determination unit 16 determines not to use the image frame N as learning data (step S14). On the other hand, if the recognition results 1 to 4 do not all match, the learning use / non-use determination unit 16 determines to use the image frame N for learning (step S13).

[0054] As described above, according to this aspect, an image frame N is input to a plurality of recognizers, and based on the recognition results of the plurality of recognizers, it is determined whether or not the image frame N is to be used as training data for machine learning. This makes it possible to efficiently obtain training data that can be used for effective learning.

[0055] <Second embodiment> Next, a second embodiment of the present invention will be described. In this embodiment, learning data is determined, and a teacher label for the image frame N determined as learning data is generated from the assigned diagnosis result.

[0056] 6 is a block diagram showing the main configuration of the image processing device 10 of this embodiment. Note that the same reference numerals are used to denote the parts that have already been explained in FIG.

[0057] The image processing device 10 mainly comprises a first processor 1, a second processor (processor) 2, and a storage unit 11. The first processor 1 and the second processor 2 may be configured as the same CPU (or GPU), or may be configured as separate CPUs (or GPUs). The first processor 1 and the second processor 2 execute programs stored in the storage unit 11 to realize the functions shown in the functional blocks.

[0058] The first processor 1 includes a video acquisition unit 12, a recognition unit 14, and a learning availability determination unit 16. The second processor (processor) 2 includes a first teacher label generation unit 18, a learning control unit 20, and a learning model 22.

[0059] The first teacher label generation unit 18 generates teacher labels for the image frame N based on the assigned diagnostic results. Here, the diagnostic results are information that is assigned by a doctor or the like during an endoscopic examination, for example, and that is attached to the image frame. For example, the doctor assigns diagnostic results such as the presence or absence of a lesion, the type of lesion, and the severity of the lesion. The doctor inputs the diagnostic results using the handheld operation unit 102 of the endoscope device 500. The input diagnostic results are attached as attached information to the image frame N.

[0060] 7 is a diagram illustrating the learning use availability determining unit 16 and the first teacher label generating unit 18. Note that the same parts as those already explained in FIG. 4 are used and the explanation will be omitted.

[0061] Image frames N1 to N4, which are consecutive in time series and form a portion of the inspection moving image M, are sequentially input to the recognition unit 14. A diagnosis result (label B) is assigned to image frame N2.

[0062] When image frames N1, N3, and N4 are input, the first recognizer 14A to the fourth recognizer 14D output recognition results 1 to 4, respectively, and of the output recognition results 1 to 4, only recognition result 1 is different from the other recognition results (recognition results 2 to 4). Therefore, since the recognition results do not all match, the learning use determination unit 16 determines that image frames N1, N3, and N4 will be used as learning data for machine learning (image frame N1 is marked with a circle in the figure).

[0063] On the other hand, when image frame N2 is input, the first recognizer 14A to the fourth recognizer 14D output recognition results 1 to 4, respectively, and the output recognition results 1 to 4 are all consistent. Therefore, since the recognition results are all consistent, the learning use determination unit 16 determines that image frame N3 will not be used as learning data for machine learning (image frame N3 is marked with an "x" in the figure).

[0064] The first teacher label generation unit 18 generates a teacher label based on the diagnosis result assigned to the image frame N3. Specifically, the first teacher label generation unit 18 generates teacher labels for nearby image frames (e.g., image frames N1 to N4) based on the diagnosis result (label B) assigned to the image frame N3. Therefore, the teacher labels for image frames N1 to N4 are label B, and when any of image frames N1 to N4 is determined as training data, label B becomes the teacher label. The first teacher label generation unit 18 may assign sample weights to the teacher labels it generates. For example, the greater the variation in recognition results 1 to 4, the greater the sample weights the first teacher label generation unit 18 assigns to the teacher labels it generates. This allows machine learning to focus on training data (and teacher labels) that are easy for doctors to judge but difficult for a recognizer to judge.

[0065] FIG. 8 is a diagram illustrating a case where the first truth label generating unit 18 generates truth labels.

[0066] The first truth label generation unit 18 generates truth labels for nearby image frames based on the assigned diagnosis results. Here, the range of the neighborhood can be arbitrarily set by the user and can be changed depending on the inspection object and the frame rate of the inspection video M.

[0067] As shown in FIG. 8, when a diagnostic result is assigned to image frame N6, the first teacher label generation unit 18 generates teacher labels for, for example, two frames before and after (image frames N4 to N8) based on the diagnostic result assigned to image frame N6. Alternatively, the first teacher label generation unit 18 may generate teacher labels for, for example, five frames before and after (image frames N1 to N11) based on the diagnostic result assigned to image frame N6. Note that a sample weight may be assigned to the teacher label corresponding to each image frame. This sample weight may be assigned according to the temporal distance from image frame N6 to which the diagnostic result is assigned. For example, the sample weights of image frames N5 and N7 are set lower than those of image frames N1 and N11.

[0068] The learning control unit 20 causes the learning model 22 to perform machine learning. Specifically, the learning control unit 20 inputs the image frame N that has been determined to be used as learning data by the learning use availability determining unit 16 into the learning model 22, and causes the learning model 22 to perform learning. In addition, the learning control unit 20 obtains the teacher label generated by the first teacher label generating unit 18, obtains the error between the output result output from the learning model 22 and the teacher label, and updates the parameters of the learning model 22.

[0069] 9 is a functional block diagram showing the main functions of the learning control unit 20 and the learning model 22. The learning control unit 20 includes an error calculation unit 54 and a parameter update unit 56. In addition, a teacher label S is input to the learning control unit 20.

[0070] Once machine learning is complete, the learning model 22 becomes a recognizer that performs image recognition of the position and type of a region of interest (lesion) within the image frame N. The learning model 22 has a multi-layer structure and holds a multiplicity of weight parameters. The learning model 22 changes from an unlearned model to a trained model by updating the weight parameters from their initial values ​​to optimal values.

[0071] This learning model 22 includes an input layer 52A, an intermediate layer 52B, and an output layer 52C. The input layer 52A, the intermediate layer 52B, and the output layer 52C each have a structure in which a plurality of "nodes" are connected by "edges." A synthetic image C, which is the learning target, is input to the input layer 52A.

[0072] The intermediate layer 52B extracts features from the image input from the input layer 52A. The intermediate layer 52B includes multiple sets of convolutional layers and pooling layers, and a fully connected layer. The convolutional layer performs a convolution operation using a filter on nearby nodes in the previous layer to obtain a feature map. The pooling layer reduces the feature map output from the convolutional layer to create a new feature map. The fully connected layer connects all nodes in the previous layer (here, the pooling layer). The convolutional layer is responsible for feature extraction, such as edge extraction from the image, and the pooling layer is responsible for providing robustness to the extracted features so that they are not affected by translation, etc. Note that the intermediate layer 52B is not limited to cases where a convolutional layer and a pooling layer are combined into one set, but may also include cases where convolutional layers are consecutive and a normalization layer.

[0073] The output layer 52C is a layer that outputs the recognition results of the position and type of the region of interest in the image frame N based on the features extracted by the intermediate layer 52B.

[0074] The trained learning model 22 outputs the recognition results of the position of the region of interest and the type of the region of interest.

[0075] The coefficients of the filters applied to each convolution layer of the learning model 22 before learning, the offset values, and the weights of the connections to the next layer in the fully connected layer are set to arbitrary initial values.

[0076] The error calculation unit 54 acquires the recognition result output from the output layer 52C of the learning model 22 and the teacher label S corresponding to the image frame N, and calculates the error between them. Possible methods for calculating the error include, for example, softmax cross entropy or least squared error (MSE). Note that, if sample weights are assigned to the teacher labels, the error calculation unit 54 calculates the error based on the sample weights.

[0077] The parameter update unit 56 adjusts the weight parameters of the learning model 22 using the error backpropagation method based on the error calculated by the error calculation unit 54.

[0078] This parameter adjustment process is repeated, and learning is repeated until the difference between the output of the learning model 22 and the teacher label S becomes small.

[0079] The learning control unit 20 uses a data set of at least image frames N and teacher labels S to optimize each parameter of the learning model 22. The learning by the learning control unit 20 may use a mini-batch method in which a certain number of data sets are extracted, batch processing of machine learning is performed using the extracted data sets, and this process is repeated.

[0080] As described above, in this embodiment, an image frame N to be used as learning data is determined, and a teacher label corresponding to the image frame N is generated based on the assigned diagnostic result. As a result, this aspect generates a teacher label by effectively using the assigned diagnostic result, and can perform effective machine learning based on the image frame N determined to be used as learning data and the teacher label.

[0081] <Third embodiment> Next, a third embodiment of the present invention will be described. In this embodiment, training data is determined, and a teacher label for an image frame N determined as training data is generated based on the distribution of recognition results from a plurality of recognizers.

[0082] 10 is a block diagram showing the main configuration of the image processing device 10 of this embodiment. Note that the same reference numerals are used for the parts that have already been explained, and explanations thereof will be omitted.

[0083] The image processing device 10 mainly comprises a first processor 1, a second processor (processor) 2, and a storage unit 11. The first processor 1 and the second processor 2 may be configured as the same CPU (or GPU), or may be configured as separate CPUs (or GPUs). The first processor 1 and the second processor 2 execute programs stored in the storage unit 11 to realize the functions shown in the functional blocks.

[0084] The first processor 1 is composed of a video acquisition unit 12, a recognition unit 14, and a learning availability determination unit 16. The second processor (processor) 2 is composed of a second teacher label generation unit 24, a learning control unit 20, and a learning model 22.

[0085] The second teacher label generation unit 24 generates teacher labels for machine learning based on the distribution of recognition results of the multiple recognizers that make up the recognition unit 14.

[0086] The second teacher label generation unit 24 can generate teacher labels for machine learning using various methods based on the distribution of recognition results from multiple recognizers. For example, the second teacher label generation unit 24 generates the label that is most frequently output in the recognition results (the majority label) as the teacher label. The second teacher label generation unit 24 may also use the average value of the scores that are the recognition results from the multiple recognizers as a pseudo-label. The second teacher label generation unit 24 can assign sample weights to the teacher labels it generates. The second teacher label generation unit 24 can change the sample weights assigned to the teacher labels depending on the variability of the recognition results. For example, the second teacher label generation unit 24 increases the sample weight as the variability of the recognition results decreases, and decreases the sample weight as the variability of the recognition results increases. If the variability of the recognition results is too large, the generated teacher labels do not need to be used for machine learning.

[0087] Fig. 11 is a diagram for explaining the learning use availability determining unit 16 and the second teacher label generating unit 24. Note that the same reference numerals are used to denote the parts that have already been explained in Fig. 4, and explanations thereof will be omitted.

[0088] The recognition unit 14 receives image frames N1 to N4 that are successive in time series.

[0089] 11 shows a case where image frame N3 is input to recognition unit 14. Note that image frame N3 is determined by learning use availability determining unit 16 to be used as learning data.

[0090] When image frame N3 is input to the recognition unit 14, recognition results 1 to 4 are output from the first recognizer 14A to the fourth recognizer 14D. When image frame N3 is input, the first recognizer 14A outputs recognition result 1 (label A). When image frame N3 is input, the second recognizer 14B outputs recognition result 2 (label A). When image frame N3 is input, the third recognizer 14C outputs recognition result 3 (label B). When image frame N4 is input, the fourth recognizer 14D outputs recognition result 4 (label A). Since recognition results 1 to 4 do not all match, the learning use determination unit 16 decides to use image frame N3 as learning data (image frame N3 is shown with a circle).

[0091] Furthermore, it is also determined that image frame N1 and image frame N4 will be used as learning data in the same manner as image frame N3 (image frame N1 and image frame N4 are indicated by "◯").

[0092] Furthermore, the second teacher label generation unit 24 generates a teacher label based on the distribution of the recognition results 1 to 4. Specifically, since the recognition result 1 is label A, the recognition result 2 is label A, the recognition result 3 is label B, and the recognition result 4 is label A, the recognition results are most widely distributed in the label A. Therefore, the second teacher label generation unit 24 generates the teacher label as label A. Note that the teacher label is also generated as label A for the image frames N1 and N4, similar to the image frame N3.

[0093] 12 shows a case where image frame N2 is input to recognition unit 14. Note that learning use availability determining unit 16 determines that image frame N2 will not be used as learning data.

[0094] When image frame N2 is input to the recognition unit 14, recognition results 1 to 4 are output from the first to fourth recognizers 14A to 14D. When image frame N2 is input, the first recognizer 14A outputs recognition result 1 (label A). When image frame N2 is input, the second recognizer 14B outputs recognition result 2 (label A). When image frame N2 is input, the third recognizer 14C outputs recognition result 3 (label A). When image frame N2 is input, the fourth recognizer 14D outputs recognition result 4 (label A). Because recognition results 1 to 4 all match, the learning use determination unit 16 determines not to use image frame N2 as learning data (image frame N2 is shown with an "x" attached).

[0095] In this embodiment, as described above, the learning frame N to be used as learning data is determined by the learning use determination unit 16. Also, as described above, a teacher label is generated by the second teacher label generation unit 24. Thereafter, as shown in FIG. 9, the learning frame N is input to the learning model 22, and the teacher label is input to the learning control unit 20. The learning control unit 20 The image frame N that has been determined to be used as learning data by the learning use determination unit 16 is input to the learning model 22. Furthermore, the learning control unit 20 is input with the teacher label S generated by the second teacher label generation unit 24. The learning control unit 20 uses a data set of at least the image frame N and the teacher label S to optimize each parameter of the learning model 22.

[0096] As described above, in this embodiment, an image frame N to be used as training data is determined, and a teacher label corresponding to the image frame N is generated based on the distribution of the recognition results. As a result, this aspect can generate a teacher label based on the recognition result even if a diagnosis result from a doctor or the like is not attached, and can perform effective machine learning based on the image frame N determined to be used as training data and the teacher label.

[0097] <Modification> Next, modified examples will be described. The following modified examples can be applied to the first to third embodiments described above.

[0098] <<Modification of the recognition unit>> The following describes modified examples of the recognition unit 14. Although an example of the recognition unit 14 has been described in Fig. 3, the present invention is not limited to this. The following describes modified examples of the recognition unit 14.

[0099] FIG. 13 is a diagram showing a modified example of the recognition unit 14. In FIG.

[0100] The recognition unit 14 is composed of a first recognizer 15A, a second recognizer 15B, a second recognizer 15C, and a second recognizer 15D. The first recognizer 15A is composed of an average trained model (recognition model) common to all countries that is directly used by users. The second recognizer 15B, the second recognizer 15C, and the second recognizer 15D are each composed of trained models trained with biased training data. By configuring the recognition unit 14 in this way, it is possible to determine the image frame N to be used as training data based on the average recognition result common to all countries and the biased recognition result.

[0101] <<Learning Use Decision Unit>> Next, a modified example of the learning use availability determining unit 16 will be described. The learning use availability determining unit 16 in the first to third embodiments determines whether or not to use the image frame N as learning data depending on the variation (distribution) of the recognition results of the first to fourth recognizers 14A to 14D for each image frame N. However, the learning use availability determining unit 16 is not limited to this. Below, a modified example of the learning use availability determining unit 16 will be described.

[0102] FIG. 14 is a diagram illustrating a modified example of the learning availability determining unit 16. In FIG.

[0103] In this example, multiple recognizers are made to perform a process of recognizing lesions on chronologically consecutive image frames, and chronologically consecutive recognition results are obtained from each of the multiple recognizers. Fig. 14 shows the recognition results when chronologically consecutive image frames N1 to N12 are input to each of the first to fourth recognizers 14A to 14D.

[0104] The learning use determination unit 16 determines whether or not to use the image frame for machine learning based on the recognition results of each of the plurality of time-series consecutive recognizers.

[0105] The first recognizer 14A outputs a recognition result α based on the input image frames N1 to N12. Specifically, the first recognizer 14A outputs a recognition result α for each of the image frames N1 to N12. Similarly to the first recognizer 14A, the third recognizer 14C and the fourth recognizer 14D also output the recognition result α based on the input image frames N1 to N12.

[0106] On the other hand, the second recognizer 14B outputs a recognition result α and a recognition result β for the input image frames N1 to N12. Specifically, the second recognizer 14B outputs the recognition result α when the image frame N1, image frames N5 to N8, and image frames N10 to N12 are input. Furthermore, the second recognizer 14B outputs the recognition result β when the image frames N2 to N4 and image frame N9 are input.

[0107] In this example, the learning availability determining unit 16 determines whether to use image frames as learning data, taking into consideration the chronologically consecutive recognition results. Specifically, the recognition result β continues for three image frames from image frames N2 to N4. Because the recognition results vary over a certain number of image frames (image frames N2 to N4), it can be assumed that this variation in the recognition results is not an error, and that image frames N2 to N4 are learning data that can be used for effective learning. Therefore, the learning availability determining unit 16 determines to use image frames N2 to N4 as learning data. On the other hand, the recognition results of the first to fourth recognizers 14A to 14D for the frames before and after image frame N9 (image frames N8 and N10) are all consistent, so it can be assumed that the variation in the recognition results for image frame N9 is an error. Therefore, the learning availability determining unit 16 determines not to use image frame N9 as learning data.

[0108] As described above, according to the learning use determination unit 16 of this example, whether or not to use an image frame N as learning data is determined based not only on the variation in the recognition results for each image frame N but also on the variation in the recognition results over time, so that learning data that can perform effective machine learning can be determined more efficiently.

[0109] <Overall configuration of the endoscope device> The inspection video M used in the technology of the present disclosure is acquired by an endoscope device (endoscope system) 500 described below, and then stored in a database DB. Note that the endoscope device 500 described below is an example and is not limited to this.

[0110] FIG. 15 is a diagram showing the overall configuration of an endoscope device 500. As shown in FIG.

[0111] The endoscope device 500 includes an endoscope body 100, a processor device 200, a light source device 300, and a display device 400. In addition, in the same drawing, a part of the hard tip portion 116 provided in the endoscope body 100 is illustrated in an enlarged manner.

[0112] The endoscope main body 100 includes a handheld control unit 102 and a scope 104. A user holds and operates the handheld control unit 102, and inserts the insertion unit (scope) 104 into the body of a subject to observe the inside of the subject's body. The term "user" is synonymous with a doctor, an operator, etc. The term "subject" used here is synonymous with a patient and an examinee.

[0113] The handheld operation unit 102 includes an air / water supply button 141, a suction button 142, a function button 143, and an image capture button 144. The air / water supply button 141 receives an air supply instruction and a water supply instruction.

[0114] The suction button 142 receives a suction instruction. Various functions are assigned to the function button 143. The function button 143 receives instructions for various functions. The image capture button 144 receives an image capture instruction operation. Image capture includes moving image capture and still image capture.

[0115] The scope (insertion section) 104 includes a flexible section 112, a bending section 114, and a tip rigid section 116. The flexible section 112, bending section 114, and tip rigid section 116 are arranged in this order from the side of the handheld operation section 102. That is, the bending section 114 is connected to the base end side of the tip rigid section 116, the flexible section 112 is connected to the base end side of the bending section 114, and the handheld operation section 102 is connected to the base end side of the scope 104.

[0116] The user can bend the bending section 114 by operating the hand operation section 102, thereby changing the orientation of the tip rigid section 116 up, down, left, or right. The tip rigid section 116 includes an imaging section, an illumination section, and a forceps port 126.

[0117] 15 illustrates a photographing lens 132 that constitutes the imaging unit. Also, the same figure illustrates an illumination lens 123A and an illumination lens 123B that constitute the illumination unit. The imaging unit is shown in FIG. 16 with the reference numeral 130. The illumination unit is shown in FIG. 16 with the reference numeral 123.

[0118] During observation and treatment, at least one of white light (normal light) and narrowband light (special light) is output via illumination lens 123A and illumination lens 123B in response to operation of operation unit 208 shown in FIG.

[0119] When the air / water supply button 141 is operated, cleaning water is released from the water supply nozzle, or gas is released from the air supply nozzle. The cleaning water and gas are used to clean the illumination lens 123A, etc. The water supply nozzle and air supply nozzle are not shown in the drawings. The water supply nozzle and the air supply nozzle may be a common nozzle.

[0120] The forceps port 126 communicates with the duct. A treatment tool is inserted into the duct. The treatment tool is supported so that it can move forward and backward as appropriate. When removing a tumor or the like, the treatment tool is applied to perform the necessary treatment. Note that reference numeral 106 in FIG. 15 denotes a universal cable. Reference numeral 108 denotes a light guide connector.

[0121] 16 is a functional block diagram of an endoscope device 500. The endoscope body 100 includes an imaging unit 130. The imaging unit 130 is disposed inside the rigid tip portion 116. The imaging unit 130 includes a photographing lens 132, an imaging element 134, a driving circuit 136, and an analog front end 138. AFE is an abbreviation for Analog Front End.

[0122] The photographing lens 132 is disposed on the tip end surface 116A of the tip rigid portion 116. An imaging element 134 is disposed on the opposite side of the photographing lens 132 from the tip end surface 116A. A CMOS type image sensor is applied as the imaging element 134. A CCD type image sensor may also be applied as the imaging element 134. Note that CMOS is an abbreviation for Complementary Metal-Oxide Semiconductor. CCD is an abbreviation for Charge Coupled Device.

[0123] A color image sensor is used as the image sensor 134. An example of a color image sensor is an image sensor equipped with color filters corresponding to RGB. RGB stands for red, green, and yellow, respectively.

[0124] A monochrome imaging element may be applied to the imaging element 134. When a monochrome imaging element is applied to the imaging element 134, the imaging unit 130 can switch the wavelength band of light incident on the imaging element 134 to perform frame sequential or color sequential imaging.

[0125] The drive circuit 136 supplies various timing signals required for the operation of the image sensor 134 to the image sensor 134 based on control signals sent from the processor unit 200 .

[0126] The analog front end 138 includes an amplifier, a filter, and an AD converter. AD is an abbreviation of analog and digital, respectively. The analog front end 138 performs processing such as amplification, noise removal, and analog-to-digital conversion on the output signal of the image sensor 134. The output signal of the analog front end 138 is transmitted to the processor device 200. AFE shown in FIG. 16 is an abbreviation for Analog Front End, which is the English term for analog front end.

[0127] An optical image of the object to be observed is formed on the light receiving surface of the image sensor 134 via the photographing lens 132. The image sensor 134 converts the optical image of the object to an electrical signal. The electrical signal output from the image sensor 134 is transmitted to the processor device 200 via a signal line.

[0128] The illumination unit 123 is disposed on the distal end hard portion 116. The illumination unit 123 includes an illumination lens 123A and an illumination lens 123B. The illumination lenses 123A and 123B are disposed adjacent to the photographing lens 132 on the distal end surface 116A.

[0129] The illumination unit 123 includes a light guide 170. The light guide 170 has an exit end located on the opposite side to the tip end surfaces 116A of the illumination lenses 123A and 123B.

[0130] The light guide 170 is inserted into the scope 104, the handheld operation unit 102, and the universal cable 106 shown in FIG.

[0131] The processor unit 200 includes an image input controller 202, an image signal processing unit 204, and a video output unit 206. The image input controller 202 acquires an electrical signal transmitted from the endoscope body 100 and corresponding to an optical image of the observation target.

[0132] The imaging signal processing unit 204 generates an endoscopic image of the observation object and an inspection moving image M based on the imaging signal, which is an electrical signal corresponding to an optical image of the observation object.

[0133] The imaging signal processing unit 204 may perform image quality correction by applying digital signal processing such as white balance processing and shading correction processing to the imaging signal. The imaging signal processing unit 204 may also add supplementary information defined by the DICOM standard to the image frames constituting the endoscopic image or the examination video M. DICOM is an abbreviation for Digital Imaging and Communications in Medicine.

[0134] The video output unit 206 transmits a display signal representing the image generated by the imaging signal processing unit 204 to the display device 400. The display device 400 displays the image of the object to be observed.

[0135] When the imaging button 144 shown in FIG. 15 is operated, the processor device 200 operates the image input controller 202, the imaging signal processing unit 204, etc. in response to an imaging command signal transmitted from the endoscope body 100.

[0136] When the processor device 200 receives a freeze command signal representing the capture of a still image from the endoscope body 100, it applies the image capture signal processing unit 204 to generate a still image based on a frame image at the operation timing of the image capture button 144. The processor device 200 causes the display device 400 to display the still image.

[0137] The processor device 200 includes a communication control unit 205. The communication control unit 205 controls communication with devices communicably connected via an in-hospital system and an in-hospital LAN, etc. The communication control unit 205 can apply a communication protocol conforming to the DICOM standard. An example of an in-hospital system is an HIS (Hospital Information System). LAN is an abbreviation for Local Area Network.

[0138] The processor device 200 includes a storage unit 207. The storage unit 207 stores endoscopic images and inspection videos M generated using the endoscope main body 100. The storage unit 207 may store various information associated with the endoscopic images and inspection videos M. Specifically, the storage unit 207 stores operation information such as an operation log for capturing the endoscopic images and inspection videos M. The operation information such as the endoscopic images, inspection videos M, and operation logs stored in the storage unit 207 is saved in a database DB.

[0139] The processor device 200 includes an operation unit 208. The operation unit 208 outputs a command signal in response to an operation by a user. The operation unit 208 may be a keyboard, a mouse, a joystick, or the like.

[0140] The processor device 200 includes an audio processing unit 209 and a speaker 209A. The audio processing unit 209 generates an audio signal representing information to be notified as audio. The speaker 209A converts the audio signal generated by the audio processing unit 209 into audio. Examples of audio output from the speaker 209A include messages, audio guidance, and warning sounds.

[0141] The processor device 200 includes a CPU 210, a ROM 211, and a RAM 212. Note that ROM is an abbreviation for Read Only Memory, and RAM is an abbreviation for Random Access Memory.

[0142] The CPU 210 functions as an overall control unit of the processor device 200. The CPU 210 functions as a memory controller that controls the ROM 211 and RAM 212. ROM The section 211 stores various programs and control parameters that are applied to the processor device 200.

[0143] The RAM 212 is used as a temporary storage area for data in various processes and as a processing area for arithmetic processing using the CPU 210. The RAM 212 can be used as a buffer memory when an endoscopic image is acquired.

[0144] <<Hardware configuration of the processor unit>> A computer may be used as the processor device 200. The computer may use the following hardware and execute a specified program to realize the functions of the processor device 200. Note that the term "program" is synonymous with software.

[0145] The processor device 200 may employ various processors as a signal processing unit that performs signal processing. Examples of processors include a CPU and a GPU (Graphics Processing Unit). A CPU is a general-purpose processor that executes programs and functions as a signal processing unit. A GPU is a processor specialized for image processing. The processor hardware employs an electric circuit that combines electric circuit elements such as semiconductor elements. Each control unit includes a ROM in which programs and the like are stored and a RAM that serves as a working area for various calculations.

[0146] Two or more processors may be applied to one signal processing unit. The two or more processors may be the same type of processor or different types of processors. Also, one processor may be applied to multiple signal processing units. The processor device 200 described in the embodiment corresponds to an example of an endoscope control unit.

[0147] <<Example of light source device configuration>> The light source device 300 includes a light source 310, an aperture 330, a condenser lens 340, and a light source control unit 350. The light source device 300 causes observation light to enter the light guide 170. The light source 310 includes a red light source 310R, a green light source 310G, and a blue light source 310B. The red light source 310R, the green light source 310G, and the blue light source 310B emit narrowband light of red, green, and blue, respectively.

[0148] The light source 310 can generate illumination light by combining any of the narrowband lights of red, green, and blue. For example, the light source 310 can generate white light by combining the narrowband lights of red, green, and blue. The light source 310 can also generate narrowband light by combining any two of the narrowband lights of red, green, and blue. Here, white light is the light used in normal endoscopic examinations and is referred to as normal light, while narrowband light is referred to as special light.

[0149] The light source 310 may generate narrow-band light using any one of red, green, and blue narrow-band light. The light source 310 may selectively emit white light or narrow-band light. The light source 310 may include an infrared light source that emits infrared light and an ultraviolet light source that emits ultraviolet light.

[0150] The light source 310 may have an embodiment including a white light source that emits white light, a filter that passes white light, and a filter that passes narrow-band light. The light source 310 in this embodiment can selectively emit either white light or narrow-band light by switching between the filter that passes white light and the filter that passes narrow-band light.

[0151] The filter that passes the narrow-band light may include a plurality of filters corresponding to different bands, and the light source 310 may selectively switch among the plurality of filters corresponding to different bands to selectively emit a plurality of narrow-band lights having different bands.

[0152] The type and wavelength band of the light source 310 can be adapted according to the type of object to be observed and the purpose of the observation. Examples of the type of light source 310 include a laser light source, a xenon light source, and an LED light source. LED is an abbreviation for Light-Emitting Diode.

[0153] When the light guide connector 108 is connected to the light source device 300, the observation light emitted from the light source 310 passes through the aperture 330 and the condenser lens 340 and reaches the incident end of the light guide 170. The observation light passes through the light guide 170 and the illumination lens 123A, etc., and is irradiated onto the observation object.

[0154] The light source control unit 350 transmits control signals to the light source 310 and the diaphragm 330 based on a command signal transmitted from the processor device 200. The light source control unit 350 controls the illuminance of the observation light emitted from the light source 310, switching of the observation light, and turning the observation light on and off.

[0155] <<Change light source>> The endoscope device 500 can use, as a light source, white light or normal light obtained by irradiating light of multiple wavelength bands as white light. On the other hand, the endoscope device 500 can also irradiate light of a specific wavelength band (special light). Specific examples of the specific wavelength band will be described below.

[0156] <<Example 1>> A first example of the specific wavelength band is the blue or green band of the visible range. The first example wavelength band includes a wavelength band from 390 to 450 nanometers or from 530 to 550 nanometers, and the first example light has a peak wavelength within the wavelength band from 390 to 450 nanometers or from 530 to 550 nanometers.

[0157] <<Second Example>> A second example of the specific wavelength band is the red band of the visible range, and the wavelength band of the second example includes a wavelength band from 585 to 615 nanometers or from 610 to 730 nanometers, and the light of the second example has a peak wavelength within the wavelength band from 585 to 615 nanometers or from 610 to 730 nanometers.

[0158] <<Example 3>> A third example of the specific wavelength band includes a wavelength band where the absorption coefficients of oxygenated hemoglobin and reduced hemoglobin differ, and the light of the third example has a peak wavelength in the wavelength band where the absorption coefficients of oxygenated hemoglobin and reduced hemoglobin differ. The wavelength band of this third example includes a wavelength band of 400±10 nanometers, 440±10 nanometers, 470±10 nanometers, or a wavelength band of 600 nanometers to 750 nanometers, and the light of the third example has a peak wavelength in the wavelength band of 400±10 nanometers, 440±10 nanometers, 470±10 nanometers, or a wavelength band of 600 nanometers to 750 nanometers.

[0159] <<Example 4>> A fourth example of a specific wavelength band is the wavelength band of excitation light used to observe fluorescence emitted by fluorescent substances in living organisms and to excite these fluorescent substances. For example, the wavelength band is from 390 nanometers to 470 nanometers. Note that fluorescence observation is sometimes called fluorescence observation.

[0160] <<Example 5>> A fifth example of the specific wavelength band is the wavelength band of infrared light, which includes a wavelength band of 790 to 820 nanometers or a wavelength band of 905 to 970 nanometers, and the light in this fifth example has a peak wavelength in the wavelength band of 790 to 820 nanometers or a wavelength band of 905 to 970 nanometers.

[0161] <<Example of special light image generation>> The processor unit 200 may generate a special light image having information of a specific wavelength band based on a normal light image captured using white light. Note that "generate" here includes "acquire." In this case, the processor unit 200 functions as a special light image acquisition unit. The processor unit 200 then acquires the signal of the specific wavelength band by performing a calculation based on the color information of red, green, and blue, or cyan, magenta, and yellow, contained in the normal light image. Note that cyan, magenta, and yellow are sometimes abbreviated as CMY, using the initials of their English names: Cyan, Magenta, and Yellow.

[0162] <Other> In the above embodiment, the hardware structure of the processing units (first processor 1 and second processor 2) that execute various processes is the following various processors: The various processors include a CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) and functions as various processing units, a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacture, and a dedicated electrical circuit, such as an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing specific processes.

[0163] The first processor 1 and / or the second processor 2 may be configured with one of these various processors, or may be configured with two or more processors of the same or different types (for example, multiple FPGAs, or a combination of a CPU and an FPGA). Furthermore, multiple processing units may be configured with a single processor. Examples of multiple processing units configured with a single processor include, first, a configuration in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server, and this processor functions as multiple processing units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a system-on-chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.

[0164] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.

[0165] The above-described configurations and functions can be realized by any hardware, software, or a combination of both. For example, the present invention can be applied to a program that causes a computer to execute the above-described processing steps (processing procedures), a computer-readable recording medium (non-transitory recording medium) on which such a program is recorded, or a computer on which such a program can be installed.

[0166] <Other> In the above embodiment, the hardware structure of the processing units (first processor 1 and second processor 2) that execute various processes is the following various processors: The various processors include a CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) and functions as various processing units, a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacture, and a dedicated electrical circuit, such as an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing specific processes.

[0167] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types (for example, multiple FPGAs, or a combination of a CPU and an FPGA). Also, multiple processing units may be configured with a single processor. Examples of multiple processing units configured with a single processor include, first, a configuration in which one processor is configured with a combination of one or more CPUs and software, as typified by client or server computers, and this processor functions as multiple processing units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a System on Chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.

[0168] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.

[0169] The above-described configurations and functions can be realized by any hardware, software, or a combination of both. For example, the present invention can be applied to a program that causes a computer to execute the above-described processing steps (processing procedures), a computer-readable recording medium (non-transitory recording medium) on which such a program is recorded, or a computer on which such a program can be installed.

[0170] Although examples of the present invention have been described above, it goes without saying that the present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the present invention. [Explanation of symbols]

[0171] 1: First processor 2: Second processor 10: Image processing device 11: Storage section 12: Video acquisition unit 14:Recognition part 14A: 1st recognizer 14B:Second recognizer 14C: 3rd recognizer 14D: 4th recognizer 16: Learning use decision unit 18: First teacher label generation unit 20: Learning control unit 22: Learning model 24: Second teacher label generation unit

Claims

1. An image processing device comprising a processor and a plurality of recognizers, The processor: Obtaining video captured by medical equipment, causing the plurality of recognizers to perform a process of recognizing lesions on image frames constituting the video, and obtaining recognition results from each of the plurality of recognizers; When the recognition results of the plurality of recognizers do not match, the image frame is used as learning data for machine learning, and when the recognition results of the plurality of recognizers match, the image frame is not used as learning data for machine learning. Image processing device.

2. The image processing device according to claim 1 , wherein the plurality of recognizers are different from each other in at least one of structure, type, and parameters of the recognizers.

3. The image processing device according to claim 1 or 2, wherein the plurality of recognizers are trained using different training data.

4. The image processing device according to claim 3 , wherein the plurality of recognizers are each subjected to machine learning using the different training data obtained by different medical devices.

5. The image processing device according to claim 4 , wherein the plurality of recognizers are each subjected to machine learning using the different learning data obtained at facilities in different countries or regions.

6. The image processing device according to claim 3 , wherein the plurality of recognizers each perform machine learning using the different learning data captured under different imaging conditions.

7. The image processing device according to claim 1 , wherein when the processor determines that an image frame to which a diagnostic result has been assigned is training data, the processor generates a teacher label for the training data based on the diagnostic result.

8. The image processing device according to claim 1 , wherein the learning data determined by the processor is used to train a learning model that performs the machine learning.

9. The image processing device according to claim 8 , wherein the processor causes the learning model to learn the training data with sample weights determined based on a distribution of the recognition results of each of the plurality of recognizers.

10. The image processing device according to claim 1 , wherein the processor generates a teacher label for the machine learning based on a distribution of the recognition results.

11. The image processing device according to claim 10 , wherein the processor changes sample weights in the machine learning in accordance with the magnitude of variation in the recognition results.

12. The processor: causing the plurality of recognizers to perform a process of recognizing lesions on the time-series consecutive image frames, and obtaining the recognition results of each of the plurality of recognizers; The image processing device according to claim 1 , further comprising: determining whether or not to use the image frame in the machine learning process based on the recognition results of the plurality of time-series consecutive recognizers.

13. 13. The image processing device according to claim 1, wherein at least one of the plurality of recognizers outputs the recognition result while the video is being acquired, and the other recognizers output the recognition result after a first time has elapsed after the video is acquired.

14. An image processing method for an image processing device having a processor and a plurality of recognizers, the processor: acquiring a video captured by a medical device; causing the plurality of recognizers to perform a process of recognizing lesions on image frames constituting the video, and acquiring recognition results from each of the plurality of recognizers; a step of using the image frame as training data to be used for machine learning when the recognition results of the plurality of recognizers do not match, and not using the image frame as training data to be used for machine learning when the recognition results of the plurality of recognizers match; An image processing method that performs the above.

15. A program for executing an image processing method of an image processing device including a processor and a plurality of recognizers, the processor, acquiring a video captured by a medical device; causing the plurality of recognizers to perform a process of recognizing lesions on image frames constituting the video, and acquiring recognition results from each of the plurality of recognizers; a step of using the image frame as training data to be used for machine learning when the recognition results of the plurality of recognizers do not match, and not using the image frame as training data to be used for machine learning when the recognition results of the plurality of recognizers match; A program that performs the following.

Citation Information

Patent Citations

  • Fortran translating system

    JP1986082242A

  • Advanced computer-aided diagnosis of pulmonary nodules

    JP2010504129A

  • Learning support device, operation method for learning support device, learning support program, learning support system, and terminal device

    JP2019061579A

  • Information processing apparatus and information processing method

    JP2019109553A

  • Method for domain adaptation based on adversarial learning and apparatus thereof

    US20200321118A1