Processing system, image processing method, learning method, and processing device
By using machine learning models to detect regions of interest in endoscopic images and automatically adjusting control parameters, the problem of low estimation probability is solved, and highly reliable auxiliary information is provided.
Patent Information
- Application Number
- CN202080098316.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-11
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2040-03-11
AI Technical Summary
Existing technologies in endoscopic image diagnosis fail to provide reliable auxiliary information when the estimated probability is low, requiring users to manually adjust parameters such as light sources to improve image quality.
The learned model generated by machine learning detects regions of interest in endoscopic images and calculates estimated probabilities. It automatically adjusts the control parameters of the endoscopic device to improve the estimated probabilities, including the control of the light source, camera element, and image processing.
It improves the estimation probability of regions of interest in endoscopic images, reduces the user burden, and provides highly reliable auxiliary information.
Smart Images

Figure CN115279247B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to processing systems, image processing methods, learning methods, and processing devices. Background Technology
[0002] Previously, in image diagnostic aids for assisting doctors in diagnosing endoscopic images, methods for detecting lesions using machine learning and methods for estimating probabilities representing the likelihood of detection are known. Neural networks are known as learned models generated through machine learning. For example, Patent Document 1 discloses a system that estimates the name and location of lesions and their probability information based on a convolutional neural network (CNN), and overlays the estimated information onto the endoscopic image to provide auxiliary information.
[0003] Existing technical documents
[0004] Patent documents
[0005] Patent Document 1: International Publication No. 2019 / 088121 Summary of the Invention
[0006] The problem that the invention aims to solve
[0007] In existing methods such as Patent Document 1, countermeasures such as not displaying auxiliary information are taken when the estimated probability is low. However, since no countermeasures to improve the estimated probability are included, auxiliary information with sufficient reliability is sometimes not provided.
[0008] According to several embodiments of this application, a processing system, an image processing method, and a learning method that can obtain highly reliable information can be provided.
[0009] Methods for solving problems
[0010] One aspect of this application relates to a processing system comprising: an acquisition unit that acquires an image of a detection object captured by an endoscope; a control unit that controls the endoscope according to control information; and a processing unit that, based on the detection object image and a learned model obtained through machine learning for calculating estimated probability information representing the probability of a region of interest within an input image, detects the region of interest contained in the detection object image, calculates the estimated probability information associated with the detected region of interest, the processing unit determines, based on the detection object image, control information for improving the estimated probability information associated with the region of interest within the detection object image, and the control unit controls the endoscope according to the determined control information.
[0011] Other aspects of this application relate to an image processing method that acquires a detection object image captured by an endoscope device, detects the region of interest contained in the detection object image based on a learned model obtained through machine learning for calculating estimated probability information representing the probability of a region of interest within an input image and the detection object image, calculates the estimated probability information associated with the detected region of interest, and, when information for controlling the endoscope device is set as control information, determines control information for improving the estimated probability information associated with the region of interest within the detection object image based on the detection object image.
[0012] Another aspect of this application relates to a learning method that takes an image captured by an endoscope as an input image, and when information for controlling the endoscope is set as control information, takes first control information as the control information when the input image is acquired, and takes second control information as the control information for improving the estimated probability information, the estimated probability information representing the probability of a region of interest detected from the input image, and generates a learned model by performing machine learning on the relationship between the input image, the first control information, and the second control information. Attached Figure Description
[0013] Figure 1 This is a structural example of a processing system.
[0014] Figure 2 This is an external view of the endoscope system.
[0015] Figure 3 This is a structural example of an endoscope system.
[0016] Figure 4 This is an example of controlling information.
[0017] Figure 5 This is an example of the structure of a learning device.
[0018] Figure 6 (A) Figure 6 (B) is a structural example of a neural network.
[0019] Figure 7 (A) is an example of the training data used by NN1. Figure 7 (B) is an example of the input and output of NN1.
[0020] Figure 8 This is a flowchart illustrating the learning process of NN1.
[0021] Figure 9 (A) is an example of the training data used by NN2. Figure 9(B) is a data example used to obtain training data. Figure 9 (C) is an example of the input and output of NN2.
[0022] Figure 10 This is a flowchart illustrating the learning process of NN2.
[0023] Figure 11 This is a flowchart illustrating the detection process and the determination and processing of control information.
[0024] Figure 12 This is an example of the time point at which the displayed image and the image of the detected object are acquired.
[0025] Figure 13 (A) Figure 13 (B) is an example of a display screen.
[0026] Figure 14 (A) is an example of the training data used by NN2. Figure 14 (B) is a data example used to obtain training data. Figure 14 (C) is an example of the input and output of NN2.
[0027] Figure 15 This is a flowchart illustrating the detection process and the determination and processing of control information.
[0028] Figure 16 (A) is an example of the training data used by NN2. Figure 16 (B) is an example of the input and output of NN2.
[0029] Figure 17 This is a flowchart illustrating the detection process and the determination and processing of control information. Detailed Implementation
[0030] In the following disclosure, various different implementations and embodiments for carrying out different features of the presented subject matter are provided. These are, of course, merely examples and are not intended to be limiting. Furthermore, in this application, reference numbers and / or characters are sometimes repeated in various examples. Such repetition is for the sake of brevity and clarity, and does not itself need to be related to the various implementations and / or the structures described. Moreover, when described as a first element and a second element "connected" or "linked," such description includes implementations where the first and second elements are directly connected or linked to each other, and also includes implementations where the first and second elements are indirectly connected or linked to each other by having one or more other elements in between.
[0031] 1. System Structure
[0032] In prior systems such as Patent Document 1, the estimated probability, representing the likelihood of using the learned model, is displayed. Furthermore, when the estimated probability is low, the estimation result is not displayed to suppress the indication of low accuracy to the user. However, existing methods do not employ measures to improve the estimated probability of the learned model.
[0033] For example, consider a case where the estimated probability is low because the lesion is captured too darkly in the image. In this case, by increasing the target dimming value in the dimming process to capture a brighter image of the lesion, the estimated probability could potentially be improved. However, existing methods only display the estimated probability, or omit the display altogether due to a low estimated probability. Therefore, to control the light source, the user needs to determine whether a change in the light intensity is necessary based on the displayed image, and then perform specific operations to change the light intensity based on that determination. In other words, existing systems only provide the processing results for the input image; whether an image optimal for lesion detection can be captured depends on the user.
[0034] Figure 1 This diagram illustrates the structure of the processing system 100 according to this embodiment. The processing system 100 includes an acquisition unit 110, a processing unit 120, and a control unit 130.
[0035] The acquisition unit 110 is an interface for acquiring images. For example, the acquisition unit 110 is an interface circuit that acquires signals from the imaging element 312 via signal lines included in the universal cable 310c. Alternatively, the acquisition unit 110 may also include a circuit using… Figure 3 The preprocessing unit 331 will be described later. Alternatively, the processing system 100 may also be included in an information processing device that acquires the image signal output from the lens unit 310 via a network. In this case, the acquisition unit 110 is a communication interface such as a communication chip.
[0036] The processing unit 120 and the control unit 130 are configured with the following hardware. The hardware may include at least one of circuitry for processing digital signals and circuitry for processing analog signals. For example, the hardware may consist of one or more circuit devices or circuit elements mounted on a circuit board. The one or more circuit devices may be, for example, an integrated circuit (IC), an FPGA (field-programmable gate array), etc. The one or more circuit elements may be, for example, resistors, capacitors, etc.
[0037] Alternatively, the processing unit 120 and the control unit 130 can also be implemented using the processor described below. The processing system 100 includes a memory for storing information and a processor that performs operations based on the information stored in the memory. The information includes, for example, programs and various types of data. The processor includes hardware. The processor can be various types of processors such as CPU (Central Processing Unit), GPU (Graphics Processing Unit), and DSP (Digital Signal Processor). The memory can be a semiconductor memory such as SRAM (Static Random Access Memory) or DRAM (Dynamic Random Access Memory), or it can be a register, a magnetic storage device such as HDD (Hard Disk Drive), or an optical storage device such as an optical disc drive. For example, the memory stores commands that can be read by a computer, and the processor executes these commands to perform the functions of the processing unit 120 and the control unit 130. Furthermore, the commands here can be commands that constitute a program's command set, or commands that instruct the processor's hardware circuitry to perform actions. Furthermore, the processing unit 120 and the control unit 130 can be implemented by a single processor or by different processors. Alternatively, the functions of the processing unit 120 can be implemented through distributed processing using multiple processors. The same applies to the control unit 130.
[0038] The acquisition unit 110 acquires images of the object being inspected, captured by the endoscopic device. The endoscopic device used here is, for example, a... Figure 2 Part or all of the endoscope system 300 described later. The endoscope device includes, for example, a portion of the endoscope body 310, the light source device 350, and the processing device 330 in the endoscope system 300.
[0039] The control unit 130 controls the endoscope device based on control information. (For example, when using...) Figure 4 As will be described later, the control unit 130 controls the endoscope device, including the control of the light source 352 included in the light source device 350, the control of the shooting conditions using the imaging element 312, and the control of image processing in the processing device 330.
[0040] The processing unit 120 detects regions of interest contained in the target image based on the learned model and the target image, and calculates estimated probability information representing the likelihood of the detected regions of interest. Furthermore, the processing unit 120 in this embodiment determines control information for improving the estimated probability information related to the regions of interest within the target image based on the target image. The control unit 130 controls the endoscope device based on the determined control information.
[0041] The learned model here is a model obtained through machine learning, used to calculate the estimated probability information of regions of interest within an input image. More specifically, the learned model performs machine learning based on a dataset that maps the input image to information identifying the regions of interest contained within that input image. Given an input image, the learned model outputs the detection result of the region of interest in that image and the estimated probability information as the probability of that detection result. The following example illustrates that the estimated probability information is the estimated probability itself, but the estimated probability information can be any information that serves as an indicator of the estimated probability, or it can be information different from the estimated probability.
[0042] Furthermore, in this embodiment, the area of interest refers to the area that has a relatively higher priority for observation than other areas for the user. If the user is a doctor performing diagnosis or treatment, the area of interest corresponds, for example, to the area where the lesion is photographed. However, if the doctor wants to observe blisters or feces, the area of interest could also be the area where the blisters or feces were photographed. That is, the object of the user's attention varies depending on the purpose of the observation, but when conducting the observation, the area that has a relatively higher priority for observation than other areas for the user becomes the area of interest.
[0043] The processing system 100 of this embodiment not only performs the processing of calculating the estimated probability, but also performs the processing of determining control information for improving the estimated probability. For example, using... Figure 4 As will be described later, the control information includes information on various parameter categories. For example, the processing unit 120 determines which parameter category can be changed to improve the estimated probability, or what value to set the parameter of that category.
[0044] The control unit 130 performs control using the control information determined by the processing unit 120. Therefore, in the new image of the detected object obtained as a result of this control, the estimated probability of the region of interest can be expected to increase compared to before the control information was changed. In other words, the processing system 100 of this embodiment can perform cause analysis and control condition changes for the system itself to improve the estimated probability. Therefore, it is possible to suppress the increase in user burden and provide information with high reliability.
[0045] Figure 2This diagram illustrates the structure of an endoscope system 300 including a processing system 100. The endoscope system 300 includes a scope body 310, a processing unit 330, a display unit 340, and a light source device 350. For example, the processing system 100 is included in the processing unit 330. A physician uses the endoscope system 300 to perform endoscopic examinations on patients. However, the structure of the endoscope system 300 is not limited to... Figure 2 It can be modified in various ways, such as omitting some structural elements or adding other structural elements. Furthermore, while the following example illustrates a flexible endoscope used for diagnosing digestive organs, the endoscope body 310 of this embodiment can also be a rigid endoscope used for laparoscopic surgery. Additionally, the endoscope system 300 is not limited to a medical endoscope for observing living organisms; it can also be an industrial endoscope.
[0046] In addition, Figure 2 The diagram shows an example of a device where the processing unit 330 is connected to the mirror body 310 via connector 310d, but it is not limited to this. For example, part or all of the structure of the processing unit 330 may also be constructed from other information processing devices such as PCs (Personal Computers) and server systems that can be connected via a network. For example, the processing unit 330 may also be implemented via cloud computing. The network here may be a private network such as an intranet or a public communication network such as the Internet. In addition, the network may be wired or wireless. That is, the processing system 100 is not limited to the structure contained in the device connected to the mirror body 310 via connector 310d; part or all of its functions may be implemented by other devices such as PCs or by cloud computing.
[0047] The mirror body 310 includes an operation section 310a, a flexible insertion section 310b, and a universal cable 310c containing signal lines, etc. The mirror body 310 is a tubular insertion device into which the tubular insertion section 310b is inserted into a body cavity. A connector 310d is provided at the front end of the universal cable 310c. The mirror body 310 is connected to the light source device 350 and the processing device 330 in a detachable manner via the connector 310d. Furthermore, when using... Figure 3 As will be described later, a light guide 315 is inserted into the universal cable 310c, and the mirror body 310 allows the illumination light from the light source device 350 to be emitted from the front end of the insertion part 310b through the light guide 315.
[0048] For example, the insertion portion 310b has a front end, a bendable curved portion, and a flexible tube portion extending from its front end toward its base end. The insertion portion 310b is inserted into the subject. The front end of the insertion portion 310b is the front end of the lens body portion 310, and is a relatively rigid front end portion. The objective lens optical system 311 and the imaging element 312, described later, are provided, for example, at the front end.
[0049] The bending section can bend in a desired direction according to the operation of the bending operation component provided on the operation unit 310a. The bending operation component includes, for example, a left-right bending operation knob and a right-up-down bending operation knob. In addition, the operation unit 310a may also be equipped with various operation buttons such as a release button and an air / water supply button, in addition to the bending operation component.
[0050] The processing unit 330 is a video processor that performs prescribed image processing on the received capture signal to generate a captured image. The image signal of the generated captured image is output from the processing unit 330 to the display unit 340, and the live captured image is displayed on the display unit 340. The structure of the processing unit 330 will be described later. The display unit 340 is, for example, a liquid crystal display, an EL (Electro-Luminescence) display, etc.
[0051] The light source device 350 is a light source device capable of emitting normal light for normal light observation mode. In addition, when the endoscope system 300 has a special light observation mode in addition to the normal light observation mode, the light source device 350 selectively emits normal light for normal light observation mode and special light for special light observation mode.
[0052] Figure 3 This is a diagram illustrating the structure of each part of the endoscope system 300. Additionally, in Figure 3 In this paper, a portion of the structure of the mirror body 310 is omitted or simplified.
[0053] The light source device 350 includes a light source 352 that emits illumination light. The light source 352 can be a xenon light source, an LED (light emitting diode), or a laser light source. Alternatively, the light source 352 can be other light sources, and the method of light emission is not limited.
[0054] The insertion unit 310b includes an objective lens optical system 311, an imaging element 312, an illumination lens 314, and a light guide 315. The light guide 315 guides illumination light from the light source 352 to the front end of the insertion unit 310b. The illumination lens 314 illuminates the subject with the illumination light guided by the light guide 315. The objective lens optical system 311 images the reflected light from the subject into an image of the subject. The objective lens optical system 311 may include, for example, a focusing lens, and the position of the image formed can be changed according to the position of the focusing lens. For example, the insertion unit 310b may also include an actuator (not shown) that drives the focusing lens based on control from the control unit 332. In this case, the control unit 332 performs AF (Auto Focus) control.
[0055] The imaging element 312 receives light from the subject via the objective lens optical system 311. The imaging element 312 can be a monochrome sensor or an element equipped with a color filter. The color filter can be a well-known Bayer filter, a complementary color filter, or other filters. Complementary color filters include filters for cyan, magenta, and yellow.
[0056] The processing device 330 performs image processing and overall system control. The processing device 330 includes a preprocessing unit 331, a control unit 332, a storage unit 333, a detection processing unit 334, a control information determination unit 335, and a post-processing unit 336. For example, the preprocessing unit 331 corresponds to the acquisition unit 110 of the processing system 100. The detection processing unit 334 and the control information determination unit 335 correspond to the processing unit 120 of the processing system 100. The control unit 332 corresponds to the control unit 130 of the processing system 100.
[0057] The preprocessing unit 331 performs A / D conversion, converting the analog signals sequentially output from the imaging element 312 into digital images, and various correction processes on the A / D converted image data. Alternatively, an A / D conversion circuit can be provided in the imaging element 312, omitting the A / D conversion in the preprocessing unit 331. The correction processes here include, for example, using... Figure 4 As described later, this includes color matrix correction processing, structure enhancement processing, noise reduction processing, and AGC (automatic gain control). Additionally, the preprocessing unit 331 can also perform other correction processing such as white balance processing. The preprocessing unit 331 outputs the processed image as the detection target image to the detection processing unit 334 and the control information determination unit 335. Furthermore, the preprocessing unit 331 outputs the processed image as the display image to the post-processing unit 336.
[0058] The detection processing unit 334 performs detection processing to detect regions of interest from the image of the object to be detected. Furthermore, the detection processing unit 334 outputs an estimated probability representing the likelihood of the detected regions of interest. For example, the detection processing unit 334 performs detection processing according to information from a learned model stored in the storage unit 333.
[0059] Furthermore, the region of interest in this embodiment can be of one type. For example, the region of interest can be a polyp, and the detection process can be the process of determining the location and size of the polyp in the image of the object to be detected. Alternatively, the region of interest in this embodiment can also include multiple types. For example, methods for classifying polyps according to their state into TYPE1, TYPE2A, TYPE2B, and TYPE3 are known. The detection process in this embodiment can not only detect the location and size of the polyp, but may also include the process of classifying the polyp into which of the above categories. In this case, the estimated probability represents information about the location, size, and probability of the classification result of the region of interest.
[0060] The control information determination unit 335 performs processing based on the image of the detected object to determine control information for improving the estimated probability. Details of the processing in the detection processing unit 334 and the control information determination unit 335 will be described later.
[0061] The post-processing unit 336 performs post-processing based on the detection results of the region of interest in the detection processing unit 334, and outputs the post-processed image to the display unit 340. This post-processing, for example, involves appending the detection results based on the detected object image to the displayed image. For example, using... Figure 12 As described later, the display image and the detected object image are acquired alternately. Additionally, if using... Figure 13 (A) and Figure 13 As described later in (B), the detection result of the region of interest narrowly refers to information used to assist the user in diagnosis and other procedures. Therefore, the detection object image in this embodiment can also be referred to as an auxiliary image used in assisting the user. Furthermore, the detection result of the region of interest can also be referred to as auxiliary information.
[0062] The control unit 332 is interconnected with the imaging element 312, the preprocessing unit 331, the detection processing unit 334, the control information determination unit 335, and the light source 352, and controls each part. Specifically, the control unit 332 controls each part of the endoscope system 300 according to the control information.
[0063] Figure 4 This is a diagram illustrating an example of control information. The control information includes light source control information for controlling the light source 352, shooting control information for controlling the shooting conditions in the imaging element 312, and image processing control information for controlling the image processing in the processing device 330.
[0064] Light source control information includes parameters such as wavelength, light intensity ratio, light quantity, duty cycle, and light distribution as parameter categories. Shooting control information includes the shooting frame rate. Image processing control information includes color matrix, structure enhancement, noise reduction, and AGC. Control information is, for example, a set of parameter categories and specific parameter values associated with those categories. Specific examples of each parameter category and parameter value are explained below.
[0065] Wavelength indicates the band of illumination light. For example, the light source device 350 can illuminate both ordinary light and special light. For example, the light source device 350 includes multiple light sources such as red LEDs, green LEDs, blue LEDs, green narrowband LEDs, and blue narrowband LEDs. The light source device 350 illuminates ordinary light, including R-light, G-light, and B-light, by illuminating the red LEDs, green LEDs, and blue LEDs. For example, the wavelength of B-light is 430nm–500nm, the wavelength of G-light is 500nm–600nm, and the wavelength of R-light is 600nm–700nm. Additionally, the light source device 350 illuminates special light, including G2-light and B2-light, by illuminating the green narrowband LEDs and blue narrowband LEDs. For example, the wavelength of B2-light is 390nm–445nm, and the wavelength of G2-light is 530nm–550nm. This special light is the illumination light used for NBI (Narrow Band Imaging). However, as special types of light, other wavelengths of light, such as infrared light, are also known, and these can be widely used in this embodiment.
[0066] The parameter value representing the wavelength is, for example, binary information determining whether it is normal light or NBI (Non-Bipolar Illumination). Alternatively, the parameter value can also be information used to determine whether to turn on or off multiple LEDs individually. In this case, the wavelength-related parameter value becomes data with a number of bits corresponding to the number of LEDs. The control unit 332 controls the turning on / off of multiple light sources based on the wavelength-related parameter value. Alternatively, the light source device 350 can also be structured as follows: it includes a white light source and a filter, and switches between normal light and special light based on the insertion / retraction or rotation of the filter. In this case, the control unit 332 controls the filter based on the wavelength parameter value.
[0067] Light intensity refers to the intensity of light emitted by the light source device 350. For example, the endoscope system 300 of this embodiment can also perform known automatic dimming processing. Automatic dimming processing refers to, for example, determining brightness based on an captured image and automatically adjusting the light intensity of the illumination light based on the determination result. In dimming processing, a dimming target value is set as a target value for brightness. The parameter value representing light intensity is, for example, the target light intensity value. By adjusting the target light intensity value, the brightness of the area of interest on the image can be optimized.
[0068] The light intensity ratio represents the ratio of the intensities of multiple lights that can be illuminated by the light source device 350. For example, the light intensity ratio of ordinary light is the ratio of the intensities of R, G, and B light. Parameter values related to the light intensity ratio are, for example, values that are normalized based on the luminous intensities of each of the multiple LEDs. Based on the aforementioned light intensity and light intensity ratio parameters, the control unit 332 determines the luminous intensity of each light source, such as the current value supplied to each light source. Furthermore, the light intensity can also be adjusted via the duty cycle, as described later. In this case, the light intensity ratio is, for example, the duty cycle of each RGB light source. By controlling the light intensity ratio, color deviations caused by individual patient differences can be eliminated, for example.
[0069] Duty cycle refers to the relationship between the on-time and off-time of the light source 352, and in a narrow sense, it represents the ratio of the on-time to the period corresponding to one frame of the image. Parameter values related to duty cycle can be, for example, values representing the aforementioned ratio or values representing the on-time. For example, when the light source 352 emits pulsed light by repeatedly turning on and off, duty cycle control refers to controlling the change in the on-time of each frame. Alternatively, duty cycle control can also involve switching the continuously emitting light source 352 to pulsed light emission. Furthermore, if the on-time is shortened without changing the light intensity per unit time, the total light intensity per frame decreases, thus darkening the image. Therefore, when the duty cycle is reduced, the control unit 332 can also control the total light intensity per frame by increasing the aforementioned dimming target value. By reducing the duty cycle, image jitter can be suppressed. Conversely, by increasing the duty cycle, the total light intensity can be maintained even when suppressing the light intensity per unit time, thus suppressing LED degradation.
[0070] Light distribution refers to the intensity of light corresponding to a direction. For example, if multiple illumination ports are provided at the front end of the insertion part 310b to illuminate different areas respectively, the light distribution changes by adjusting the amount of light irradiated from each illumination port. For example, parameter values related to light distribution refer to information that determines the amount of light at each illumination port. Furthermore, by using optical systems such as lenses and filters, it is also possible to control the amount of light irradiated in each direction, and various modifications can be made by changing the specific structure of the light distribution.
[0071] For example, when the insertion part 310b moves within a tubular subject, the subject located near the center of the image is farther from the front end of the insertion part 310b than a subject located at the periphery of the image. As a result, the image becomes brighter at the periphery than at the center. Conversely, when the front end of the insertion part 310b is approximately directly opposite the wall of the subject, the change in distance from the front end to the subject is not significant between the center and periphery of the image. Therefore, the image becomes brighter at the center than at the periphery. Thus, the direction of brightness and darkness can vary depending on the situation. By controlling the light distribution, the brightness of the desired area of the image can be optimized.
[0072] Frame rate represents the number of images captured per unit of time. Parameters related to frame rate include, for example, the number of frames per second. Increasing the frame rate increases the number of images captured per unit of time, thus suppressing image jitter. Conversely, decreasing the frame rate extends the illumination time of the light source 352, making it easier to obtain bright images.
[0073] A color matrix is, for example, a matrix used to calculate the corrected pixel values based on the original R, G, and B pixel values. The parameter values associated with the color matrix are a set of values representing the elements of that matrix. However, the parameter values are not limited to the matrix itself; other information that can adjust the hue by transforming the RGB pixel values can also be used. By using a color matrix, similar to controlling the light intensity ratio of light source 352, color deviation can be suppressed.
[0074] Structure enhancement processing is, for example, digital enhancement filtering. The parameters associated with structure enhancement refer to information that determines the filtering characteristics of the digital filter, such as the spatial filter size and a set of values representing the filter's various elements. By performing structure enhancement processing, the shape of the region of interest can be clearly defined. However, excessive structure enhancement processing can potentially introduce artifacts.
[0075] Noise reduction processing is, for example, a smoothing filter. The parameter values associated with noise reduction processing refer to information that determines the smoothing filter. This information could be, for example, the value of σ in a Gaussian filter, or a set of values representing the various elements of the filter. Furthermore, the degree of noise reduction can be adjusted by varying the number of times a given filter is applied. In this case, the parameter values associated with noise reduction are numerical values representing the number of times the smoothing filter is applied. By performing noise reduction processing, the noise contained in the image is reduced, thus improving the visual recognizability of the region of interest. However, if noise reduction processing is excessive, it can lead to blunting of edges in the region of interest.
[0076] The parameter values related to AGC represent gain. Increasing the gain can brighten the image, but sometimes it also increases noise. Conversely, while decreasing the gain can suppress noise, sometimes the image brightness is insufficient.
[0077] As described above, the control information includes information on various parameter categories. By adjusting the parameter values of each parameter category, the characteristics of the obtained image change. The control information in this embodiment may include information on all the aforementioned parameter categories, or some may be omitted. Additionally, the control information may also include information on other parameter categories related to the light source, shooting, and image processing.
[0078] The optimal parameter value for improving the estimated probability of the area of interest varies depending on the situation. This situation refers to factors such as the patient, organs, type of area of interest, and the relative position and posture of the tip of the insertion part 310b to the subject. Therefore, determining the optimal control information places a heavy burden on the user. Regarding this, according to the method of this embodiment, as described above, the processing system 100 can automatically determine the control information deemed appropriate.
[0079] 2. Processing flow
[0080] The processing described in this embodiment will be explained in detail. After explaining the learning process for generating the learned model, the derivation process using the learned model will be explained. Here, the derivation process is the region of interest detection process performed in the detection processing unit 334. In addition, the example of using the learned model will also be described below for the control information determination process performed in the control information determination unit 335.
[0081] 2.1 Learning Processing
[0082] Figure 5 This is a structural example of a learning device 400. The learning device 400 includes an acquisition unit 410 and a learning unit 420. The acquisition unit 410 acquires training data for learning. Each piece of training data is data that maps input data to corresponding correct answer labels. The learning unit 420 performs machine learning based on the acquired training data, thereby generating a learned model. Details of the training data and the specific process of the learning process will be described later.
[0083] The learning device 400 can be, for example, an information processing device such as a PC or a server system. Alternatively, the learning device 400 can be implemented through distributed processing using multiple devices. For example, the learning device 400 can be implemented using cloud computing with multiple servers. Furthermore, the learning device 400 can be integrated with the processing system 100, or it can be a separate device.
[0084] An overview of machine learning will be provided. Hereinafter, machine learning using neural networks will be described, but the method of this embodiment is not limited to this. In this embodiment, for example, machine learning using other models such as SVM (support vector machine) can be performed, or machine learning using methods developed from various methods such as neural networks and SVM can be performed.
[0085] Figure 6(A) is a schematic diagram illustrating a neural network. A neural network has an input layer that takes input data, intermediate layers that perform calculations based on the output from the input layer, and an output layer that outputs data based on the output from the intermediate layers. Figure 6 In example (A), a network with two intermediate layers is shown, but the intermediate layers can be one or more layers. Furthermore, the number of nodes in each layer is not limited to a certain value. Figure 6 The example of (A) can be adapted in various ways. Furthermore, considering accuracy, the learning in this embodiment preferably utilizes deep learning with multi-layer neural networks. Here, "multi-layer" is narrowly defined as four layers or more.
[0086] like Figure 6 As shown in (A), the nodes in a given layer are connected to nodes in adjacent layers. Weighting coefficients are assigned to each connection. Each node multiplies the output of the preceding node by the weighting coefficient, and the sum of the multiplication results is obtained. Then, each node adds a bias to the sum, applies an activation function to the addition result, and thus obtains the output of that node. By performing this process sequentially from the input layer to the output layer, the output of the neural network is obtained. Furthermore, various activation functions such as the sigmoid function and the ReLU function are known, and these functions can be widely used in this embodiment.
[0087] Learning in a neural network involves determining appropriate weighting coefficients. These weighting coefficients include biases. Specifically, the learning device 400 inputs the training data into the neural network and calculates the output by performing forward operations using the weighting coefficients at that time. The learning unit 420 of the learning device 400 calculates an error function based on the output and the positive label from the training data. Then, the weighting coefficients are updated to reduce the error function. For example, the backpropagation method, which updates the weighting coefficients from the output layer to the input layer, can be used to update the weighting coefficients during the weighting process.
[0088] In addition, a neural network can be, for example, a CNN (Convolutional Neural Network). Figure 6 (B) is a schematic diagram illustrating a CNN. A CNN consists of convolutional layers and pooling layers that perform convolutional operations. Convolutional layers are layers that perform filtering. Pooling layers are layers that perform pooling operations to reduce the vertical and horizontal dimensions. Figure 6 The example shown in (B) is a network that, after multiple operations based on convolutional and pooling layers, yields its output through operations based on fully connected layers. A fully connected layer is a layer that performs operations by connecting all nodes of the previous layer to nodes of a given layer. Figure 6 (A) corresponds to the operations of each layer described above. Furthermore, although in Figure 6Not illustrated in (B), but also similar when using CNN. Figure 6 Similarly, (A) performs activation function-based computation. Various CNN architectures are known, and these architectures can be widely applied in this embodiment. For example, the CNN of this embodiment can utilize the well-known RPN (Region Proposal Network).
[0089] When using CNNs, the processing steps are also different. Figure 6 The same applies to (A). That is, the learning device 400 inputs the input data from the training data into the CNN and obtains the output by performing filtering or pooling operations using the filtering characteristics at that time. Based on the output and the positive label, an error function is calculated, and the weighting coefficients including the filtering characteristics are updated to reduce the error function. When updating the weighting coefficients of the CNN, for example, the backpropagation method can also be used.
[0090] Figure 7 (A) is an example of training data used to generate the learned model used in the detection processing in the detection processing unit 334. Hereinafter, the learned model used for the detection processing will be denoted as NN1. Figure 7 (B) is a diagram illustrating the input and output of NN1.
[0091] like Figure 7 As shown in (A), the training data includes an input image and annotation data assigned to that input image. The input image is an image captured by the endoscope 310, specifically an image inside a living organism. The annotation data is information identifying regions of interest within the input image, such as information assigned by a user with specialized knowledge, like a physician. For example, as... Figure 7 As shown in (A), the annotation data is information that determines the position and size of a rectangular region containing the region of interest. This rectangular region will be labeled as a detection box. For example, the annotation data is information combining the coordinates of the upper left and lower right endpoints of the detection box. Alternatively, the annotation data can also be information that determines the region of interest in pixels. For example, the annotation data could also be a mask image where pixels within the region of interest are designated as the first pixel value, and pixels outside the region of interest are designated as the second pixel value.
[0092] like Figure 7As shown in (B), the NN1 takes the input image as input, performs forward computation, and outputs detection results and estimated probability information. For example, the NN1 sets a predetermined number of candidate detection boxes on the input image and performs a process to obtain the probability that each candidate is a detection box. In this case, candidate detection boxes with a probability above a predetermined value are output as detection boxes. Furthermore, the probability corresponding to the candidate detection box used as a detection box is the estimated probability. Alternatively, the NN1 calculates the probability that each pixel in the input image is included in a region of interest. In this case, the set of pixels with high probabilities represents the detection result of the region of interest. The estimated probability information is information determined based on the set of probabilities corresponding to pixels classified as regions of interest, such as statistical values representing multiple values of probability. These statistical values can be the average, the median, or other values.
[0093] Furthermore, NN1 can also perform region-of-interest classification. For example, NN1 can accept an input image as input, perform forward computation, and output the location and size of the region of interest, as well as the type of the region of interest, as the detection result. In addition, NN1 outputs estimated probability information representing the likelihood of the detection result. For example, for each candidate detection box, NN1 calculates the probability that the subject contained in the candidate detection box is a polyp of TYPE1, a polyp of TYPE2A, a polyp of TYPE2B, a polyp of TYPE3, or normal mucosa. That is, the detection result of the region of interest in this embodiment can also include the type of the region of interest.
[0094] Figure 8 This is a flowchart illustrating the learning process of NN1. First, in steps S101 and S102, the acquisition unit 410 acquires the input image and the annotation data assigned to that input image. For example, the learning device 400 stores the training data corresponding to the input image and the annotation data in a storage unit (not shown). The processing in steps S101 and S102 is, for example, the process of reading out one of the training data.
[0095] In step S103, the learning unit 420 calculates the error function. Specifically, the learning unit 420 inputs the input image into NN1 and performs a forward operation based on the weighting coefficients at that time. Then, the learning unit 420 calculates the error function based on a comparison between the operation result and the annotation data.
[0096] In step S104, the learning unit 420 performs a process to update the weighting coefficients to reduce the error function. As described above, the process in step S104 can utilize methods such as error backpropagation. The processes in steps S101 to S104 correspond to one learning process based on one training data set.
[0097] In step S105, the learning unit 420 determines whether to end the learning process. For example, the learning unit 420 may end the learning process after performing steps S101 to S104 a predetermined number of times. Alternatively, the learning device 400 may save a portion of the multiple training data as validation data. Validation data is used to confirm the accuracy of the learning results and is not used to update the weighting coefficients. The learning unit 420 may also end the learning process if the positive solution rate of the estimation process using the validation data exceeds a predetermined threshold.
[0098] If the result in step S105 is "No", the process returns to step S101 to continue learning based on the next training data. If the result in step S105 is "Yes", the learning process ends. The learning device 400 sends the information of the generated learned model to the processing system 100. If yes... Figure 3 In the example, the information of the learned model is stored in storage unit 333. Furthermore, various methods such as batch learning and mini-batch learning are known in machine learning, and these methods can be widely applied in this embodiment.
[0099] Figure 9 (A) is an example of training data used to generate the learned model used in the control information determination process in the control information determination unit 335. Hereinafter, the learned model used in the control information determination process will be referred to as NN2. Figure 9 (B) is used to generate Figure 9 The example of training data shown in (A) is a sample of the data. Figure 9 (C) is a diagram illustrating the input and output of NN2.
[0100] like Figure 9 As shown in (A), the training data includes a second input image and control information used when acquiring the second input image. Furthermore, both the input image and the second input image, which are inputs to NN1, are images captured using the mirror section 310; they can be the same image or different images. Additionally, in Figure 9 Example (A) shows an example where the control information includes N parameter categories. The N parameter categories will be denoted as P1 to P2. N P1~P NFor example, these could be any category from wavelength, light intensity ratio, light quantity, duty cycle, light distribution, frame rate, color matrix, structural emphasis, noise reduction, and AGC. Additionally, parameter categories P1 to P... N The parameter values are denoted as p1 to p2. N For example, p1 to p N This includes parameter values related to light source control information, shooting control information, and image processing control information, and a second input image is obtained by performing illumination, shooting, and image processing based on these parameter values. Additionally, as... Figure 9 As shown in (A), the training data contains information for determining the priority parameter categories.
[0101] The priority parameter category indicates the category of parameters that should be changed preferentially to improve the estimation probability. Improving the estimation probability means that the second estimation probability of the second image obtained using the modified control information is higher than the first estimation probability obtained using the given control information for the first image. The priority parameter category is P. k This refers to the determination that P has been changed. k Compared to parameter values of other parameter categories, P was changed. k Given the parameter values, the probability of the estimated probability is high. k is an integer between 1 and N.
[0102] For example, during the training data acquisition phase, images are acquired by continuously capturing images of a given subject containing the region of interest while changing the control information. At this time, the parameter values are changed category by category. For example, if the control information shown in C1 is set to the initial value, in C2, the parameter value of only P1 relative to the initial value is changed from p1 to p1' to acquire an image. This process is repeated thereafter. Furthermore, the estimated probabilities are calculated by inputting the images acquired at each time point into the aforementioned NN1. Here, an example is shown where the estimated probabilities are represented as 0 to 100%.
[0103] exist Figure 9 In the example shown in (B), the estimated probability is as low as 40% in the initial control information. In contrast, the characteristics of the image obtained by changing the control information change, thus changing the estimated probability of the output when that image is input into NN1. For example, as... Figure 9 As shown in (B)C3, after changing the parameter category P k When the parameter value is , the estimated probability becomes maximum. In this case, the parameter class considered to be dominant with respect to the estimated probability is P. k By prioritizing changes to P k This can improve the estimated probability. That is, the acquisition unit 410 uses the image of C1 as the second input image, and sets the parameter values of C1, i.e., p1 to p... N As control information, Pk One training data point is thus obtained by using the priority parameter category. However, the data used to obtain the training data is not limited to... Figure 9 (B) can perform various transformations.
[0104] like Figure 9 As shown in (C), NN2 accepts a second input image and control information as input, performs forward computation, and outputs the probability of recommending a change in that parameter category for each parameter category. For example, the output layer of NN2 contains N nodes and outputs N output data.
[0105] Figure 10 This is a flowchart illustrating the learning process of NN2. First, in steps S201 and S202, the acquisition unit 410 acquires and Figure 9 The image group corresponding to (B) and the control information group, which is a set of control information when acquiring each image. In step S203, the acquisition unit 410 obtains the estimated probability group by inputting each image into the generated NN1. Through the processing of steps S201 to S203, the acquisition unit obtains the estimated probability group. Figure 9 The data shown in (B) is as follows.
[0106] Next, in step S204, the acquisition unit 410 performs based on Figure 9 The data shown in (B) is used to process the training data. Specifically, the acquisition unit 410 acquires and accumulates multiple training data sets that correspond to the second input image, control information, and priority parameter categories.
[0107] In step S205, the acquisition unit 410 reads out one piece of training data. In step S206, the learning unit 420 calculates the error function. Specifically, the learning unit 420 inputs the second input image and control information from the training data into NN2, and performs forward calculations based on the weighting coefficients at this time. The calculation result is as follows: Figure 9 As shown in (C), N data points represent the probability of recommending each parameter category as a preferred parameter category. In the case that the output layer of NN2 is a known softmax layer, the N data points are probability data that sum to 1. The preferred parameter categories included in the training data are P. k In this case, the correct label is recommended as P1~P k-1 The probability of change is 0, P is recommended. k The probability of change is 1, so P is recommended. k+1 ~P N The probability of the data is 0. The learning unit 420 calculates the error function by comparing the N probability data obtained through positive calculation with the N probability data that serve as the positive solution label.
[0108] In step S207, the learning unit 420 performs a process of updating the weighting coefficients to reduce the error function. The process in step S207, as described above, can utilize methods such as backpropagation. In step S208, the learning unit 420 determines whether to end the learning process. The termination condition for the learning process can be based on the number of times the weighting coefficients have been updated, as described above, or it can be based on the positive resolution rate of the estimation process using validation data. If the result in step S208 is "No," the process returns to step S205 and continues with the learning process based on the next training data. If the result in step S208 is "Yes," the learning process ends. The learning device 400 sends the information of the generated learned model to the processing system 100.
[0109] Furthermore, the processing performed by the learning device 400 in this embodiment can also be implemented as a learning method. The learning method of this embodiment acquires an image captured by an endoscope as an input image. When using information used in controlling the endoscope as control information, it acquires control information acquired at the time the input image was acquired, acquires control information for improving the estimated probability information representing the probability of a region of interest detected from the input image, and performs machine learning on the relationship between the input image, the control information acquired at the time the input image was acquired, and the control information for improving the estimated probability information, thereby generating a learned model. Furthermore, the control information for improving the estimated probability information can also be a priority parameter category as described above. Alternatively, the control information for improving the estimated probability information can also be, as a variation described later, a group of priority parameter categories and specific parameter values. Alternatively, the control information for improving the estimated probability information can also be, as a further variation described later, a set of multiple parameter categories and parameter values for each parameter category.
[0110] 2.2 Derivation Processing
[0111] Figure 11 This is a flowchart illustrating the processing of the processing system 100 in this embodiment. First, in step S301, the acquisition unit 110 acquires an image of the detected object. For example, the acquisition unit 110 may also be used as follows: Figure 12 The target image is acquired at a ratio of 2 frames per second, as described later. Furthermore, the acquisition unit 110 can obtain control information used in acquiring the target image from the control unit 130.
[0112] In step S302, the processing unit 120 (detection processing unit 334) calculates the detection result of the region of interest and an estimated probability representing its likelihood by inputting a detection object image to the NN1. In step S303, the processing unit 120 determines whether the estimated probability is above a given threshold. If the result is "yes" in step S303, the detection result of the region of interest is sufficiently reliable. Therefore, in step S304, the processing unit 120 outputs the detection result of the region of interest. For example, the detection result of the region of interest from the detection processing unit 334 is sent to the post-processing unit 336, and displayed on the display unit 340 after post-processing by the post-processing unit 336. Furthermore, if using... Figure 13 (A) Figure 13 As described later in (B), the processing unit 120 outputs the estimated probability along with the detection results of the region of interest.
[0113] If "No" is selected in step S303, the reliability of the detection results for the region of interest is low, and even if displayed directly, it may be useless for the user's diagnosis. On the other hand, if the display itself is omitted, the user cannot be notified of the possibility of the existence of a region of interest. Therefore, the control information is changed in this embodiment.
[0114] If the answer in step S303 is "No", then in step S305, the processing unit 120 (control information determination unit 335) performs processing for updating the control information. The processing unit 120 determines the priority parameter category by inputting the detected object image and control information obtained in step S301 into NN2.
[0115] In step S306, the control unit 130 (control unit 332) controls the parameter values of the priority parameter categories to change. The acquisition unit 110 acquires the detected object image based on the changed control information. Furthermore, NN2 here determines the priority parameter category, but not the specific parameter value. Therefore, the control unit 332 controls the parameter values of the priority parameter categories to change sequentially. The acquisition unit 410 acquires a group of detected object images based on the parameter value group.
[0116] In step S307, the processing unit 120 inputs each detection object image contained in the detection object image group into NN1 to calculate the detection result of the region of interest and the estimated probability representing its probability. The detection processing unit 334 extracts the parameter value and the detection object image with the highest estimated probability from the parameter value group and the detection object image group.
[0117] In step S308, the detection processing unit 334 determines whether the estimated probability of the region of interest in the extracted detection object image is above a given threshold. If the probability is "yes" in step S308, the detection result of the region of interest is sufficiently reliable. Therefore, in step S309, the detection processing unit 334 outputs the detection result of the region of interest.
[0118] If the answer in step S308 is "No", then in step S310, the control information determination unit 335 performs processing for updating the control information. For example, the control information determination unit 335 determines the priority parameter category by inputting the detection object image extracted in step S307 and the control information used in acquiring the detection object image into NN2.
[0119] The priority parameter category determined in step S310 can be the same as the priority parameter category P determined in step S305. i Different parameter categories P j (j is an integer satisfying j≠i). Performing step S310 is equivalent to adjusting the parameter category P. i The estimated probability is not sufficiently improved. Therefore, by comparing it with P... i Different parameter categories can be used as the next priority parameter category, which can effectively improve the estimated probability. In step S311, based on the control of changing the parameter values of the priority parameter categories in sequence, the acquisition unit 410 acquires a group of detected object images based on the parameter value group.
[0120] Furthermore, in step S310, the processing using NN2 can be omitted, and the processing result in step S305 can be used instead. For example, the control information determination unit 335 can also perform the processing of selecting the parameter category determined in step S305 as the preferred parameter category.
[0121] exist Figure 11 In this embodiment, a loop process is described to clearly indicate changes in the object image to be detected, the control information when acquiring the object image, and the priority parameter category. However, as shown in steps S301 to S305, the processing in this embodiment repeatedly performs a loop including acquiring the object image and control information, detection processing, threshold determination of the estimated probability, and determination of the control information, until the condition that the estimated probability is above the threshold is met. Therefore, the same loop process continues after step S311. Furthermore, if the estimated probability is not sufficiently improved even if the loop is continued, the processing unit 120 terminates the processing, for example, by changing all parameter categories or by repeating the loop a predetermined number of times.
[0122] Additionally, the above uses Figure 11The processing described focuses on processing a single region of interest. After detecting a region of interest, it cannot be said that continuing to excessively process that region of interest is preferable. This is because, assuming that depending on the operation of the insertion unit 310b, the field of view of the camera unit changes, the region of interest shifts from the detected object image, or a new region of interest is detected within the detected object image. As described above, by limiting the maximum number of executions of the loop processing, it is possible to smoothly switch from processing a given region of interest to processing different regions of interest. Especially... Figure 11 The following Figure 15 In addition, considering the small number of parameter categories to be changed and the time required for the estimated probability to improve, it is important to suppress the number of loops.
[0123] As explained above, the processing unit 120 of the processing system 100 operates according to the learned model, thereby detecting regions of interest contained in the object image and calculating estimated probability information related to the detected regions of interest. The learned model here corresponds to NN1. Furthermore, the processing unit 120 can also determine control information for improving the estimated probability by operating according to the learned model.
[0124] The operations performed in the processing unit 120 of the learned model, that is, the operations for outputting output data based on input data, can be performed either by software or by hardware. In other words, in Figure 6 The product sum operation performed in each node of (A), the filtering process performed in the convolutional layers of the CNN, etc., can also be performed in software. Alternatively, the above operations can also be performed by circuit devices such as FPGAs. In addition, the above operations can also be performed by a combination of software and hardware. In this way, the operation of the processing unit 120 according to the instructions from the learned model can be implemented in various ways. For example, the learned model includes a derivation algorithm and weighting coefficients used in the derivation algorithm. The derivation algorithm is an algorithm that performs filtering operations, etc., based on the input data. In this case, the derivation algorithm and the weighting coefficients can be stored in a storage unit, and the processing unit 120 performs derivation processing in software by reading the derivation algorithm and the weighting coefficients. The storage unit is, for example, the storage unit 333 of the processing device 330, but other storage units can also be used. Alternatively, the derivation algorithm can be implemented by an FPGA, etc., and the storage unit stores the weighting coefficients. Alternatively, the derivation algorithm including weighting coefficients can be implemented by an FPGA, etc. In this case, the storage unit storing the information of the learned model is, for example, the built-in memory of the FPGA.
[0125] Furthermore, the processing unit 120 calculates first estimated probability information based on the first detected object image obtained using control with the first control information and the learned model. If the processing unit 120 determines that the probability represented by the first estimated probability information is lower than a given threshold, it performs processing to determine second control information as control information for improving the estimated probability information. This processing is, for example, similar to... Figure 11 The corresponding processing steps S302, S303, and S305 are as follows.
[0126] In this way, even when the estimated probability is insufficient, i.e. when the reliability of the detected region of interest is low, we can try to obtain more reliable information.
[0127] In existing methods, if information with a low estimated probability is provided, the user may make a misdiagnosis. On the other hand, setting a high estimated probability threshold for whether to provide information to the user can suppress misdiagnosis, but it reduces the amount of auxiliary information provided and increases the chance of overlooking lesions. The method according to this embodiment solves this trade-off problem. As a result, it provides information that enables users to make high-precision diagnoses and suppresses overlooked lesions.
[0128] Furthermore, the processing unit 120 can also calculate second estimated probability information based on the second detected object image obtained according to the control using the second control information and the learned model. If the processing unit 120 determines that the probability represented by the second estimated probability information is lower than a given threshold, it performs processing to determine third control information as control information for improving the estimated probability information. This processing is, for example, similar to... Figure 11 The corresponding processing steps are S307, S308, and S310.
[0129] In this way, even when the estimated probability is insufficient after changing the control information, the control information can be further modified, thus improving the probability of obtaining highly reliable information. Specifically, it is possible to broadly search for appropriate control information that makes the estimated probability exceed a given threshold.
[0130] In addition, such as Figure 4As shown, the control information includes at least one of the following: light source control information for controlling the light source 352 that illuminates the subject; shooting control information for controlling the shooting conditions for capturing an image of the target; and image processing control information for controlling the image processing of the captured image signal. As described above, an image of the target is acquired by performing the steps of emitting light from the light source 352, receiving light from the subject by the imaging element 312, and processing the image signal as a result of the light reception. By performing control related to any of these steps, the characteristics of the acquired image of the target can be changed. As a result, the estimated probability can be adjusted. However, as described above, the control unit 130 does not need to perform all control of the light source, shooting, and image processing to improve the estimated probability; any one or two of these steps can be omitted.
[0131] Alternatively, the control information can also represent at least one of the following: hue, brightness, and position within the image, all related to the region of interest. In other words, the light source 352, etc., can be controlled to make at least one of the following—hue, brightness, and position—of the region of interest in the detection object image approach a desired value. For example, by making the hue associated with the region of interest similar to the hue of the input image included in the training data, or, more specifically, the hue of the portion corresponding to the region of interest in the input image, the estimation probability can be improved. Or, by controlling the light source 352 to make the brightness associated with the region of interest in the detection object image the optimal brightness for that region of interest, the estimation probability can be improved.
[0132] Furthermore, the control information may also include first to Nth (N is an integer greater than or equal to 2) parameter categories. The processing unit 120 can also determine the second control information by changing the i-th (i is an integer less than or equal to 1 ≤ i ≤ N) parameter category among the first to Nth parameter categories included in the first control information. In this way, parameter categories that contribute significantly to the estimated probability can be changed, thus improving the estimated probability through efficient control. As will be described later, control is easier compared to simultaneously changing the parameter values of multiple parameter categories.
[0133] Furthermore, the processing unit 120 calculates second estimated probability information based on the second captured image obtained using control with the second control information and the learned model. If it is determined that the probability represented by the second estimated probability information is lower than a given threshold, the processing unit 120 can determine third control information by changing the j-th parameter category (j is an integer satisfying 1≤j≤N, j≠i) among the first to N-th parameter categories included in the second control information.
[0134] In this way, even if changing the parameter value of a given parameter category does not sufficiently improve the estimated probability, it is possible to try changing the parameter value of different parameter categories. Therefore, the probability of improving the estimated probability can be increased.
[0135] Furthermore, the processing unit 120 can also perform processing to determine second control information, which is used to improve the estimated probability information, based on the first control information and the first image of the detected object. That is, when determining the control information, not only the image of the detected object can be used, but also the control information used when acquiring the image of the detected object can be used. For example, such as... Figure 9 As shown in (C), NN2, as a learned model used to determine control information, accepts a set of images and control information as input.
[0136] In this way, processing can be performed not only on the object image but also on the image itself, taking into account the conditions under which it was acquired. Therefore, processing accuracy can be improved compared to using only the object image as input. However, the first control information is not necessarily required in determining the second control information. For example, it can also be derived from... Figure 9 Control information is omitted from the training data shown in (A). In this case, fewer inputs are needed in the processing to determine the control information, thus reducing the processing load. Furthermore, the reduced amount of training data also reduces the learning processing load.
[0137] Additionally, the processing unit 120 can also perform processing to determine control information for improving the estimated probability information based on the second learned model and the detected object image. The second learned model is a model obtained by machine learning the relationship between the second input image and the control information for improving the estimated probability information. The second learned model corresponds to NN2 mentioned above.
[0138] In this way, machine learning can be utilized not only in the detection process but also in the determination process of control information. Therefore, the accuracy of control information determination is improved, and the estimated probability can be rapidly increased above a threshold. However, as described later, variations can be implemented for the determination process of control information without using machine learning.
[0139] Furthermore, the processing performed by the processing system 100 in this embodiment can also be implemented as an image processing method. The image processing method of this embodiment acquires a detection object image captured by an endoscope device, detects a region of interest contained in the detection object image based on a learned model obtained through machine learning for calculating estimated probability information representing the probability of a region of interest within the input image, calculates estimated probability information related to the detected region of interest, and, when using information used in the control of the endoscope device as control information, determines control information for improving the estimated probability information related to the region of interest within the detection object image based on the detection object image.
[0140] 2.3 Background processing and display processing
[0141] Figure 12 This is a diagram illustrating the relationship between the captured frame and the images acquired in each frame. For example, the acquisition unit 110 of the processing system 100 acquires an image of a detected object that serves as a region of interest and an object for calculating the estimated probability information in a given frame, and acquires a display image used in the display unit 340 in a frame different from the given frame.
[0142] That is, in the method of this embodiment, the detection object image, which is the object of the detection processing and control information determination processing of the processing system 100, and the display image, which is the object of the prompt to the user, can be separated. In this way, the control information used in acquiring the detection object image and the display control information used in acquiring the display image can be managed separately.
[0143] As described above, the characteristics of the acquired image change when the control information is altered. When the control information for displaying the image changes frequently or abruptly, the changes in the image become more significant, potentially hindering the user's observation and diagnosis. In this regard, by distinguishing between the image of the object being detected and the displayed image, abrupt changes in the control information used for displaying the image can be suppressed. That is, the changes in the characteristics of the displayed image can be limited to a relatively gradual change, even if altered manually by the user or automatically, thus suppressing situations that hinder the user's diagnosis.
[0144] Furthermore, since the image of the object being detected does not need to be displayed, the control information can be changed drastically. Therefore, control can be performed at high speed, improving the estimation probability. For example, when adjusting the light intensity of the LED 352 as a light source using automatic dimming control, the light intensity changes in conventional dimming are limited to gradual changes compared to the light intensity changes achievable by the characteristics of the LED. This is because drastic changes in brightness can hinder the user's diagnosis. In this respect, even drastic changes in the light intensity of the image of the object being detected are unlikely to cause problems, thus fully utilizing the light intensity variation capabilities of the LED.
[0145] Additionally, obtaining part 110 can also be done as follows: Figure 12 The detection object image and the display image are acquired alternately as shown. This reduces the difference between the acquisition time of the detection object image and the acquisition time of the display image.
[0146] The processing unit 120 of the processing system 100 can also perform processing to display the detection results of the region of interest detected based on the object image and the estimated probability information calculated based on the object image when the estimated probability information calculated based on the object image is above a given threshold.
[0147] Figure 13 (A) Figure 13 (B) is an example of a display image displayed on display unit 340. Figure 13 In (A), A1 represents the displayed image of the display unit 340. Furthermore, A2 represents the region of interest captured in the displayed image. A3 represents the detection bounding box detected in the object image, and A4 represents the estimated probability. Figure 13 As shown in (A), by overlaying the detection results and estimated probability information of the region of interest onto the display image, information related to the region of interest in the display image can be presented to the user in an easily understandable way.
[0148] The image of the object being detected here is, for example, in Figure 12 The image is acquired in frame F3, and the displayed image is acquired in frame F4. This reduces the difference between the capture time of the object image and the capture time of the displayed image. Since the difference between the state of the region of interest in the object image and the state of the region of interest in the displayed image is reduced, it is easier to correlate the detection results based on the object image with the displayed image. For example, since the difference between the position and size of the region of interest in the object image and the position and size of the region of interest in the displayed image is sufficiently small, the probability that the region of interest in the displayed image is included in the detection frame can be improved when detection frames are overlapped. Furthermore, the displayed image and the object image are not limited to images acquired in consecutive frames. For example, depending on the time required for detection processing, the detection results using the object image of frame F3 can be overlaid on display images acquired at times after F4, such as frames F6 and F8 (not shown).
[0149] In addition, the processing unit 120 may also perform processing to associate and display the region in the detected object image that contains at least the region of interest with the display image if the estimated probability information calculated based on the learned model NN1 and the detected object image is above a given threshold.
[0150] For example, such as Figure 13 As shown in (B), the processing unit 120 can also perform the process of displaying a portion of the detected object image on the display unit 340. Figure 13 B1 in (B) represents the display area of display unit 340. B2 to B5 are respectively related to... Figure 13 Similarly, A1 to A4 of (A) represent the displayed image, the region of interest in the displayed image, the detection box, and the estimated probability. As shown in B6, the processing unit 120 can also perform processing to display a portion of the detected object image in a region of the display area that is different from the region where the displayed image is displayed.
[0151] Since the image of the detection target is an image in a state where the estimated probability of the region of interest is high, it is considered an image in which it is easy for the user to recognize the region of interest. Therefore, by displaying at least the part of the image of the detection target that includes the region of interest, it is possible to assist in determining whether the region represented by the detection frame is truly the region of interest. In addition, Figure 13 In (B) of shows an example of displaying a part of the image of the detection target, but it is also possible to display the entire image of the detection target.
[0152] 3. Variations
[0153] Hereinafter, several variations will be described.
[0154] 3.1 Structure of the learned model
[0155] As described above, an example has been described in which the learned model NN1 for performing the detection process and the learned model NN2 for performing the determination process of the control information are different models. However, NN1 and NN2 can also be implemented by a single learned model NN3.
[0156] For example, NN3 is a network that accepts the image of the detection target and the control information as inputs and outputs the detection result of the region of interest, the estimated probability information, and the priority parameter category. Various structures can be considered for the specific structure of NN3. For example, NN3 can also be a model including a feature quantity extraction layer commonly used in the detection process and the determination process of the control information, a detection layer for performing the detection process, and a control information characteristic layer for performing the determination process of the control information. The feature quantity extraction layer is a layer that accepts the image of the detection target and the control information as inputs and outputs the feature quantity. The detection layer is a layer that accepts the feature quantity from the feature quantity extraction layer as an input and outputs the detection result of the region of interest and the estimated probability information. The control information determination layer is a layer that accepts the feature quantity from the feature quantity extraction layer as an input and outputs the priority parameter category. For example, based on Figure 7 the training data shown in (A) of, the weighting coefficients included in the feature quantity extraction layer and the detection layer are learned. In addition, based on Figure 9 the training data shown in (A) of, the weighting coefficients included in the feature quantity extraction layer and the control information determination layer are learned. Depending on the form of the training data, the detection layer and the control information determination layer can also be learning targets at the same time.
[0157] 3.2 Variations related to the process of determining the control information
[0158] <Variation 1 of NN2>
[0159] Figure 14 (A) of is another example of the training data for generating the learned model used in the determination process of the control information. Figure 14 (B) of is for generating Figure 14 The example of training data shown in (A) is a sample of the data. Figure 14 (C) is a diagram illustrating the input and output of NN2.
[0160] like Figure 14 As shown in (A), the training data includes a second input image, control information used when acquiring the second input image, a priority parameter category, and a recommended value for that priority parameter category. The control information and... Figure 9 The example (A) is the same, concerning parameter categories P1 to P2. N The parameter values p1~p N The priority parameter category is also related to... Figure 9 Similarly, (A) is information that determines any parameter category, i.e., P. k The recommended value refers to the category P as a parameter. k The recommended value for the parameter is p. k '.
[0161] For example, in the training data acquisition phase, images are acquired by continuously capturing images of a given subject containing the region of interest while changing the control information. At this time, data related to various parameter values is acquired for a given parameter category. For simplicity, the following explanation will focus on parameter values p1 to p... N This indicates that there are M candidate values for each parameter. For example, parameter value p1 can be selected from p... 11 ~p 1M Choose from M candidate values. For parameter values p2 to p... N The same applies. However, the number of candidate parameter values can vary depending on the parameter category.
[0162] For example, suppose the initial value of the parameter is p as shown in D1. 11 p 21 , ···, p N1 For the data within the range shown in D2, only the parameter values of parameter category P1 are changed sequentially to p. 12 ~p 1M Parameter categories P2~P N The parameter values are fixed. Similarly, within the range shown in D3, only the parameter values for parameter category P2 change sequentially to p. 22 ~p 2M The parameter values for other parameter categories are fixed. This will continue thereafter. Figure 14 In the example shown in (B), starting from the initial state shown in D1, the parameter values for the N parameter categories change to M-1 different types. Therefore, based on N×(M-1) different control information, N×(M-1) images are obtained. The acquisition unit 410 calculates the estimated probability by inputting these N×(M-1) images into NN1 respectively.
[0163] For example, suppose that in classifying parameter P k parameter value p k Change to p k At time 'p', the estimated probability becomes the maximum. k 'for p k2 ~p kM Any value in it. In this case, the acquisition unit 410 takes the image shown in D1 as the second input image, and takes the parameter value shown in D1, i.e., p 11 ~p N1 As control information, parameter category P k As a priority parameter category, p k As a recommended value, one training data point is thus obtained. However, the data used to obtain the training data is not limited to... Figure 14 (B) can implement various transformations. For example, it can perform transformations that omit a portion to take into account the processing load, without needing to obtain all of the N×(M-1) additional data.
[0164] like Figure 14 As shown in (C), NN2 accepts a second input image and control information as input, performs positive operations, and outputs priority parameter categories and recommended values.
[0165] For example, the output layer of an NN2 layer can also contain N×M nodes, outputting N×M output data. Since N×M output nodes contain output data, P1 should be set to p. 11 The probability data should be set as p1. 12 The probability data, ..., should set P1 as p 1M The probability data has M nodes. Similarly, the output node contains the output, and P2 should be set to p. 21 The probability data should be set as p2. 22 The probability data, ..., should set P2 as p 2M The NN2 in this variant outputs M nodes representing the probability data. Similarly, for all combinations of N priority parameter categories and M parameter values, it outputs the probability data that should be recommended. If the node with the largest probability data can be determined, then the priority parameter category and the recommended value can be determined. However, the specific structure of NN2 is not limited to this; other structures capable of determining the priority parameter category and the recommended value can also be used.
[0166] The learning process of NN2 and Figure 10The same. However, in step S206, the learning unit 420 inputs the second input image and the control information in the training data into NN2, performs a forward operation according to the weighting coefficient at this time, and thereby obtains N×M probability data. The learning unit 420 obtains an error function based on the comparison process of the N×M probability data with the priority parameter category and the recommended value in the training data. For example, when the priority parameter category included in the training data is P1 and the recommended value is p 11 in the case of, the correct label is the information that the probability data with P1 set to p 11 is 1 and all other probability data are 0.
[0167] Figure 15 is a flowchart illustrating the processing of the processing system 100 in the present embodiment. Steps S401 to S404 are the same as Figure 11 steps S301 to S304 of. That is, the processing unit 120 obtains an estimated probability by inputting the detection target image into NN1, and performs an output process if the estimated probability is above the threshold, otherwise performs a process of changing the control information.
[0168] When the result in step S403 is "No", in step S405, the processing unit 120 (control information determination unit 335) performs a process for updating the control information. The processing unit 120 determines the priority parameter by inputting the detection target image and the control information obtained in step S401 into NN2. As used Figure 14 in (C) of, in this modification example, not only the priority parameter category can be determined, but also the recommended parameter value can be determined.
[0169] In step S406, the control unit 130 (control unit 332) performs control to change the parameter value of the priority parameter category to the recommended value. The acquisition unit 110 acquires the detection target image based on the determined control information. Different from the example shown in Figure 11 , since the recommended value is determined, there is no need to perform control to sequentially change the parameter values of the priority parameter category. As a result, compared with the example in Figure 11 , the estimated probability can be improved in a short time. Steps S407 to S411 are repetitions of the same process, so the description is omitted.
[0170] <Variant Example 2 of NN2>
[0171] Figure 16 of (A) is another example of the training data for generating the learned model used in the control information determination process. Figure 16 of (B) is a diagram illustrating the input and output of NN2.
[0172] As Figure 16As shown in (A), the training data includes a second input image, control information used when acquiring the second input image, and recommended control information. The control information and... Figure 9 The example (A) is the same, concerning parameter categories P1 to P2. N The parameter values p1~p N Recommended control information refers to information for parameter categories P1 to P2. N Recommended parameter values p1' to p1' for improving the estimated probability. N '.
[0173] For example, in the training data acquisition phase, images are obtained by continuously capturing a given subject containing the region of interest while changing the control information. At this point, for each of the N parameter categories, the parameter value can have M-1 possible changes, thus obtaining (M-1) combinations of these values. N Then, the acquisition unit 410 acquires one training data point by using the control information that maximizes the estimated probability as the recommended control information. However, depending on the values of N and M, a large number of images need to be taken to acquire the training data. Therefore, the collection target can also be limited to (M-1). N A portion of the data. Furthermore, regarding the process of obtaining... Figure 16 The data collected from the training data shown in (A) enables the implementation of various variations.
[0174] like Figure 16 As shown in (B), NN2 accepts a second input image and control information as input, performs forward computation, and outputs recommended values for the parameter values of each parameter category. For example, the output layer of NN2 contains N nodes and outputs N output data.
[0175] The learning process of NN2 and Figure 10 The process is the same. However, in step S206, the learning unit 420 inputs the second input image and control information from the training data into NN2, performs a forward calculation based on the weighting coefficients at this time, and thereby obtains N recommended values. The learning unit 420 calculates the error function based on the comparison between these N recommended values and the N parameter values contained in the recommended control information of the training data.
[0176] Figure 17 This is a flowchart illustrating the processing of the processing system 100 in this embodiment. Steps S501 to S504 are... Figure 11 Steps S301 to S304 are the same. That is, the processing unit 120 calculates the estimated probability by inputting the image of the detected object into NN1. If the estimated probability is above a threshold, display processing is performed; otherwise, the control information is changed.
[0177] If the answer in step S503 is "No", then in step S505, the processing unit 120 (control information determination unit 335) performs processing for updating the control information. The processing unit 120 determines the recommended control information by inputting the detected object image and control information obtained in step S501 into NN2. As described above, the control information here is a set of recommended parameter values in each parameter category.
[0178] In step S506, the control unit 130 (control unit 332) performs control to change the parameter values of multiple parameter categories to recommended values. The acquisition unit 110 acquires an image of the detected object based on the determined control information. Figure 11 , Figure 15 The examples shown are different, allowing for the simultaneous modification of multiple parameter categories. Additionally, as... Figure 17 As shown, in this variant, the processing can also be terminated without performing output processing (steps S504, S509), and the same loop processing can continue.
[0179] If used Figure 16 , Figure 17 As described above, the control information may include first to Nth (N is an integer of 2 or more) parameter categories, and the processing unit 120 determines the second control information by changing two or more parameter categories from the first to Nth parameter categories included in the first control information. By centrally changing the parameter values of multiple parameter categories in this way, [the system]... Figure 11 , Figure 15 Compared to previous examples, this method can improve the estimated probability in a short period of time.
[0180]
[0181] Furthermore, the above examples illustrate the application of machine learning in the processing of control information, but machine learning is not mandatory. For example, processing unit 120 may also determine a priority order for multiple parameter categories included in the control information and modify the control information according to that priority order. For example, in Figure 11 In the flowchart shown, step S305 is omitted, and the parameter value of the parameter category with the highest priority is changed.
[0182] Alternatively, the processing system 100 may store a database that maps the current image and control information to priority parameter categories such as those that increase the estimated probability. Then, the processing unit 120 performs a process to determine the similarity between the image of the target object and the control information obtained when acquiring that image, and the images and control information stored in the database. The processing unit 120 then performs a process to determine the control information by changing the priority parameter category that maps to the data with the highest similarity.
[0183] 3.3 Operation Information
[0184] Furthermore, the acquisition unit 110 of the processing system 100 can also acquire operation information based on user operations on the endoscope device. The processing unit 120 determines control information to improve the estimated probability information based on the detected object image and the operation information. In other words, operation information can also be added as input to the process of determining the control information.
[0185] For example, when a user performs a user operation such as pressing the zoom button or bringing the front end of the insertion unit 310b close to the subject, it is assumed that the user desires a more detailed examination of the subject. For instance, it is assumed that the user desires not only auxiliary information indicating the presence or absence of polyps, but also auxiliary information to assist in the classification and identification of polyps. In this case, the processing unit 120, for example, changes the illumination light to NBI. This improves the estimation probability of the detection result of the region of interest, and more specifically, improves the estimation accuracy of the detection result, including the classification result of the region of interest. That is, when a specified user operation is performed, the processing unit 120 determines control information based on that user operation, enabling control that reflects the user's intention. Furthermore, when performing user operations such as zooming or approaching, it is assumed that the camera unit is directly facing the subject. Therefore, it is assumed that controlling the light distribution as control information can also improve the estimation probability. Moreover, various variations can be implemented for the specific parameter categories and parameter values of the control information determined based on the operation information.
[0186] For example, the storage unit of the processing system 100 may store the learned model NN2_1 when a user operation is performed and the learned model NN2_2 when no user operation is performed as learned models for determining control information. The processing unit 120 switches the learned model used in determining control information based on the user operation. For example, when the processing unit 120 detects that a specified user operation has been performed, such as pressing a zoom button or a magnification button, it determines that a learned model has been performed. Alternatively, the processing unit 120 determines that a user operation has been performed to bring the insertion unit 310b closer to the subject based on the amount of illumination and the brightness of the image. For example, when the amount of illumination is small but the image is bright, it can be determined that the tip of the insertion unit 310b is close to the subject.
[0187] For example, NN2_1 is learned as a parameter value related to the wavelength of the light source control information, making it easy to select parameter values for NBI selection. This facilitates detailed observation of the subject using NBI during user operations. For instance, as the detection result of the region of interest, a classification result according to the NBI classification criteria is output. Examples of NBI classification criteria include VS classification for gastric lesions, and JNET, NICE, and EC classifications for large intestine lesions.
[0188] On the other hand, NN2_2 can be learned, for example, as a parameter value related to the wavelength of the light source control information, and it is easy to select parameter values for selecting ordinary light. Alternatively, NN2_2 can also be learned as a model that increases the probability of determining the control information for emitting amber and violet light from the light source device 350. Amber light has a peak wavelength in the band 586nm to 615nm, and violet light has a peak wavelength in the band 400nm to 440nm. These lights are, for example, narrow-band lights with a half-width of tens of nm. Violet light is suitable for obtaining the characteristics of superficial blood vessels or glandular structures of mucosa. Amber light is suitable for obtaining the characteristics of deep blood vessels or redness, inflammation, etc. of mucosa. That is, by irradiating amber and violet light, lesions that can be detected based on the characteristics of superficial blood vessels or glandular structures of mucosa, or lesions that can be detected based on the characteristics of deep blood vessels or redness, inflammation, etc. of mucosa, can be detected as regions of interest. Without user intervention, the ease of using violet and amber light improves the estimated probability of a wide range of lesions, including cancer and inflammatory diseases.
[0189] Furthermore, the features that differentiate the control information determined based on operational information are not limited to switching the learned model used in the deterministic process. For example, NN2 is a model used both when a user operation is performed and when no user operation is performed, and operational information can also be used as input to NN2.
[0190] 3.4 Switching between learned models used for detection and processing
[0191] Furthermore, the processing unit 120 can also perform a first processing and a second processing. The first processing is the detection of regions of interest contained in the detection object image based on a first learned model and the detection object image. The second processing is the detection of regions of interest contained in the detection object image based on a second learned model and the detection object image. In other words, multiple learned models NN1 can be set for the detection processing. Hereinafter, the first learned model will be referred to as NN1_1, and the second learned model will be referred to as NN1_2.
[0192] Then, the processing unit 120 performs a process of selecting any one of the multiple learned models, including the first learned model NN1_1 and the second learned model NN1_2. If the first learned model NN1_1 is selected, the processing unit 120 performs the following processing: based on the first learned model NN1_1 and the detected object image, it performs a first process and calculates estimated probability information, and determines control information for improving the calculated estimated probability information. If the second learned model NN1_2 is selected, the processing unit 120 performs the following processing: based on the second learned model NN1_2 and the detected object image, it performs a second process and calculates estimated probability information, and determines control information for improving the calculated estimated probability information.
[0193] In this way, the learned model used for detection processing can be switched according to the situation. Furthermore, the process of determining control information that improves the estimated probability information output as the first learned model and the process of determining control information that improves the estimated probability information output as the second learned model can be the same process or different processes. For example, as described above, multiple learned models for determining control information can also be used. Several specific examples of switching learned models are described below.
[0194] For example, the processing unit 120 can also operate in any of a plurality of judgment modes, including: an existence judgment mode, which determines whether a region of interest is contained in the detection object image based on a first learned model NN1_1 and the detection object image; and a qualitative judgment mode, which determines the state of the region of interest contained in the detection object image based on a second learned model NN1_2 and the detection object image. This allows for processing that prioritizes either the determination of the presence or absence of a region of interest or the determination of its state. For example, the processing unit 120 can switch the learned model used in the detection processing depending on whether the goal is to detect lesions without omission or to correctly classify the stages of the detected lesions.
[0195] Specifically, the processing unit 120 can also determine whether to switch to the qualitative determination mode based on the detection results of the region of interest in the existence determination mode. For example, if the size of the region of interest detected in the existence determination mode is large, the position is close to the center of the image of the object to be detected, or the estimated probability is above a predetermined threshold, the processing unit 120 determines that it should switch to the qualitative determination mode.
[0196] When the processing unit 120 determines that it needs to switch to a qualitative judgment mode, it calculates estimated probability information based on the second learned model NN1_2 and the image of the detected object, and determines control information to improve the calculated estimated probability information. In this way, since a learned model appropriate to the situation is selected, the result desired by the user can be obtained as the detection result, and by determining the control information, the reliability of the detection result can be improved.
[0197] For example, NN1_1 and NN1_2 are learned models that were trained using training data with different characteristics. For instance, the first learned model NN1_1 was learned based on training data that mapped a first learning image to information determining the presence or location of a region of interest within that image. The second learned model NN1_2 was learned based on training data that mapped a second learning image to information determining the state of a region of interest within that image.
[0198] In this way, by making the positive label in the training data different, the characteristics of NN1_1 and NN1_2 can be made different. Therefore, as NN1_1, a model can be generated specifically for determining whether a region of interest exists in the image of the object being detected. Furthermore, as NN1_2, a model can be generated specifically for determining the state of the region of interest in the image of the object being detected, such as determining which of the aforementioned lesion classification criteria it conforms to.
[0199] Furthermore, the first learning image can also be an image captured using white light. The second learning image can also be an image captured using special light with a wavelength different from white light, or an image captured with the subject magnified compared to the first learning image. This allows the information used as input in the training data to be different. Consequently, the characteristics of NN1_1 and NN1_2 can be different. Additionally, considering such a case, the processing unit 120 can also change the light source control information in the control information when switching from the existence determination mode to the qualitative determination mode. Specifically, the processing unit 120 performs the following control: illuminating ordinary light as illumination light in the existence determination mode, and illuminating special light as illumination light in the qualitative determination mode. This allows the switching of the learned model to be linked with the change of the control information.
[0200] Furthermore, the triggering of switching the learned model is not limited to the detection results of the region of interest in the existence determination mode. For example, the acquisition unit 110 may acquire operation information based on the user's operation of the endoscope device as described above, and the processing unit 120 may perform processing to select the learned model based on the operation information. In this case, it is also possible to select an appropriate learned model based on the kind of observation the user desires.
[0201] Furthermore, the switching of the learned model used for detection processing is not limited to viewpoints based on existence determination and qualitative determination. For example, the processing unit 120 can also perform processing to determine the photographed object captured in the detection object image, selecting a first learned model when the photographed object is a first photographed object, and selecting a second learned model when the photographed object is a second photographed object. Here, the photographed object is, for example, the organ being photographed. For example, the processing unit 120 selects the first learned model when the large intestine is photographed, and selects the second learned model when the stomach is photographed. However, the photographed object can also be a part where a single organ is further subdivided. For example, the learned model can be selected based on whether it is the ascending colon, transverse colon, descending colon, or S-shaped colon. In addition, the photographed object can also be distinguished by classifications other than organs.
[0202] Thus, by using learned models that vary depending on the subject being photographed, the accuracy of the detection process can be improved, i.e., the estimated probability can be increased. Therefore, in Figure 11 In the processing of parameters, it is easy to achieve the expected estimated probability.
[0203] In this scenario, the first learned model is learned based on training data obtained by mapping a first learning image of a first subject to information identifying regions of interest within that first learning image. The second learned model is learned based on training data obtained by mapping a second learning image of a second subject to information identifying regions of interest within that second learning image. This allows for the generation of learned models specifically designed for detecting regions of interest in each subject.
[0204] 3.5 Request for change of camera position and orientation
[0205] Furthermore, the above is an example of determining control information as a process to improve the estimated probability, but different processes can also be performed to improve the estimated probability.
[0206] For example, if it is determined that the probability represented by the estimated probability information is below a given threshold, the processing unit 120 may also perform a prompting process, which requests the user to change at least one of the position and orientation of the camera unit of the endoscope relative to the area of interest. Here, the camera unit is, for example, the imaging element 312. The change in the position and orientation of the camera unit corresponds to the change in the position and orientation of the front end of the insertion unit 310b.
[0207] For example, consider the case where the imaging element 312 captures the region of interest from an oblique direction. An oblique direction, for example, indicates that the difference between the optical axis direction of the objective lens optical system 311 and the normal direction of the subject surface is greater than a predetermined threshold. In this case, the shape of the region of interest in the image of the detected object is distorted; for example, its size in the image becomes smaller. In this case, due to the low resolution of the region of interest, even if the control information is changed, the estimation probability may not be sufficiently improved. Therefore, the processing unit 120 instructs the user to orient the imaging unit towards the subject. This instruction may also be displayed on the display unit 340, for example. Additionally, guidance displays related to the movement direction and amount of the imaging unit may be provided. In this way, the resolution of the region of interest in the image becomes higher, thus improving the estimation probability.
[0208] Furthermore, when the distance between the area of interest and the camera is large, the area of interest in the image is small and appears dark. In this case, it may be impossible to sufficiently improve the estimated probability when adjusting the control information. Therefore, the processing unit 120 instructs the user to move the camera closer to the subject.
[0209] As described above, the processing unit 120 can trigger a request to change the position and posture of the camera unit by using the case where the estimated probability does not reach a threshold through changes in the control information as a trigger. This prioritizes changes to the control information, thus reducing the burden on the user when the situation can be handled by such changes. Furthermore, by enabling position and posture change requests, the estimated probability can be improved even when changes to the control information cannot address the issue.
[0210] Furthermore, the triggering of the prompt processing for requesting a change in position and pose is not limited to the above. For example, the NN2 mentioned above can also be a learned model that outputs both information determining control information and information determining whether to make a request for a change in position and pose.
[0211] 3.6 Display processing or storage processing
[0212] Consider the case where the estimated probability of using the first control information is less than the threshold when the control information is changed sequentially to the first control information and then to the second control information.
[0213] The processing unit 120 can also perform processing that displays the detection result of the region of interest based on the first detection object image and skips the display of the first estimated probability information. Furthermore, the processing unit 120 can also perform the following processing: calculate second estimated probability information based on the second detection object image obtained using control with second control information and NN1; and if the second estimated probability information is above a given threshold, display the detection result of the region of interest based on the second detection object image and the second estimated probability information.
[0214] Furthermore, if the estimated probability is less than a threshold when using the first and second control information, and greater than or equal to the threshold when using the third control information, the processing unit 120 displays the detection result of the region of interest in the second detection object image and skips the display of the second estimated probability information. Moreover, the processing unit 120 can also perform processing to display the detection result of the region of interest in the third detection object image obtained based on control using the third control information, and the third estimated probability information calculated based on the third detection object image.
[0215] In this embodiment, even if the estimated probability at a given point in time is low, the estimated probability can be improved using the method described above. For example, even if the estimated probability is less than a threshold, there is a possibility that the estimated probability may become above the threshold in the future by changing the control information. Therefore, in this embodiment, the detection result of the region of interest can be displayed even when the estimated probability is less than the threshold. This suppresses the possibility of missing the region of interest. At this time, the value of the estimated probability changes over time by repeatedly performing the loop processing of the decision control information. As a user, you want to know whether the displayed region of interest is sufficiently reliable, so you are less concerned about the temporal changes in the estimated probability. Therefore, the processing unit 120, for example, displays only the detection box and not the estimated probability in the control information update loop, and displays the estimated probability together with the detection box when the estimated probability is above the threshold. This allows for a display that is easy for the user to understand.
[0216] Furthermore, while the method of this embodiment can improve the estimated probability, if the original estimated probability is too low, even if the control information is changed, the estimated probability may not reach a threshold. Therefore, the processing unit 120 may also set a second threshold smaller than the aforementioned threshold. The processing unit 120 performs the following processing: if the estimated probability is higher than or lower than the second threshold, only the detection box is displayed; if the estimated probability is higher than the threshold, both the detection box and the estimated probability are displayed.
[0217] Furthermore, the above examples illustrate how to match the detection results of the region of interest with the displayed image, but the processing performed when the estimated probability is above a threshold is not limited to display processing.
[0218] For example, the processing unit 120 can also store the detected object image if the estimated probability information calculated based on the learned model NN1 and the detected object image is above a given threshold. This allows for the accumulation of images with high visual recognizability of areas considered to be of interest. Typically, during observation using the endoscope system 300, still images are not saved unless the user explicitly presses the shutter button. Therefore, it is possible not to store images obtained by capturing areas of interest such as lesions. Alternatively, storing dynamic images is also considered, but in this case, the number of images becomes large, resulting in a heavy processing load for searching for areas of interest. In this regard, by performing storage processing based on an estimated probability exceeding a threshold, information related to the area of interest can be appropriately stored.
[0219] Furthermore, the processing unit 120 can also process and store the display image corresponding to the detected object image if the estimated probability information calculated based on the learned model NN1 and the detected object image is above a given threshold. In this way, images viewed by the user can also be used as objects for storage processing.
[0220] The embodiments and their variations have been described above. However, this application is not directly limited to these embodiments and their variations. During implementation, the constituent elements can be modified and specified without departing from the spirit of the invention. Furthermore, multiple constituent elements disclosed in the above embodiments and variations can be appropriately combined. For example, several structural elements can be deleted from all structural elements described in the embodiments and variations. Moreover, the constituent elements described in different embodiments and variations can be appropriately combined. Thus, various modifications and applications can be realized without departing from the spirit of the invention. Additionally, in the specification or drawings, any term described at least once with a more general or synonymous term can be replaced with its different term at any point in the specification or drawings.
[0221] Label Explanation
[0222] 100…Processing system, 110…Acquisition unit, 120…Processing unit, 130…Control unit, 300…Endoscope system, 310…Endoscope body, 310a…Operating unit, 310b…Insertion unit, 310c…Universal cable, 310d…Connector, 311…Objective lens optical system, 312…Image sensor, 314…Illumination lens, 315…Light guide, 330…Processing device, 331…Preprocessing unit, 332…Control unit, 333…Storage unit, 334…Detection processing unit, 335…Control information determination unit, 336…Postprocessing unit, 340…Display unit, 350…Light source device, 352…Light source, 400…Learning device, 410…Acquisition unit, 420…Learning unit
Claims
1. A processing system, characterized by, The processing system includes: an acquisition unit that acquires a detection target image captured by an endoscope device; and a processing unit that detects a region of interest included in the detection target image based on the detection target image and a learned model for detection processing that is obtained through machine learning and that calculates estimation probability information representing a likelihood of the region of interest within an input image, calculates the estimation probability information related to the detected region of interest, the processing unit determines control information for improving the estimation probability information related to the region of interest within the detection target image based on the detection target image, the processing system further includes a control unit, the control unit controls the endoscope device in accordance with the determined control information.
2. The processing system according to claim 1, wherein the processing unit calculates first estimation probability information based on a first detection target image obtained based on control using first control information and the learned model, the processing unit performs processing of determining second control information for improving the estimation probability information in a case where it is determined that the likelihood represented by the first estimation probability information is lower than a given threshold value.
3. The processing system according to claim 2, wherein the processing unit calculates second estimation probability information based on a second detection target image obtained based on control using the second control information and the learned model, the processing unit performs processing of determining third control information for improving the estimation probability information in a case where it is determined that the likelihood represented by the second estimation probability information is lower than the given threshold value.
4. The processing system according to claim 1, wherein the control information includes at least one of light source control information for controlling a light source that irradiates an object with illumination light, imaging control information for controlling an imaging condition for imaging the detection target image, and image processing control information for controlling image processing for an imaged image signal.
5. The processing system according to claim 4, wherein the control information is information representing at least one of a hue, a brightness, and a position within a screen related to the region of interest.
6. The processing system according to claim 2, wherein the control information includes first to Nth parameter categories, where N is an integer of two or more, the processing unit determines the second control information by changing two or more of the first to Nth parameter categories included in the first control information.
7. The processing system according to claim 2, wherein the control information includes first to Nth parameter categories, where N is an integer of two or more, the processing unit determines second control information by changing an i-th parameter category included in the first control information, where i is an integer satisfying 1 ≤ i ≤ N.
8. The processing system according to claim 7, wherein the processing section calculates second estimation probability information based on a second detection target image obtained based on control using the second control information and the learned model, the processing section determines third control information by changing a j-th parameter category among the first to N-th parameter categories included in the second control information, where j is an integer satisfying 1 ≤ j ≤ N, j ≠ i, in a case where it is determined that the likelihood indicated by the second estimation probability information is lower than the given threshold value.
9. The processing system according to claim 1, wherein the processing section calculates the estimation probability information based on a first detection target image obtained based on control using first control information and the learned model, the processing section performs processing of determining the control information, that is, second control information, for improving the estimation probability information, based on the first control information and the first detection target image.
10. The processing system according to claim 1, wherein the processing section is capable of performing: first processing of detecting the attention region included in the detection target image based on a first learned model and the detection target image; and second processing of detecting the attention region included in the detection target image based on a second learned model and the detection target image, the processing section performs processing of selecting an arbitrary learned model from among a plurality of learned models including the first learned model and the second learned model, in a case where the first learned model is selected, the processing section performs the first processing based on the first learned model and the detection target image and calculates the estimation probability information, and determines the control information for improving the calculated estimation probability information, in a case where the second learned model is selected, the processing section performs the second processing based on the second learned model and the detection target image and calculates the estimation probability information, and determines the control information for improving the calculated estimation probability information.
11. The processing system according to claim 10, wherein the acquisition section acquires operation information based on a user operation of the endoscope device, the processing section performs the processing of selecting the learned model based on the operation information.
12. The processing system according to claim 10, wherein the processing section is capable of operating in an arbitrary mode among a plurality of judgment modes, the plurality of judgment modes include a presence determination mode of determining the presence or absence of the attention region included in the detection target image based on the first learned model and the detection target image, and a qualitative determination mode of determining a state of the attention region included in the detection target image based on the second learned model and the detection target image.
13. The processing system according to claim 12, wherein The processing section determines whether to shift to the qualitative determination mode based on a detection result of the region of interest in the presence determination mode, The processing section calculates the estimation probability information based on the second learned model and the detection target image when it is determined that the qualitative determination mode is to be shifted to, and determines the control information for improving the calculated estimation probability information.
14. The processing system according to claim 12, wherein a first learning image for learning of the first learned model is an image captured using white light, a second learning image for learning of the second learned model is an image captured using special light having a wavelength band different from that of the white light or an image captured in a state where a subject is enlarged compared to the first learning image.
15. The processing system according to claim 10, wherein the processing section performs processing of a captured object captured in the detection target image, the processing section selects the first learned model when the captured object is a first captured object and selects the second learned model when the captured object is a second captured object.
16. The processing system according to claim 1, wherein the acquisition section acquires the detection target image as a detection target of the region of interest and a calculation target of the estimation probability information in a given frame, the acquisition section acquires a display image for display in a display section in a frame different from the given frame.
17. The processing system according to claim 1, wherein the processing section performs processing of displaying a detection result of the region of interest and the estimation probability information when the estimation probability information calculated based on the learned model and the detection target image is equal to or higher than a given threshold value.
18. The processing system according to claim 2, wherein the processing section performs processing of displaying a detection result of the region of interest based on the first detection target image and skipping display of the first estimation probability information, the processing section performs processing of calculating second estimation probability information based on a second detection target image acquired based on control using the second control information and the learned model, and displaying a detection result of the region of interest based on the second detection target image and the second estimation probability information when the second estimation probability information is equal to or higher than a given threshold value.
19. The processing system according to claim 16, wherein the processing section performs processing of displaying a region of the detection target image in which the region of interest is included in association with the display image when the estimation probability information calculated based on the learned model and the detection target image is equal to or higher than a given threshold value.
20. The processing system according to claim 1, wherein The processing section performs prompting processing that requests a user to change at least one of a position and a posture of an imaging section of the endoscope device with respect to the region of interest, in a case where the likelihood indicated by the estimation probability information is determined to be lower than a given threshold value.
21. An image processing method, characterized by, acquiring a detection target image captured by an endoscope device, detecting a region of interest included in the detection target image based on the detection target image and a learned model for detection processing that is acquired through machine learning and that is used to calculate estimation probability information indicating a likelihood of the region of interest within an input image, calculating the estimation probability information related to the detected region of interest, when information for controlling the endoscope device is set as control information, determining the control information for improving the estimation probability information related to the region of interest within the detection target image based on the detection target image.
22. A learning method, characterized by, acquiring an image captured by an endoscope device as an input image, when information for controlling the endoscope device is set as control information, acquiring first control information that is the control information at the time of acquiring the input image, acquiring second control information that is the control information for improving estimation probability information indicating a likelihood of a region of interest detected from the input image, generating a learned model used in determination processing of the control information by machine learning of a relationship among the input image, the first control information, and the second control information.
23. A processing device, comprising: The processing device includes: an acquisition section that acquires a detection target image captured by an endoscope device; and a processing section that detects a region of interest included in the detection target image based on the detection target image and a learned model for detection processing that is acquired through machine learning and that is used to calculate estimation probability information indicating a likelihood of the region of interest within an input image, calculates the estimation probability information related to the detected region of interest, the processing section determines control information for improving the estimation probability information related to the region of interest within the detection target image based on the detection target image, and outputs the determined control information, wherein the determined control information is used to control the endoscope device.
Citation Information
Patent Citations
Image diagnosis assistance apparatus, data collection method, image diagnosis assistance method, and image diagnosis assistance program
WO2019088121A1
Endoscope system and method for operating same
CN110325100A
Medical image processing device, medical image processing method, and medical image processing program
WO2019054045A1