Using loss function to learn detection model

By using the weighted loss function in image analysis to calculate the error between the output image and the annotated image, the problem of insufficient accuracy of the detection model in the prior art is solved, and higher detection accuracy and model learning effect are achieved.

CN113614780BActive Publication Date: 2025-05-06INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080020091.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-10
Filing Date
2020-03-13
Publication Date
2025-05-06
Estimated Expiration
2040-03-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively learn detection models in image analysis, especially in areas of interest where high accuracy is required.

Method used

The detection model is updated by calculating the error between the output image and the annotated image using the loss function and weighting the error within the region of interest to reduce the error to a greater extent.

Benefits of technology

The accuracy of the detection model in the region of interest is improved, the recognition ability of the target region is enhanced, and the detection model can be effectively learned even when there are errors in the annotated image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113614780B_ABST
    Figure CN113614780B_ABST
Patent Text Reader

Abstract

It is desirable to accurately learn a detection model (130). A computer-implemented method is provided, comprising: obtaining an input image; obtaining an annotated image that specifies a region of interest in the input image; inputting the input image into a detection model (130), the detection model generating an output image from the input image showing the target region; calculating an error between the output image and the annotated image using a loss function, the loss function weighting errors within the region of interest more than errors outside the region of interest; and updating the detection model in a manner that reduces the error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention involves using a loss function to learn a detection model. Background Art

[0002] Image analysis is important in many fields. One non-limiting example is analyzing X-ray images to identify tumors, where having an effective detection model to detect diseases such as lung tumors from X-ray images is useful.

[0003] Therefore, there is a need in the art to solve the above problems. Summary of the invention

[0004] From a first aspect, the present invention provides a computer-implemented method, the method comprising: obtaining an input image; obtaining an annotated image specifying a region of interest in the input image; inputting the input image into a detection model, the detection model generating an output image showing the target region from the input image; using a loss function to calculate the error between the output image and the annotated image, the loss function weighting errors within the region of interest more heavily than errors outside the region of interest; and updating the detection model in a manner that reduces the error.

[0005] From another aspect, the present invention provides a device comprising: a processor or a programmable circuit; and one or more computer-readable media, which collectively include instructions, which, in response to being executed by the processor or the programmable circuit, cause the processor or the programmable circuit to: acquire an input image; acquire an annotated image specifying a region of interest in the input image; input the input image into a detection model, the detection model generating an output image showing the target region from the input image; using a loss function to calculate the error between the output image and the annotated image, the loss function weighting errors within the region of interest more heavily than errors outside the region of interest; and updating the detection model in a manner that reduces the error.

[0006] Viewed from another aspect, the present invention provides a computer program product for image analysis, the computer program product comprising a computer readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit to perform a method for performing the steps of the present invention.

[0007] Viewed from another aspect, the invention provides a computer program stored on a computer readable medium and loadable into the internal memory of a digital computer, the computer program comprising software code portions for performing the steps of the invention when the program is run on a computer.

[0008] From another aspect, the present invention provides a computer program product comprising one or more computer-readable storage media that collectively store program instructions that can be executed by a processor or a programmable circuit to cause the processor or the programmable circuit to perform operations, the operations comprising: acquiring an input image; acquiring an annotated image that specifies a region of interest in the input image; inputting the input image into a detection model that generates an output image showing the target region from the input image; using a loss function to calculate the error between the output image and the annotated image, the loss function weighting errors within the region of interest specified by the annotated image more heavily than errors outside the specified region of interest; and updating the detection model in a manner that reduces the error.

[0009] From another aspect, the present invention provides a device, comprising: a processor or a programmable circuit; and one or more computer-readable media, which collectively include instructions, which, in response to being executed by the processor or the programmable circuit, cause the processor or the programmable circuit to: implement a neural network for generating an output image showing a target area from an input image, wherein the neural network includes multiple convolutional layers, multiple pooling layers, multiple deconvolution layers, and multiple batch normalization layers between the input and the output; and the multiple batch normalization layers are arranged every predetermined number of layers in at least one of a first path including some of the multiple convolutional layers and the multiple pooling layers and a second path including the remaining convolutional layers and the multiple deconvolution layers.

[0010] From another aspect, the present invention provides a computer program product, which includes one or more computer-readable storage media, which collectively store program instructions, which can be executed by a processor or a programmable circuit to cause the processor or the programmable circuit to perform operations, including: implementing a neural network, which is used to generate an output image showing a target area from an input image, wherein the neural network includes multiple convolutional layers, multiple pooling layers, multiple deconvolution layers and multiple batch normalization layers between the input and the output; and multiple batch normalization layers are arranged every predetermined number of layers in at least one path of a first path including some of the multiple convolutional layers and the multiple pooling layers and a second path including the remaining convolutional layers and the multiple deconvolution layers.

[0011] According to an embodiment of the present invention, a computer-implemented method is provided, the method comprising: obtaining an input image; obtaining an annotated image specifying a region of interest in the input image; inputting the input image into a detection model, the detection model generating an output image showing the target region from the input image; using a loss function to calculate the error between the output image and the annotated image, the loss function weighting errors inside the region of interest more heavily than errors outside the region of interest; and updating the detection model in a manner that reduces the error.

[0012] Weighted cross entropy can be used as a loss function. In this way, the error can be weighted using weighted cross entropy.

[0013] The detection model may include one or more convolutional layers, one or more pooling layers, one or more deconvolutional layers, and one or more batch normalization layers between the input and the output. In this way, in some embodiments, the detection model accelerates the convergence of learning and limits over-learning.

[0014] The computer-implemented method may further include obtaining at least one coordinate in the input image; and specifying the region of interest based on the at least one coordinate. In this way, the region of interest can be roughly specified from the at least one coordinate.

[0015] According to another embodiment of the present invention, a device is provided, the device including a processor or a programmable circuit; and one or more computer-readable media, which collectively include instructions, in response to being executed by the processor or the programmable circuit, the instructions causing the processor or the programmable circuit to acquire an input image; acquire an annotation image specifying a region of interest in the input image; input the input image to a detection model, the detection model generating an output image showing a target region from the input image; using a loss function to calculate the error between the output image and the annotation image, the loss function weighting the error within the region of interest more than the error outside the region of interest; and updating the detection model in a manner that reduces the error. In this way, the processor or programmable circuit acquires the input image and accurately learns the detection model.

[0016] According to another embodiment of the present invention, a computer program product is provided, the computer program product comprising one or more computer-readable storage media, the one or more computer-readable storage media collectively storing program instructions, the program instructions executable by a processor or programmable circuit to cause the processor or programmable circuit to perform operations, the operations comprising: obtaining an input image; obtaining an annotation image specifying an area of ​​interest in the input image; inputting the input image to a detection model, the detection model generating an output image showing the target area from the input image; using a loss function to calculate the error between the output image and the annotation image, the loss function weighting the error within the area of ​​interest specified by the annotation image more than the error outside the specified area of ​​interest; and updating the detection model in a manner that reduces the error. In this way, the accuracy associated with the machine learning of the detection model can be increased.

[0017] According to another embodiment of the present invention, a device is provided, which includes a processor or a programmable circuit; and one or more computer-readable media, which together include instructions, which, in response to being executed by the processor or the programmable circuit, cause the processor or the programmable circuit to implement a neural network, wherein the neural network includes a plurality of convolutional layers, a plurality of pooling layers, a plurality of deconvolutional layers, and a plurality of batch normalization layers between the input and the output, and the plurality of batch normalization layers are respectively arranged after every predetermined number of layers in at least one of a first path including some of the plurality of convolutional layers and the plurality of pooling layers and a second path including the remaining convolutional layers and the plurality of deconvolutional layers.

[0018] According to another embodiment of the present invention, a computer program product is provided, the computer program product includes one or more computer-readable storage media, the one or more computer-readable storage media collectively storing program instructions, the program instructions can be executed by a processor or a programmable circuit to make the processor or the programmable circuit perform operations including the following: implement a neural network, the neural network is used to generate an output image showing a target area from an input image, wherein the neural network includes a plurality of convolutional layers, a plurality of pooling layers, a plurality of deconvolutional layers, and a plurality of batch normalization layers between the input and the output, and in at least one of a first path including some of the convolutional layers and the plurality of pooling layers and a second path including the remaining convolutional layers and the plurality of deconvolutional layers of the plurality of convolutional layers, a plurality of batch normalization layers are respectively arranged after each predetermined number of layers. In this way, in some embodiments, the detection model accelerates the convergence of learning and limits over-learning.

[0019] Each summary of the invention does not necessarily describe all necessary features of the embodiments of the present invention. The present invention may also be a sub-combination of the above features. The above and other features and advantages of the present invention will become more apparent through the following description of the embodiments in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The present invention will now be described, by way of example only, with reference to preferred embodiments as shown in the following drawings:

[0021] Figure 1 is a functional block diagram illustrating a computing environment according to an exemplary embodiment of the present invention, in which a system for learning a detection model utilizes a loss function.

[0022] Figure 2 According to at least one embodiment of the present invention, Figure 1 An exemplary network of a detection model within an environment.

[0023] Figure 3 According to at least one embodiment of the present invention, Figure 1 The operation process of the detection model is executed on the computing device within the environment.

[0024] Figure 4 An example of comparison between an output image according to the present embodiment and an output image obtained by a conventional manner is shown.

[0025] Figure 5 A functional block diagram illustrating a computing environment showing modifications to a system for learning a detection model using a loss function in accordance with at least one embodiment of the present invention is shown.

[0026] Figure 6 An example of specifying a region of interest by the apparatus 100 according to the present modification is shown.

[0027] Figure 7 A functional block diagram illustrating a computing environment that is yet another modification of a system for learning a detection model using a loss function in accordance with at least one embodiment of the present invention is shown.

[0028] Figure 8 According to at least one embodiment of the present invention Figure 1 A block diagram of components of one or more computing devices within a computing environment is depicted in FIG. DETAILED DESCRIPTION

[0029] Hereinafter, some embodiments of the present invention will be described. The embodiments do not limit the invention as described in the claims, and all combinations of features described in the embodiments may include, but are not limited to, embodiments provided by features disclosed in the specification.

[0030] Figure 1 1 is a functional block diagram of the device 100 according to the present embodiment. The device 100 can be a computer, such as a PC (personal computer), a tablet computer, a smart phone, a workstation, a server computer or a general-purpose computer, a computer system connecting multiple computers, or any programmable electronic device or a combination of programmable electronic devices that can execute a system for learning a detection model using a loss function. Such a computer system is also a computer in a broad sense. In addition, the device 100 can be implemented by one or more virtual computer environments that can be executed in a computer system. On the contrary, the device 100 can be a special-purpose computer designed to update the detection model, or can include special-purpose hardware utilized by a special-purpose circuit. In addition, if the device 100 can be connected to the Internet, the device 100 can be used in cloud computing.

[0031] The device 100 uses a loss function to calculate the error between the annotated image and the output image generated by inputting the input image into the detection model, and updates the detection model in a manner that reduces this error. In doing so, the device 100 according to the present embodiment learns the detection model by weighting the error in the region of interest (ROI) more heavily than the error outside the region of interest (ROI) as a loss function. In the present embodiment, as an example, a case is shown in which the device 100 learns a detection model for detecting diseases such as lung tumors from an X-ray image of an animal. However, the detection model includes this example, but is not limited to the scope of identifying the features of tissue structures (e.g., tumors in animals). The device 100 can learn detection models for detecting various regions from various images. The device 100 includes an input image acquisition unit 110, an annotation image acquisition unit 120, a detection model 130, an error calculation unit 140, a weight storage unit 150, and an update unit 160.

[0032] The input image acquisition unit 110 retrieves a plurality of input images, which are considered "raw" images in some embodiments for the purpose of clarity of understanding. For example, in one embodiment, the input image acquisition unit 110 retrieves a plurality of X-ray images of an animal. The input image acquisition unit 110 communicates with the detection model 130 so that the detection model 130 is provided with a plurality of acquired input images.

[0033] The annotated image acquisition unit 120 retrieves an image of an annotation that specifies a region of interest in the input image for each of the plurality of input images. For example, the annotated image acquisition unit 120 may retrieve an image of each annotation in which a mark indicating that a region is different from other regions is attached to the region of interest in the input image. The annotated image acquisition unit 120 transmits the acquired annotated image to the error calculation unit 140.

[0034] The detection model 130 receives a plurality of input images retrieved from the input image acquisition unit 110, and generates each output image showing the target area based on these input images. In the present embodiment, a multilayer neural network described further below is used as the algorithm of the detection model 130. However, the detection model 130 is not limited to the present embodiment. On the contrary, in addition to the models described below, neural networks such as CNN, FCN, SegNet and U-Net or any algorithm capable of detecting the target area such as support vector machines (SVM) and determination trees can be used as the detection model 130. The detection model 130 transmits the generated output image to the error calculation unit 140.

[0035] The error calculation section 140 calculates the corresponding error between the output image supplied from the detection model 130 and the annotation image supplied from the annotation image acquisition section 120 using a loss function. When doing so, the error calculation section 140 uses weighted cross entropy as a loss function to weight the error within the region of interest more heavily than the error outside the region of interest. In other words, as an example, the error calculation section 140 weights the error within the region designated by the doctor as a lung tumor more than the error outside the region. This is described in further detail below.

[0036] The weight storage section 150 stores weights in advance, and supplies the weights to the error calculation section 140. The error calculation section 140 calculates errors using the weights using the weighted cross entropy.

[0037] The error calculation section 140 applies the weights supplied from the weight storage section 150 to the weighted cross entropy to calculate errors between the output image and the annotated image, and supplies the errors to the updating section 160 .

[0038] The updating unit 160 updates the detection model 130 in a manner that reduces the error provided from the error calculation unit 140. For example, the updating unit 160 updates each parameter of the detection model in a manner that minimizes the error using an error back propagation technique. Those skilled in the art will recognize that the error back propagation technique itself is widely accepted as an algorithm used when learning a detection model.

[0039] Figure 2An exemplary network of the detection model 130 according to the present embodiment is shown. As an example, the detection model 130 includes one or more convolutional layers, one or more pooling layers, one or more deconvolutional layers, and one or more batch normalization layers between the input and the output. In the present figure, as an example, the detection model 130 includes, between the input layer 210 and the output layer 260, a plurality of convolutional layers 220a to 220w (collectively referred to as "convolutional layers 220"), a plurality of pooling layers 230a to 230e (collectively referred to as "pooling layers 230"), a plurality of deconvolutional layers 240a to 240e (collectively referred to as "deconvolutional layers 240"), and a plurality of batch normalization layers 250a to 250i (collectively referred to as "batch normalization layers 250").

[0040] The convolution layers 220 each output a feature map by performing a convolution operation applied to an input image while sliding a kernel (filter) having a predetermined size.

[0041] The pooling layers 230 each compress and downsample the information to deform the input image into a shape that is easier to process. In this case, for example, each pooling layer 230 may use maximum pooling to select and compress the maximum value of each range, or may use average pooling to calculate and compress the average value of each range.

[0042] The deconvolution layers 240 each expand the size by adding blanks around and / or between each element in a feature map input thereto, and then perform a deconvolution operation applied while sliding a kernel (filter) having a predetermined size.

[0043] The batch normalization layers 250 each replace the output of each unit with a new value normalized for each mini-batch. In other words, each batch normalization layer 250 performs normalization so that the elements of each image in the plurality of images have a normalized distribution.

[0044] The detection model 130 according to the present embodiment has a first path and a second path, the first path is an encoding path including a plurality of convolutional layers 220a to 220l and a plurality of pooling layers 230a to 230e, and the second path is a decoding path including a plurality of convolutional layers 220m to 220w and a plurality of deconvolutional layers 240a to 240e, for example. The encoding path and the decoding path each connect feature maps having the same dimension.

[0045] In addition, in the detection model 130 according to the present embodiment, in at least one of a first path including some of the plurality of convolutional layers 220 and a plurality of pooling layers 230 and a second path including the remaining convolutional layers 220 and a plurality of deconvolutional layers 240, a plurality of batch normalization layers 250 are arranged every predetermined number of layers.

[0046] For example, a plurality of batch normalization layers 250a to 250d are respectively arranged at each predetermined number of layers in the encoding path, and a plurality of batch normalization layers 250e to 250i are respectively arranged at each predetermined number of layers in the decoding path. More specifically, as shown in the present figure, in the encoding path, a plurality of batch normalization layers 250a to 250d are respectively arranged after convolutional layers 220d, 220f, 220h, and 220j connected to the decoding path. In addition, in the decoding path, a plurality of batch normalization layers 250e to 250h are respectively arranged after convolutional layers 220n, 220p, 220r, 220t, and 220v.

[0047] In this way, according to the detection model 130 of the present embodiment, at least one batch normalization layer 250 is set in at least one of the encoding path and the decoding path, so that a large change in the internal variable distribution (internal covariate shift) can be prevented and the convergence of learning can be accelerated, and over-learning can also be restricted. In addition, by using such a detection model 130, when learning the correct label classification from a label containing an error, the device 100 according to the present embodiment can learn the detection model 130 more effectively.

[0048] Figure 3 The operation process of the device 100 learning the detection model 130 according to this embodiment is shown.

[0049] In step 310, the device 100 acquires a plurality of input images. For example, the input image acquisition unit 110 acquires a plurality of X-ray images of an animal. If the sizes of the images are different from each other, the input image acquisition unit 110 may acquire images that have each undergone preprocessing (e.g., normalization of pixel values, cropping to a predetermined shape, and resizing to a predetermined size) as input images. At this time, for example, the input image acquisition unit 110 may acquire a plurality of input images via the Internet, via user input, or via a memory device capable of storing data. The input image acquisition unit 110 provides the plurality of acquired input images to the detection model 130.

[0050] In step 320, the apparatus 100 acquires an annotation image that specifies a region of interest in the input image for each of the plurality of input images. For example, the annotation image acquisition unit 120 acquires each annotation image, wherein, for example, pixels inside a region designated as a lung tumor by a doctor are labeled with class c=1, and pixels outside the region are labeled with class c=0. In the above description, an example is shown in which each pixel is labeled according to two classes (e.g., c=0 and c=1) based on whether the pixel is designated as inside or outside a region of a lung tumor, but the present embodiment is not limited thereto. For example, the annotation image acquisition unit 120 may acquire each annotation image, wherein pixels inside a lung region and pixels inside a lung tumor region are labeled with class c=2, pixels inside a lung region but outside a lung tumor region are labeled with class c=1, and pixels outside a lung region and outside a lung tumor region are labeled with class c=0. In other words, the annotation image acquisition unit 120 may acquire each annotation image labeled according to three or more classes (e.g., c=0, c=1, and c=2). At this time, the annotation image acquisition section 120 may acquire each annotation image via a network, via user input, or via a memory device capable of storing data, etc. The annotation image acquisition section 120 provides the acquired annotation image to the error calculation section 140 .

[0051] In step 330, the device 100 inputs each of the multiple input images acquired in step 310 into the detection model 130, which generates an output image of the target area from these input images. For example, for each of the multiple input images, the detection model 130 generates an output image, wherein the pixels within the area are predicted to be the target area, that is, the lung tumor area in the input image is labeled with class c=1, and the pixels outside the area are labeled with class c=0. Then, the detection model 130 provides each generated output image to the error calculation unit 140. In the above description, an example is shown in which the device 100 generates an output image after acquiring an annotated image, but the device 100 is not limited thereto. The device 100 may acquire an annotated image after generating the output image. In other words, step 320 may be performed after step 330.

[0052] In step 340, the device 100 calculates the corresponding error between the output image acquired in step 330 and the annotated image acquired in step 320 using a loss function. At this time, the error calculation unit 140 uses cross entropy as a loss function. Generally, cross entropy is a scale defined between two probability distributions as a probability distribution and a reference fixed probability distribution, has a minimum value when the probability distribution and the reference fixed distribution are the same, and has a larger value as the probability distribution is different from the reference fixed distribution. Here, when calculating the error between the output image and the annotated image, if the cross entropy is used without weighting any pixel in the input image, the error is strongly affected by a pixel group with a wide area. When this happens, for example, if an error is included in the annotated image, there is a case where the detection model for detecting the target area cannot be accurately learned even if the detection model is updated in order to minimize the error.

[0053] Therefore, in the present embodiment, the error calculation part 140 is used to calculate the error of the loss function for the error within the region of interest more than the weight of the error outside the region of interest. At this time, the error calculation part 140 can use weighted cross entropy as the loss function. For example, the weight storage part 150 pre-stores the weights to be applied to the weighted cross entropy, and provides these weights to the error calculation part 140. The error calculation part 140 applies the weights provided from the weight storage part 150 to the weighted cross entropy to calculate the error between the output image and the annotated image. Here, when the error inside the region of interest is weighted more heavily than the error outside the region of interest, in the weighted cross entropy, the weight in the target region can be set to a value greater than the weight outside the target region. On the contrary, in the weighted cross entropy, the weight inside the region of interest can be set to a value greater than the weight outside the region of interest. This is explained using a mathematical expression.

[0054] For example, the error calculation section 140 uses the weighted cross entropy shown by the following expression as the loss function. It should be noted that X is the set of all pixels "i" in the input image, C is the set of all classes c, and W c is the weight of class c, p i c is the value of category c at pixel “i” in the annotated image, and q i c is the value of category c at pixel “i” in the output image.

[0055] Expression 1:

[0056]

[0057] Here, the following example is described in which the annotation image acquisition unit 120 acquires an annotation image in which pixels inside a region of interest (ROI) designated by a doctor as a lung tumor are labeled as class c=1 and pixels outside the region of interest are labeled as class c=0, and the detection model 130 generates an output image in which class c=1 is given to pixels within a target region predicted to be a lung tumor region in the input image, and class c=0 is given to pixels outside the target region.

[0058] In this case, Expression 1 is expanded to the expression shown below. Specifically, the cross entropy is expanded to the sum of four components, which are (1) the component in which the pixels inside the region of interest are marked with c=1 and predicted to be inside the target area, (2) the component in which the pixels inside the region of interest are marked with c=0 and predicted to be outside the target area, (3) the component in which the pixels outside the region of interest are marked with c=1 and predicted to be inside the target area, and (4) the component in which the pixels outside the region of interest are marked with c=0 and predicted to be outside the target area.

[0059] Expression 2:

[0060]

[0061] However, in the above case, the probability that the pixel inside the region of interest is marked with c=0 is 0, and therefore the second component in Expression 2 is 0. Similarly, the probability that the pixel outside the region of interest is marked with c=1 is 0, and therefore the third component in Expression 2 is 0. The probability that the pixel inside the region of interest is marked with c=1 and the probability that the pixel outside the region of interest is marked with c=0 are both 1, and therefore Expression 2 can be expressed as the following expression. Specifically, the cross entropy is shown as the sum of two components, which are a component in which the pixel inside the region of interest is predicted to be inside the target region and a component in which the pixel outside the region of interest is predicted to be outside the target region.

[0062] Expression 3:

[0063]

[0064] In the present embodiment, the loss function weights the error inside the region of interest more heavily than the error outside the region of interest. In other words, in Expression 3, the weight Wc of the first component is set to a value greater than the weight Wc of the second component. Here, in the weighted cross entropy, by setting the weight inside the target region to a larger value than the weight outside the target region, the error inside the region of interest can be weighted more heavily than the error outside the region of interest. In other words, by setting the weights of the first component and the third component in Expression 2 to a larger value than the weights of the second component and the fourth component in Expression 2, the weight of the first component in Expression 3 can be set to a larger value than the weight of the second component in Expression 3. Conversely, in the weighted cross entropy, by setting the weight inside the region of interest to a larger value than the weight outside the region of interest, the error inside the region of interest can be weighted more heavily than the error outside the region of interest. In other words, by setting the weights of the first and second components in Expression 2 to values ​​greater than the weights of the third and fourth components in Expression 2, the weight of the first component in Expression 3 can be set to a value greater than the weight of the second component in Expression 3. Then, the error calculation section 140 provides the updating section 160 with the error between the output image and the annotation image calculated in this manner.

[0065] In step 350, the apparatus 100 updates the detection model 130 in a manner to reduce the error calculated in step 340. For example, the updating unit 160 updates the error using an error back propagation technique in a manner to minimize each parameter of the detection model 130, and then ends the process.

[0066] The apparatus 100 may repeatedly perform the process from step 330 to step 350 according to different learning methods (eg, batch learning, mini-batch learning, and online learning).

[0067] Figure 4An example of comparison between an output image according to the present embodiment and an output image obtained by a conventional method is shown. In the present figure, an input image, an output image obtained by a conventional method, an output image obtained according to the present embodiment, and an overlapping image in which the output image obtained according to the present embodiment is overlapped on the input image are shown in the order described from left to right. As shown in the present figure, using a detection model according to a conventional method, it is impossible to classify the input image and make predictions between the lung tumor area and other areas. In contrast, using the detection model 130 learned by the device 100 according to the present embodiment, an output in which the lung tumor area (the area shown in white in the figure) and other areas (the area shown in black in the figure) are divided according to categories can be obtained. When observing the image in which the output is superimposed on the input image, the target area predicted by the detection model 130 matches the actual lung tumor area for the most part, and the reliability of the detection model 130 can be confirmed.

[0068] In this way, according to the device 100 of the present embodiment, when the loss function is used to calculate the error between the output image and the annotated image, the loss function weights the error within the region of interest more heavily than the error outside the region of interest. Thus, the device 100 can (i) enhance the effect caused by the pixels inside the region of interest, and (ii) calculate the error between the output image and the annotated image. Then, the detection model is updated in a manner that minimizes the error calculated in this manner, and thus, for example, even if an error is included in the annotated image, the device 100 can accurately learn the detection model 130 for detecting the target area. In other words, the device 100 can learn the detection model 130 by weakly supervised learning that learns the correct label classification from the label including the error.

[0069] Figure 5 FIG. 1 shows a functional block diagram of a modified device 100 according to this embodiment. Figure 1 Components having the same functions and configurations are given the same reference numerals, and the following description includes only the different points. The apparatus 100 according to the present modification further includes a coordinate acquisition section 510 and a region of interest designation section 520 .

[0070] The coordinate acquisition section 510 acquires at least one coordinate in the input image. The coordinate acquisition section 510 provides the at least one acquired coordinate to the region of interest designation section 520. Here, the coordinate acquisition section 510 may acquire only the coordinate of a single point in the input image, or may acquire a set of multiple coordinates in the input image.

[0071] The region of interest designation section 520 designates the region of interest based on at least one coordinate provided from the coordinate acquisition section 510. Here, the region of interest designation section 520 may designate a predetermined range as the region of interest using at least one coordinate as a reference. Instead or in addition, the region of interest designation section 520 may designate a range having a texture similar to that at at least one coordinate as the region of interest. Instead or in addition, the region of interest designation section 520 may designate a region surrounded by a set of multiple coordinates as the region of interest. The region of interest designation section 520 provides information about the designated region of interest to the annotation image acquisition section 120. The annotation image acquisition section 120 retrieves the annotation image based on the region of interest designated by the region of interest designation section 520.

[0072] Figure 6 An example of specifying a region of interest by the device 100 according to the present modification is shown. For example, the device 100 displays an input image in a screen, and receives an input caused by manipulation by a user (for example, but not limited to, a physician, a veterinarian, a laboratory technician, or a student). The coordinate acquisition portion 510 retrieves at least one coordinate in the input image based on the user input, for example, the coordinates at the position indicated by the cross mark shown in the left image of the present figure. The region of interest designation portion 520 then specifies the region of interest as a predetermined range based on the at least one coordinate, for example, an elliptical range of a specified size centered on the coordinates of the position indicated by the cross mark shown in the right image of the present figure (the area indicated in white in the figure).

[0073] However, the method for specifying the region of interest is not limited thereto. As described above, the region of interest designation unit 520 may designate a range of textures having a texture similar to that at the position indicated by the cross mark as the region of interest. As another example, the user may use a mouse operation or the like to perform an operation such as surrounding a partial area in the input image, and the coordinate acquisition unit 510 may acquire a set of multiple coordinates surrounding the area. The region of interest designation unit 520 may then designate the region surrounded by the multiple coordinate sets as the region of interest. In addition, the region of interest designation unit 520 may use a combination of the above-mentioned multiple designation methods. In other words, the region of interest designation unit 520 may designate a range of textures having a texture similar to that of at least one coordinate within a predetermined range using at least one coordinate as a reference as the region of interest. Alternatively, the region of interest designation unit 520 may designate a range including similar textures within a region surrounded by a set of multiple coordinates as the region of interest.

[0074] In this way, the device 100 according to the present modification roughly specifies the region of interest by causing the user to input at least one coordinate in the input image. Then, according to the device 100 according to the present modification, even when an annotation image according to the region of interest roughly specified in this way, i.e., an annotation image containing errors, is acquired, correct label classification is learned from the labels containing errors, and thus the detection model 130 for detecting the target region can be accurately learned.

[0075] Figure 7 An exemplary block diagram of a device 100 according to another modification of this embodiment is shown. Figure 5 Components having the same functions and configurations are given the same reference numerals, and the following description includes only the different points. The apparatus 100 of the present modification further includes a region of interest display section 710 and a modification section 720.

[0076] For example, the region of interest display section 710 displays the region of interest specified by the region of interest specifying section 520 on a monitor or the like.

[0077] The modifying section 720 receives a user input and modifies the region of interest specified by the region of interest specifying section 520 when the region of interest displaying section 710 displays the region of interest.

[0078] In this way, the device 100 according to the present modification has a function of displaying the region of interest roughly designated by the region of interest designation section 520 and allowing the user to modify the region of interest. In this way, according to the device 100 according to the present modification, the modification based on the user's experience can be reflected in the region of interest designated mechanically, and an annotation image with less error can be acquired.

[0079] As an example, the above embodiment can be modified in the following manner. For example, the device 100 includes a plurality of detection models, and retrieves a plurality of annotated images, wherein the region of interest for the same input image is set to have different ranges. The device 100 can then use the plurality of annotated images to learn each of the plurality of detection models, and use the region where the target region output by each detection model overlaps as the target region output by the detection model 130. For example, the device 100 includes two detection models, and for the same input image, retrieves an image in which the region of interest is set to an annotation with a relatively wide range and an image in which the region of interest is set to an annotation with a relatively narrow range. The device 100 then uses the image in which the region of interest is set to an annotation with a relatively wide range to learn one of the detection models, and uses the image in which the region of interest is set to an annotation with a relatively narrow range to learn another detection model. Then, the region predicted as the target region by both of the two detection models can be used as the target region output by the detection model 130.

[0080] As another example, the above embodiment may be modified in the following manner. For example, if the region of interest designation section 520 designates the region of interest based on at least one coordinate, the annotation image acquisition section 120 may set an annotation image corresponding to the at least one coordinate. The above description is an example of a case where the annotation image acquisition section 120 acquires an annotation image, wherein pixels inside the region of interest are labeled with class c=1 and pixels outside the region of interest are labeled with class c=0. In other words, the annotation image acquisition section 120 acquires an annotation image in which all pixels inside the region of interest are labeled with the same class. However, the annotation image acquisition section 120 may acquire an annotation image in which pixels inside the region of interest are labeled with a class corresponding to the positions of these pixels. In other words, as an example, if the region of interest designation section 520 has designated a region of interest centered on at least one coordinate, the annotation image acquisition section 120 may acquire an annotation image in which pixels at at least one coordinate are labeled with class c=1, and the farther these pixels are from the at least one coordinate, the other pixels are labeled with a class closer to 0. In this way, the device 100 can further enhance the effect of pixels near the location specified by the user (even within the region of interest) to learn the detection model.

[0081] Different embodiments of the present invention may be described with reference to flowcharts and block diagrams, the blocks of which may represent: (1) steps of a process for performing an operation, and (2) portions of an apparatus responsible for performing the operation. Certain steps and portions may be implemented by dedicated circuits, programmable circuits provided with computer-readable instructions stored on a computer-readable medium, and / or processors provided with computer-readable instructions stored on a computer-readable medium. Dedicated circuits may include digital and / or analog hardware circuits, and may include integrated circuits (ICs) and / or discrete circuits. Programmable circuits may include reconfigurable hardware circuits, including logical AND, OR, XOR, NAND, NOR and other logical operations, flip-flops, registers, memory elements, etc., such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), etc.

[0082] The present invention may be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon, the computer-readable program instructions being used to cause a processor to perform aspects of the present invention.

[0083] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device (such as a punch card) or a raised structure in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (e.g., a light pulse by a fiber optic cable), or an electrical signal transmitted by a wire.

[0084] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.

[0085] The computer-readable program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages, such as "C" programming language or similar programming languages. The computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer, partially on a remote computer, or completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (for example, by using the Internet of an Internet service provider). In certain embodiments, an electronic circuit (including, for example, a programmable logic circuit, a field programmable gate array (FPGA) or a programmable logic array (PLA)) can be executed by utilizing the state information of the computer-readable program instructions to individualize the electronic circuit, so as to perform aspects of the present invention.

[0086] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each box of the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0087] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine that, when executed by a processor of the computer or other programmable data processing apparatus, creates a device for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer, a programmable data processing apparatus, and / or other device to function in a particular manner, such that a computer-readable storage medium having instructions stored therein includes an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0088] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other apparatus that causes a series of operating steps to be performed on a computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0089] The flow chart and block diagram in the accompanying drawings show the architecture, function and operation of the possible implementation of the system, method and computer program product according to various embodiments of the present invention. To this end, each box in the flow chart or block diagram can represent a part of a module, segment or instruction, which includes one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the function marked in the box may not occur in the order marked in the figure. For example, depending on the function involved, the two boxes shown in succession can actually be executed substantially at the same time, or these boxes can sometimes be executed in the opposite order. It will also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a system based on special-purpose hardware, and the system based on special-purpose hardware performs a specified function or action or performs a combination of special-purpose hardware and computer instructions.

[0090] Although the embodiments of the present invention have been described, the technical scope of the present invention is not limited to the above-described embodiments. It is obvious that those skilled in the art can add various changes and improvements to the above-described embodiments. It is also obvious from the scope of the claims that the embodiments to which such changes or improvements have been added may be included in the technical scope of the present invention.

[0091] The operations, processes, steps and phases of each process performed by the device, system, program and method shown in the claims, embodiments or figures can be performed in any order as long as the order is not indicated by "before", "before", etc. and as long as the output from the previous process is not used in the latter process. Even if a phrase (such as "first" or "next" in the claims, embodiments or schematics) is used to describe the process flow, it does not necessarily mean that the process must be performed in this order.

[0092] Figure 8 An exemplary hardware configuration of a computer configured to perform the above-mentioned operations according to an embodiment of the present invention is shown. The program installed in the computer 700 can make the computer 700 function as or perform operations associated with the apparatus of the embodiment of the present invention or one or more parts thereof (including modules, components, elements, etc.), and / or make the computer 700 perform the process of the embodiment of the present invention or its steps. Such a program can be executed by the CPU 700-12 to make the computer 700 perform certain operations associated with some or all of the blocks of the flowcharts and block diagrams described herein.

[0093] The computer 700 according to the present embodiment includes a CPU 700-12, a RAM 700-14, a graphic controller 700-16, and a display device 700-18 connected to each other through a host controller 700-10. The computer 700 also includes input / output units such as a communication interface 700-22, a hard disk drive 700-24, a DVD drive 700-26, and an IC card drive, which are connected to the host controller 700-10 via an input / output controller 700-20. The computer also includes conventional input / output units such as a ROM 700-30 and a keyboard 700-42, which are connected to the input / output controller 700-20 through an input / output chip 700-40.

[0094] The CPU 700-12 operates according to the programs stored in the ROM 700-30 and the RAM 700-14, thereby controlling each unit. The graphic controller 700-16 obtains image data generated by the CPU 700-12 on a frame buffer or the like provided in the RAM 700-14 or on itself, and displays the image data on the display device 700-18.

[0095] The communication interface 700-22 communicates with other electronic devices via the network 700-50. The hard disk drive 700-24 stores programs and data used by the CPU 700-12 in the computer 700. The DVD drive 700-26 reads programs or data from the DVD-ROM 700-01 and provides the programs or data to the hard disk drive 700-24 via the RAM 700-14. The IC card drive reads programs and data from the IC card and / or writes programs and data to the IC card.

[0096] The ROM 700-30 stores therein a boot program or the like executed by the computer 700 upon activation, and / or a program depending on the hardware of the computer 700. The input / output chip 700-40 may also connect various input / output units to the input / output controller 700-20 via a parallel port, a serial port, a keyboard port, a mouse port, and the like.

[0097] The program is provided by a computer-readable medium such as a DVD-ROM 700-01 or an IC card. The program is read from the computer-readable medium, installed in a hard disk drive 700-24, a RAM 700-14, or a ROM 700-30, which is also an example of a computer-readable medium, and executed by a CPU 700-12. The information processing described in these programs is read into the computer 700, resulting in cooperation between the program and the various types of hardware resources described above. The device or method can be constructed by implementing the operation or processing of information such as a receiving buffer provided on a recording medium according to the use of the computer 700-50.

[0098] For example, when communication is performed between the computer 700 and an external device, the CPU 700-12 may execute a communication program loaded on the RAM 700-14 to instruct the communication interface 700-22 on the communication processing based on the processing described in the communication program. The communication interface 700-22 reads the transmission data stored on the transmission buffer area set in the recording medium such as the RAM 700-14, the hard disk drive 700-24, the DVD-ROM 700-01 or the IC card under the control of the CPU 700-12, and transmits the read transmission data to the network 700-50 or writes the reception data received from the network 700-50 to the reception buffer area set on the recording medium, etc.

[0099] In addition, the CPU 700-12 can cause all or a necessary part of a file or database to be read into the RAM 700-14, which has been stored in an external recording medium such as the hard disk drive 700-24, the DVD-drive 700-26 (DVD-ROM 700-01), an IC card, etc., and perform different types of processing on the data on the RAM 700-14. Then, the CPU 700-12 can write the processed data back to the external recording medium.

[0100] Various types of information (e.g., various types of programs, data, tables, and databases) can be stored in a recording medium to undergo information processing. CPU700-12 can perform various types of processing on data read from RAM700-14, including various types of operations, information processing, conditional judgments, conditional branches, unconditional branches, search / replace information, etc., as described throughout this disclosure and specified by the instruction sequence of the program, and write the results back to RAM700-14. In addition, CPU700-12 can search for information in files, databases, etc. in a recording medium. For example, when multiple entries (each entry having an attribute value of a first attribute) are associated with an attribute value of a second attribute, CPU700-12 can search for entries that match a condition specifying an attribute value of the first attribute, read the attribute value of the second attribute stored in the entry from the multiple entries, thereby obtaining the attribute value of the second attribute associated with the first attribute that meets the predetermined condition.

[0101] The above-mentioned program or software module may be stored in a computer-readable medium on or near the computer 700. In addition, a recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet may be used as a computer-readable medium to provide the program to the computer 700 via the network.

[0102] The present invention may be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon, the computer-readable program instructions being used to cause a processor to perform aspects of the present invention.

[0103] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device (such as a punch card) or a raised structure in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (e.g., a light pulse by a fiber optic cable), or an electrical signal transmitted by a wire.

[0104] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.

[0105] The computer-readable program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages, such as "C" programming language or similar programming languages. The computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer, partially on a remote computer, or completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (for example, by using the Internet of an Internet service provider). In certain embodiments, an electronic circuit (including, for example, a programmable logic circuit, a field programmable gate array (FPGA) or a programmable logic array (PLA)) can be executed by utilizing the state information of the computer-readable program instructions to individualize the electronic circuit, so as to perform aspects of the present invention.

[0106] Aspects of the present invention are described herein with reference to flowcharts of methods, devices (systems) and computer program products according to embodiments of the present invention and / or block diagrams. It should be understood that each box in the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0107] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine that, when executed by a processor of the computer or other programmable data processing apparatus, creates a device for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer, a programmable data processing apparatus, and / or other device to function in a particular manner, such that a computer-readable storage medium having instructions stored therein includes an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0108] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other apparatus that causes a series of operating steps to be performed on a computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0109] The flow chart and block diagram in the accompanying drawings show the architecture, function and operation of the possible implementation of the system, method and computer program product according to various embodiments of the present invention. To this end, each box in the flow chart or block diagram can represent a part of a module, segment or instruction, which includes one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the function marked in the box may not occur in the order marked in the figure. For example, depending on the function involved, the two boxes shown in succession can actually be executed substantially at the same time, or these boxes can sometimes be executed in the opposite order. It will also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a system based on special-purpose hardware, and the system based on special-purpose hardware performs a specified function or action or performs a combination of special-purpose hardware and computer instructions.

[0110] Although the embodiments of the present invention have been described, the technical scope of the present invention is not limited to the above-described embodiments. It is obvious that those skilled in the art can add various changes and improvements to the above-described embodiments. It is also obvious from the scope of the claims that the embodiments to which such changes or improvements have been added may be included in the technical scope of the present invention.

[0111] The operations, processes, steps and phases of each process performed by the device, system, program and method shown in the claims, embodiments or figures can be performed in any order as long as the order is not indicated by "prior to", "before", etc., and as long as the output from the previous process is not used in the next process. Even if a phrase (such as "first" or "next" in the claims, embodiments or schematics) is used to describe the process flow, it does not necessarily mean that the process must be performed in this order.

Claims

1. A computer-implemented method for image analysis, the method comprising: Get the input image; obtaining an annotation image specifying a region of interest in the input image, wherein pixels of the annotation image are labeled with corresponding values ​​of a class associated with the region of interest; Inputting the input image to a detection model, the detection model generating an output image showing a target region from the input image, wherein pixels of the output image are labeled with corresponding values ​​based on a class of the target region; calculating an error between the output image and the annotation image using a loss function, the loss function weighting errors within the region of interest more heavily than errors outside the region of interest, a weighted cross entropy being used as the loss function, the weighted cross entropy weights inside the target region being set to have a larger value than those outside the target region, the loss function comprising multiplying a value of a label of a pixel of the annotation image by a logarithm of a value of a label of a corresponding pixel of the output image, and applying a weight associated with a class; and The detection model is updated in a manner that reduces the error.

2. The computer-implemented method of claim 1, wherein the weighted cross entropy weights inside the region of interest are set to have larger values ​​than the weights outside the region of interest.

3. The computer-implemented method of claim 2, wherein the weighted cross entropy is given by the following expression, where X is the set of all pixels i in the input image, C is the set of all classes c, and W c is the weight of the class c, p i c is the value of the class c at the pixel i in the annotation image, and q i c is the value of the class c at the pixel i in the output image; Expression 1:

4. A computer-implemented method as described in any one of claims 1 to 3, wherein the detection model includes one or more convolutional layers, one or more pooling layers, one or more deconvolution layers, and one or more batch normalization layers between the input and the output.

5. The computer-implemented method of claim 4, wherein the plurality of batch normalization layers are arranged every predetermined number of layers in at least one of a first path and a second path, the first path comprising some of the plurality of convolutional layers and the plurality of pooling layers, and the second path comprising the remaining convolutional layers of the plurality of convolutional layers and the plurality of deconvolutional layers.

6. The computer-implemented method of any one of claims 1 to 3, further comprising: Obtaining at least one coordinate in the input image; as well as The region of interest is specified according to the at least one coordinate.

7. The computer-implemented method of claim 6, wherein specifying the region of interest based on the at least one coordinate comprises: A predetermined range based on the at least one coordinate is designated as the region of interest.

8. The computer-implemented method of claim 6, wherein the designating the region of interest according to the at least one coordinate comprises designating a range having a texture similar to the texture at the at least one coordinate as the region of interest. 9 . The computer-implemented method of claim 6 , wherein the specifying the region of interest according to the at least one coordinate comprises specifying a range surrounded by a set of multiple coordinates as the region of interest.

10. The computer-implemented method of claim 6, further comprising: displaying the designated region of interest; as well as While displaying the region of interest, receiving user input and modifying the region of interest.

11. An apparatus comprising: processor or programmable circuit; as well as One or more computer-readable media that collectively include instructions that, in response to being executed by the processor or the programmable circuit, cause the processor or the programmable circuit to: Get the input image; obtaining an annotation image specifying a region of interest in the input image, wherein pixels of the annotation image are labeled with corresponding values ​​of a class associated with the region of interest; Inputting the input image to a detection model, the detection model generating an output image showing a target region from the input image, wherein pixels of the output image are labeled with corresponding values ​​based on a class of the target region; calculating an error between the output image and the annotation image using a loss function, the loss function weighting errors within the region of interest more heavily than errors outside the region of interest, a weighted cross entropy being used as the loss function, the weighted cross entropy weights inside the target region being set to have a larger value than those outside the target region, the loss function comprising multiplying a value of a label of a pixel of the annotation image by a logarithm of a value of a label of a corresponding pixel of the output image, and applying a weight associated with a class; and The detection model is updated in a manner that reduces the error.

12. The apparatus of claim 11, wherein the detection model comprises one or more convolutional layers, one or more pooling layers, one or more deconvolutional layers, and one or more batch normalization layers between the input and the output.

13. The apparatus of claim 11 or 12, wherein the instructions further cause the processor or the programmable circuit to: obtaining at least one coordinate in the input image; and The region of interest is specified according to the at least one coordinate.

14. A computer program product for image analysis, the computer program product comprising: A computer readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit in order to perform the method of any one of claims 1 to 10.

15. A computer program product, comprising a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, wherein the computer program can be loaded into an internal memory of a digital computer, wherein the computer program comprises a software code portion, and when the computer program is run on the digital computer, the software code portion is used to execute the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • A device and a method for segmenting a low-field strength stomach MRI image

    CN109377497A