Information processing device, information processing method, and recording medium

By employing low-noise label maps and multiple segmentation models with fusion and background noise reduction, the accuracy of image segmentation models is enhanced, addressing the issue of noisy ground truth data and improving detection precision, especially for small objects.

WO2026154643A1PCT designated stage Publication Date: 2026-07-23NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NEC CORP
Filing Date
2025-01-17
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing image segmentation models struggle with low estimation accuracy due to noisy ground truth data, particularly when detecting small objects, leading to misclassification and background detection as targets.

Method used

The use of low-noise label maps with three or more levels of grayscale to train segmentation models, reducing noise by assigning pixel values between belonging and not belonging to the target, and employing multiple segmentation models with fusion and background noise reduction techniques to generate high-quality low-noise label maps.

Benefits of technology

Improves the estimation accuracy of segmentation models, especially for small objects, by minimizing noise and enhancing the precision of pixel classification, thereby reducing misclassification and improving detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025001415_23072026_PF_FP_ABST
    Figure JP2025001415_23072026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device comprises: a label generation means for generating a low-noise label map regarding a detection target included in an input image; a first segmentation means for using a segmentation-related first model to estimate pixels in the input image that correspond to a portion of the detection target, and generate a first label map including pixels having pixel values corresponding to a likelihood indicating how likely these pixels belong to the detection target; and a training means for training the first model using the low-noise label map and the first label map.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and recording medium

[0001] This disclosure relates to the technical fields of information processing equipment, information processing methods, and recording media.

[0002] One such device has been proposed that segments a digital image into target objects within the digital image (see Patent Document 1).

[0003] Special Publication No. 2013-524361

[0004] This disclosure aims to provide an information processing device, an information processing method, and a recording medium that improve upon the technologies described in prior art documents.

[0005] One embodiment of an information processing device includes: a label generation means for generating a low-noise label map relating to a detection target included in an input image; a first segmentation means for estimating pixels in the input image that correspond to a part of the detection target and generating a first label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the detection target, using a first segmentation model; and a learning means for learning the first model using the low-noise label map and the first label map.

[0006] One aspect of the information processing method involves generating a low-noise label map relating to the object to be detected in the input image, estimating pixels in the input image that correspond to a part of the object to be detected using a first segmentation model, generating a first label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the object to be detected, and training the first model using the low-noise label map and the first label map.

[0007] One embodiment of a recording medium contains a computer program that causes a computer to execute an information processing method that generates a low-noise label map relating to a detection target included in an input image, estimates pixels in the input image that correspond to a part of the detection target using a first segmentation model, generates a first label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the detection target, and trains the first model using the low-noise label map and the first label map.

[0008] Another aspect of the information processing device includes: a first segmentation means that uses a first model relating to segmentation to estimate pixels corresponding to any part of any of the multiple detection targets in an input image containing a plurality of detection targets, and generates an estimated label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to any of the plurality of detection targets; and a learning means that uses the estimated label map and the ground truth label map to train the first model, wherein the estimated label map includes a plurality of first regions corresponding to each of the plurality of detection targets, each of the plurality of first regions is separated from the other first regions, the ground truth label map includes a plurality of second regions corresponding to each of the plurality of detection targets, and the learning means calculates the loss between one of the plurality of first regions and one of the plurality of second regions corresponding to the one first region, and trains the first model based on the plurality of losses corresponding to each of the plurality of first regions.

[0009] This is a conceptual diagram showing an example of an information processing device according to the embodiment. This is a conceptual diagram showing the operation of an example of an information processing device according to the embodiment. This is a flowchart showing the operation of an example of an information processing device according to the embodiment. This is a block diagram showing an example of a computer configuration. This is a diagram showing an example of the configuration of a low-noise label generation means according to the embodiment. This is a diagram for explaining an example of the operation of a background noise improvement means according to the embodiment. This is a diagram for explaining another example of the operation of a background noise improvement means according to the embodiment. This is a diagram showing another example of the configuration of a low-noise label generation an example of the configuration of a learning means according to the embodiment. This is a diagram showing another example of the configuration of a learning means according to the embodiment. This is a conceptual diagram showing another example of an information processing device according to the embodiment. This is a conceptual diagram showing the operation of another example of an information processing device according to the embodiment. This is a conceptual diagram showing another example of an information processing device according to the embodiment.

[0010] Embodiments of an information processing device, an information processing method, and a recording medium will be described based on the drawings.

[0011] <First Embodiment> An information processing device, information processing method, and recording medium according to the first embodiment will be described with reference to Figures 1 to 3. In the following, the information processing device, information processing method, and recording medium according to the first embodiment will be described using the information processing device 10.

[0012] Figure 1 is a conceptual diagram showing the concept of the information processing device 10. The information processing device 10 learns a model for detecting objects from an image using image segmentation technology. In Figure 1, the information processing device 10 includes a low-noise label generation means 11, a segmentation means 12, and a learning means 13.

[0013] The segmentation means 12 estimates pixels in the input image that correspond to a part of the target to be detected, using a first model related to segmentation. Here, the first model corresponds to a model learned by the information processing device 10. The first model may be an untrained model or a trained model. The first model may have a neural network model structure.

[0014] For example, the segmentation means 12 (specifically, the first model) may determine the likelihood that each pixel in the input image corresponds to a part of the target to be detected. In other words, the segmentation means 12 (specifically, the first model) may determine the likelihood that each pixel belongs to the target to be detected. For example, the likelihood may be expressed as a value in the range of 0 to 1.

[0015] The segmentation means 12 generates a first label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the detection target. In the first label map, a collection of pixels having pixel values ​​greater than or equal to a predetermined value (in other words, a region consisting of pixels having pixel values ​​greater than or equal to a predetermined value) corresponds to the detection target detected by the segmentation means 12.

[0016] The low-noise label generation means 11 generates a low-noise label map relating to the detection targets included in the input image. The low-noise label map is used as ground truth data in training the first model described above. The low-noise label map is a label map that includes pixels having a first pixel value indicating that it belongs to the detection target, pixels having a second pixel value indicating that it does not belong to the detection target, and pixels having a pixel value between the first pixel value and the second pixel value. In other words, the low-noise label map is a label map with three or more levels of grayscale. Here, the pixel value between the first pixel value and the second pixel value corresponds to the class likelihood. That is, the pixel value of each pixel in the low-noise label map may indicate the likelihood that the pixel belongs to the detection target. The first pixel value may be "1", and the second pixel value may be "0".

[0017] In digital images, the pixel value of a pixel corresponding to a part of an object's edge is affected by both light from the object (e.g., light reflected by the object, light scattered by the object, and light emitted by the object itself) and light from the background of the object. In other words, in digital images, the boundary between the object and the background becomes ambiguous. On the other hand, in segmentation processing, each pixel is determined to be either a pixel corresponding to a part of an object (in other words, belonging to the object) or a pixel corresponding to something other than the object (in other words, not belonging to the object). Therefore, a pixel corresponding to a part of an object's edge may be determined to belong to the object or not during segmentation processing.

[0018] If a binary image is used as the ground truth data, pixels near the boundary between the object (i.e., the object to be detected) and the background in the ground truth data will be assigned either a pixel value indicating that it belongs to the object (corresponding to the first pixel value above) or a pixel value indicating that it does not belong to the object (corresponding to the second pixel value above). In this case, the pixel value can be rephrased as a label. Therefore, assigning a pixel value to a pixel can be called labeling.

[0019] For example, when a human creates a binary image as ground truth data, the following can occur: If a human is unsure whether a pixel belongs to the object to be detected, they may mistakenly classify that pixel as part of the background (i.e., not belonging to the object to be detected). In other words, a classification error (i.e., a labeling error) can occur. Alternatively, as mentioned above, digital images contain pixels where light from an object (i.e., the object to be detected) and light from the background overlap. For such pixels, a binary image must assign either a pixel value indicating that it belongs to the object or a pixel value indicating that it does not belong to the object. That is, for pixels that should be assigned a pixel value between a pixel value indicating that it belongs to the object and a pixel value indicating that it does not belong to the object, either a pixel value indicating that it belongs to the object or a pixel value indicating that it does not belong to the object must be assigned.

[0020] As a result, the boundary between the object and the background in the binary image, which serves as the ground truth data, may deviate from its original state. In supervised learning, the model is trained so that the data output by the model approaches the ground truth data. Therefore, the discrepancy between the state shown by the ground truth data and the original state can become a disruptive factor in the learning process and may affect the learning of the first model described above. This discrepancy between the state shown by the ground truth data and the original state can be referred to as "noise."

[0021] As mentioned above, a low-noise label map is a label map with three or more levels of grayscale. In a low-noise label map, pixel values ​​are assigned to the boundary region where the signal information of an object (i.e., the object being detected) and the signal information of the background are mixed. These values ​​fall between pixel values ​​indicating that the object belongs to the object and pixel values ​​indicating that the object does not belong to the object. Therefore, the discrepancy between the state shown by the low-noise label map and the original state is smaller than the discrepancy between the state shown by the binary image and the original state. Thus, a low-noise label map, which is a label map with three or more levels of grayscale, can be said to be low-noise.

[0022] The learning means 13 trains the first model using the low-noise label map and the first label map. For example, the learning means 13 may use a loss function to calculate the difference (i.e., loss) between the low-noise label map as ground truth data and the first label map as the output of the first model during the training of the first model. The learning means 13 may change the parameters related to the first model during the training of the first model so that this difference becomes smaller. Examples of loss functions include the cross-entropy loss function, the IoU (Intersection over Union) loss function, the Dice loss function, the L1 loss, the L2 loss, etc.

[0023] Next, the operation of the information processing device 10 will be explained with reference to Figures 2 and 3. Figure 2 is a conceptual diagram showing the operation of the information processing device 10. Figure 3 is a flowchart showing the operation of the information processing device 10. As shown in Figure 2, the input image Img1 as training data (or learning data) is input to the low-noise label generation means 11 and the segmentation means 12.

[0024] In Figure 3, the low-noise label generation means 11, which receives the input image Img1, generates a low-noise label map LL (see Figure 2) relating to the detection target contained in the input image Img1 (step S101). In parallel with the processing in step S101, the segmentation means 12, which receives the input image Img1, performs segmentation processing on the input image Img1 using a first model related to segmentation (step S102). The segmentation means 12 outputs a first label map EL1 (see Figure 2) based on the results of the segmentation processing (step S103). The learning means 13 trains the first model using the low-noise label map LL as ground truth data and the first label map EL1 (step S104).

[0025] Furthermore, the input image (for example, input image Img1) may be an infrared image. For example, by using an infrared image as the input image, detection of the target can be performed regardless of whether it is day or night. Also, since infrared images can avoid the influence of relatively strong light such as headlights, relatively detailed image information can be obtained using infrared images. However, the input image may also be a visible light image. Furthermore, the input image may be the image itself acquired from a camera, etc. (for example, an image that has not been preprocessed), or it may be an image acquired from a camera, etc. that has been preprocessed. As an example of preprocessing, adaptive histogram equalization (CLAHE) can be mentioned.

[0026] Thus, the information processing device 10 may perform an information processing method in which it generates a low-noise label map relating to the detection target included in the input image, estimates pixels corresponding to a part of the detection target in the input image using a first segmentation model, generates a first label map including pixels having pixel values ​​corresponding to the likelihood of belonging to the detection target, and trains the first model using the low-noise label map and the first label map.

[0027] The information processing device 10 described above may be realized by a computer reading a computer program recorded on a recording medium. In this case, the recording medium may contain a computer program that causes the computer to execute an information processing method that generates a low-noise label map relating to the detection target contained in the input image, estimates pixels corresponding to a part of the detection target in the input image using a first segmentation model, generates a first label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the detection target, and learns the first model using the low-noise label map and the first label map.

[0028] (Technical Effects) According to the inventor's research, the following has been found: When a segmentation model is trained using manually labeled binary images as ground truth data, it is difficult to sufficiently improve the estimation accuracy of the model. As a result, for example, detection of the target may fail, or the background may be mistakenly detected as the target. In contrast, in this embodiment, the first segmentation model is trained using a low-noise label map as ground truth data. As described above, a low-noise label map has relatively little noise (i.e., deviation from the original state). Therefore, the information processing device 10 according to the first embodiment can improve the estimation accuracy of the first model.

[0029] Incidentally, assuming the size of an object remains constant, the longer the distance between the camera (in other words, the imaging device) and the object, the fewer pixels make up the object in the digital image (in other words, the smaller the size of the object in the digital image). When the number of pixels making up an object is relatively small, the impact of deviation from the original state becomes greater compared to when the number of pixels making up the object is relatively large. For example, for an object made up of 1000 pixels, if the deviation from the original state is 1 pixel, the deviation ratio to the number of pixels is 0.1%. In contrast, for an object made up of 10 pixels, if the deviation from the original state is 1 pixel, the deviation ratio to the number of pixels is 10%. Therefore, when training a segmentation model (e.g., the first model) using a noisy label map as ground truth data, the estimation accuracy of the segmentation model may decrease as the number of pixels making up the target of detection decreases.

[0030] For the reasons described above, when the number of pixels constituting the target to be detected is relatively small, using a low-noise label map as ground truth data can significantly improve the estimation accuracy of the first model related to segmentation. In this case, the target to be detected may be a minute object composed of a predetermined number of pixels or less in the input image (for example, input image Img1).

[0031] <Second Embodiment> The information processing apparatus, information processing method, and recording medium according to the second embodiment will be described with reference to Figure 4 in addition to Figures 1 to 3. In the following, the information processing apparatus, information processing method, and recording medium according to the second embodiment will be described using the information processing apparatus 10. Note that in the second embodiment, explanations that overlap with the first embodiment described above will be omitted as appropriate.

[0032] The first embodiment described above explains the concept of the information processing device 10. The second embodiment will describe the information processing device 10 in detail. For example, the information processing device 10 may be implemented by the computer COM shown in Figure 4. Figure 4 is a block diagram showing an example of the configuration of a computer.

[0033] The computer COM includes an arithmetic unit 110, a storage device 120, a communication device 130, an input device 140, and an output device 150. The arithmetic unit 110, the storage device 120, the communication device 130, the input device 140, and the output device 150 are connected via a data bus 160. Note that the computer COM may not include at least one of the input device 140 and the output device 150.

[0034] The arithmetic unit 110 may have a processor. Note that the arithmetic unit 110 may have a single processor or may have a plurality of processors. That is, the arithmetic unit may have one or more processors. Note that the processor may be a multi-core processor. When the arithmetic unit 110 has a single processor that is a multi-core processor, the arithmetic unit 110 can be logically said to have a plurality of processors.

[0035] The processor may be, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), a TPU (Tensor Processing Unit), and a quantum processor.

[0036] The storage device 120 may be, for example, at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and an optical disk array. That is, the storage device 120 may be realized by a single device or may be realized by a plurality of devices.

[0037] The communication device 130 may be able to communicate with a device external to the computer COM. Note that the communication device 130 may perform wired communication or may perform wireless communication.

[0038] The input device 140 is a device capable of receiving input of information to the computer COM from the outside. The input device 140 may include an operating device (for example, a keyboard, a mouse, a touch panel, etc.) that can be operated by the user of the computer COM. The input device 140 may include a recording medium reading device capable of reading information recorded on a recording medium detachable from the computer COM, such as a USB (Universal Serial Bus) memory. In addition, when information is input to the computer COM via the communication device 130 (in other words, when the computer COM acquires information via the communication device 130), the communication device 130 may function as an input device.

[0039] The output device 150 is a device capable of outputting information to the outside of the computer COM. The output device 150 may have a display device capable of outputting visual information such as characters and images as the above information. The output device 150 may have a speaker capable of outputting auditory information such as sound as the above information. The output device 150 may have a vibration motor capable of outputting tactile information such as vibration as the above information. The output device 150 may have a printer. The output device 150 may be capable of outputting information to a recording medium detachable from the computer COM, such as a USB memory. In addition, when the computer COM outputs information via the communication device 130, the communication device 130 may function as an output device.

[0040] The storage device 120 is capable of storing desired data. The storage device 120 may store the computer program CP executed by the arithmetic device 110. The storage device 120 may temporarily store data temporarily used by the arithmetic device 110 when the arithmetic device 110 is executing the computer program CP.

[0041] Furthermore, the computer program CP may be recorded on a computer-readable and non-temporary recording medium (i.e., a recording medium different from the storage device 120). In this case, the computer program CP may be stored in the storage device 120 by reading the recording medium using a recording medium reading device (not shown) provided by the computer COM. Furthermore, at least one of the following may be used as the recording medium: an optical disc, a magnetic medium, a magneto-optical disc, a semiconductor memory, and any other medium capable of storing a program. Furthermore, the computer program CP may be obtained from an external device (not shown) of the computer COM via the communication device 130. In other words, the computer program CP may be downloaded from an external device to the storage device 120 of the computer COM.

[0042] The arithmetic unit 110 (for example, a processor) may execute the processing that the computer COM should perform together with the storage device 120 in which the computer program CP is stored (in other words, together with the storage device 120 and the computer program CP stored in the storage device 120). For example, by the arithmetic unit 110 executing the computer program CP, a logical functional block for executing the processing that the computer COM should perform may be realized within the arithmetic unit 110 (for example, within the processor).

[0043] Here, the computer program CP may include a computer program that causes the computer to execute an information processing method that generates a low-noise label map relating to the objects to be detected in the input image, estimates pixels in the input image that correspond to a part of the objects to be detected using a first segmentation model, generates a first label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the objects to be detected, and trains the first model using the low-noise label map and the first label map.

[0044] For example, the low-noise label generation means 11, the segmentation means 12, and the learning means 13 may be implemented within the arithmetic unit 110 by the execution of a computer program CP by the arithmetic unit 110. In this case, the low-noise label generation means 11, the segmentation means 12, and the learning means 13 may be implemented as the logical functional blocks described above. At least one of the low-noise label generation means 11, the segmentation means 12, and the learning means 13 may be implemented as a physical processing circuit (i.e., hardware). At least one of the low-noise label generation means 11, the segmentation means 12, and the learning means 13 may be implemented in a form in which logical functional blocks and physical processing circuits coexist.

[0045] Furthermore, when the low-noise label generation means 11, segmentation means 12, and learning means 13 are implemented as logical functional blocks, the low-noise label generation means 11, segmentation means 12, and learning means 13 may be implemented by a single processor. The low-noise label generation means 11, segmentation means 12, and learning means 13 may each be implemented by different processors. Parts of the low-noise label generation means 11, segmentation means 12, and learning means 13 may be implemented by one processor, while the remaining parts of the low-noise label generation means 11, segmentation means 12, and learning means 13 may be implemented by one or more other processors different from that one processor.

[0046] Furthermore, the low-noise label generation means 11, the segmentation means 12, and the learning means 13 may each be implemented by different computers. Alternatively, parts of the low-noise label generation means 11, the segmentation means 12, and the learning means 13 may be implemented by one computer, while the remaining parts of the low-noise label generation means 11, the segmentation means 12, and the learning means 13 may be implemented by one or more other computers different from that one computer.

[0047] (Technical Effects) According to the second embodiment, the information processing device 10 can be realized relatively easily.

[0048] <Third Embodiment> The information processing apparatus, information processing method, and recording medium according to the third embodiment will be described with reference to Figures 5 to 7. Hereinafter, the information processing apparatus, information processing method, and recording medium according to the third embodiment will be described using the information processing apparatus 10. In the third embodiment, explanations that overlap with the first and second embodiments described above will be omitted as appropriate, and common parts in the drawings will be denoted by the same reference numerals.

[0049] In the third embodiment, a specific example of the low-noise label generation means 11 will be described. Figure 5 is a diagram showing an example of the configuration of the low-noise label generation means 11. In Figure 5, the low-noise label generation means 11 includes a plurality of segmentation means 211, 212, ..., 21n, a fusion means 220, and a background noise improvement means 230.

[0050] Each of the multiple segmentation means 211, 212, ..., 21n uses a second model related to segmentation to estimate pixels in the input image that correspond to a part of the target to be detected. In the third embodiment, the multiple second models used by each of the multiple segmentation means 211, 212, ..., 21n are different models from each other. For example, the multiple second models may have the same model structure but different parameters. For example, the multiple second models may have different model structures. The multiple second models are different models from the first model. Each of the multiple second models is a trained model.

[0051] For example, the segmentation means 211 (specifically, the second model) may determine the likelihood that each pixel in the input image corresponds to a part of the target to be detected. In other words, the segmentation means 211 (specifically, the second model) may determine the likelihood that each pixel belongs to the target to be detected. For example, the likelihood may be expressed as a value in the range of 0 to 1.

[0052] The segmentation means 211 generates a second label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the detection target. In the second label map, a collection of pixels having pixel values ​​greater than or equal to a predetermined value (in other words, a region consisting of pixels having pixel values ​​greater than or equal to a predetermined value) corresponds to the detection target detected by the segmentation means 211.

[0053] Multiple segmentation means 212, ..., 21n also generate second label maps in the same way as segmentation means 211. As described above, the multiple second models are different from each other. For this reason, the pixel value of one pixel in one of the multiple second label maps generated by each of the multiple segmentation means 211, 212, ..., 21n may differ from the pixel value of the pixel corresponding to the above one pixel in another second label map. The fusion means 220 fuses the multiple second label maps generated by each of the multiple segmentation means 211, 212, ..., 21n to generate a fused label map. The fusion means 220 may determine the pixel value of one pixel in the fused label map as follows. For example, the fusion means 220 may obtain the pixel value of the pixel corresponding to one pixel in the fused label map from each of the multiple second label maps. The fusion means 220 may use the average value of the multiple obtained pixel values ​​as the pixel value of one pixel in the fused label map. In this case, the fusion means 220 may generate a fused label map that includes a pixel whose pixel value is the average of the pixel values ​​calculated based on a plurality of second label maps. Alternatively, the fusion means 220 may use the weighted average of the acquired plurality of pixel values ​​as the pixel value of one pixel in the fused label map. In this case, the fusion means 220 may generate a fused label map that includes a pixel whose pixel value is the weighted average of the pixel values ​​calculated based on a plurality of second label maps.

[0054] The weights of the above weighted average may be weights corresponding to the accuracy of the multiple second models used by each of the multiple segmentation means 212, ..., 21n. In this case, the accuracy of each second model may be determined according to the IoU loss. Alternatively, the accuracy of each second model may be determined according to the average accuracy of the maps (corresponding to the second label map) output from each second model. The fused label map generated as described above is a label map that includes pixels having a first pixel value, pixels having a second pixel value, and pixels having a pixel value between the first and second pixel values. In other words, the fused label map is a label map with three or more grayscale levels.

[0055] The background noise reduction means 230 acquires a manual label map ML. The manual label map ML is a label map that includes pixels having a first pixel value (i.e., pixels corresponding to a part of the target to be detected) and pixels having a second pixel value (i.e., pixels corresponding to areas other than the target to be detected), which are generated by a human determining pixels corresponding to a part of the target to be detected in the input image (e.g., input image Img1). In other words, the manual label map ML is a binary image of the first pixel value and the second pixel value.

[0056] The background noise improvement means 230 generates a third label map by changing the pixel values ​​of at least some pixels around a region consisting of multiple pixels having a first pixel value (corresponding to the white region in the manual label map ML shown in Figure 5) from the second pixel value to the first pixel value in the manual label map ML. The third label map is a map that includes pixels having a first pixel value and pixels having a second pixel value. The white region in the manual label map ML shown in Figure 5 (i.e., the region consisting of multiple pixels having a first pixel value) will henceforth be referred to as the "foreground region" as appropriate.

[0057] A specific example of the operation of the background noise reduction means 230 will be explained with reference to Figure 6 in addition to Figure 5. Figure 6 is a diagram illustrating an example of the operation of the background noise reduction means 230. In Figure 6, the color of the pixels having the second pixel value (black in the manual label map ML shown in Figure 5) is different from that in Figure 5 in order to make the leader lines visible. The same applies to Figure 7, which will be described later.

[0058] For example, the background noise improvement means 230 may define a bounding box BB (see Figure 6) that surrounds the foreground region of the manual label map ML (corresponding to region FA in Figure 6). The background noise improvement means 230 may generate a third label map by changing the pixel values ​​of region FC, which is the part of the region surrounded by the bounding box BB other than the foreground region (i.e., region FA), from second pixel values ​​to first pixel values.

[0059] Another specific example of the operation of the background noise improvement means 230 will be described with reference to Figure 7 in addition to Figure 5. Figure 7 is a diagram illustrating another example of the operation of the background noise improvement means 230. For example, the background noise improvement means 230 may perform an expansion process on the foreground region of the manual label map ML (corresponding to region FA in Figure 7). As a result, the pixel values ​​of pixels surrounding the foreground region of the manual label map ML (i.e., region FA) (region FC in Figure 7) are changed from second pixel values ​​to first pixel values, and a third label map is generated.

[0060] The background noise reduction means 230 generates a low-noise label map (for example, a low-noise label map LL) by calculating the product of the pixel value of one pixel in the fused label map and the pixel value of the corresponding pixel in the third label map. The second pixel value may be "0". In this case, the product of the pixel having the second pixel value in the third label map (in other words, the pixel corresponding to the background) and the pixel value of the fused label map becomes "0". Here, the reliability of the background region other than the foreground region in the manually created label map ML is sufficiently high. For this reason, the reliability of the background region of the third label map based on the manually created label map ML is also sufficiently high. As described above, by calculating the product of the pixel value of the fused label map and the pixel value of the third label map, the third label map can be prioritized over the fused label map for the background region.

[0061] (Technical Effects) In the third embodiment, a fused label map is generated by fusing multiple second label maps generated by multiple segmentation means 211, 212, ..., 21n, respectively. In the fused label map, it is expected that the pixel values ​​of pixels that are likely to correspond to a part of the detection target will be relatively large, and the pixel values ​​of pixels that are unlikely to correspond to a part of the detection target will be relatively small. However, in the fused label map, the pixel values ​​of pixels that should be the background (i.e., parts that are not the detection target) may be relatively high. In other words, there may be errors in the fused label map. In the third embodiment, the background noise improvement means 230 generates a low-noise label map using the fused label map and the third label map. Therefore, even if there are errors in the fused label map, it is possible to suppress the reflection of those errors in the low-noise label map. Accordingly, according to the third embodiment, a high-quality low-noise label map can be generated.

[0062] <Fourth Embodiment> The information processing apparatus, information processing method, and recording medium according to the fourth embodiment will be described with reference to Figure 8. Hereinafter, the information processing apparatus, information processing method, and recording medium according to the fourth embodiment will be described using the information processing apparatus 10. In the fourth embodiment, a specific example of the low-noise label generation means 11 will be described, similar to the third embodiment described above. For the fourth embodiment, explanations that overlap with the first to third embodiments described above will be omitted as appropriate, and common parts in the drawings will be denoted by the same reference numerals.

[0063] Figure 8 shows another example of the configuration of the low-noise label generation means 11. In Figure 8, the low-noise label generation means 11 includes a plurality of segmentation means 211, 212, ..., 21n, a fusion means 220, a background noise improvement means 230, and an expansion means 240.

[0064] Each of the multiple segmentation means 211, 212, ..., 21n uses a second model related to segmentation to estimate pixels in the input image that correspond to a part of the target to be detected. In the fourth embodiment, the multiple second models used by each of the multiple segmentation means 211, 212, ..., 21n may be the same model or they may be different models. Here, the multiple second models being the same means that the model structure is the same and the parameters are the same. Each of the multiple second models is a trained model.

[0065] The extension means 240 generates multiple processed images, each artificially modified from the input image, by applying predetermined image processing to the input image. Examples of predetermined image processing include noise processing to change the amount of image noise, Gaussian blur processing, and projection transformation processing. The number of processed images generated by the extension means 240 may be the same as the number of the multiple segmentation means 211, 212, ..., 21n. Each of the multiple segmentation means 211, 212, ..., 21n is input with a different processed image.

[0066] For example, the segmentation means 211 (specifically, the second model) may determine the likelihood that each pixel in the input processed image corresponds to a part of the detection target. In other words, the segmentation means 211 (specifically, the second model) may determine the likelihood that each pixel belongs to the detection target. For example, the likelihood may be expressed as a value in the range of 0 to 1.

[0067] The segmentation means 211 generates a second label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the detection target. In the second label map, a collection of pixels having pixel values ​​greater than or equal to a predetermined value (in other words, a region consisting of pixels having pixel values ​​greater than or equal to a predetermined value) corresponds to the detection target detected by the segmentation means 211.

[0068] The multiple segmentation means 212, ..., 21n also generate second label maps in the same way as segmentation means 211. As described above, each of the multiple segmentation means 211, 212, ..., 21n is input with a different processed image. For this reason, the pixel value of one pixel in one of the multiple second label maps generated by each of the multiple segmentation means 211, 212, ..., 21n may differ from the pixel value of the corresponding pixel in another second label map.

[0069] The fusion means 220 fuses the multiple second label maps generated by the multiple segmentation means 211, 212, ..., 21n to generate a fused label map. The fusion means 220 may determine the pixel value of one pixel in the fused label map as follows: For example, the fusion means 220 may obtain the pixel value of a pixel corresponding to one pixel in the fused label map from each of the multiple second label maps. The fusion means 220 may use the average value of the multiple obtained pixel values ​​as the pixel value of one pixel in the fused label map. In this case, the fusion means 220 may generate a fused label map that includes a pixel whose pixel value is the average value of the pixel values ​​calculated based on the multiple second label maps. Alternatively, the fusion means 220 may use the weighted average value of the multiple obtained pixel values ​​as the pixel value of one pixel in the fused label map. In this case, the fusion means 220 may generate a fused label map that includes a pixel whose pixel value is the weighted average value of the pixel values ​​calculated based on the multiple second label maps. The weights of the weighted mean may be weights corresponding to the intensity parameters related to predetermined image processing when the processed image is generated. Examples of intensity parameters include the standard deviation of the Gaussian blur, the standard deviation of the noise, the rotation angle associated with the projection transformation, the translation distance, etc. For example, the larger the intensity parameter, the smaller the weight of the second label generated by the segmentation means that receives the processed image generated using the intensity parameter. For example, if the intensity parameter is "x" and the weight is "w", the fusion means 220 may determine the weight based on the formula "w = exp(-x)".

[0070] The operation of the background noise reduction means 230 is the same as that of the background noise reduction means 230 according to the third embodiment described above. Therefore, a description of the background noise reduction means 230 will be omitted.

[0071] (Technical Effects) According to the fourth embodiment, a high-quality, low-noise label map can be generated, similar to the third embodiment described above.

[0072] <Fifth Embodiment> The information processing apparatus, information processing method, and recording medium according to the fifth embodiment will be described with reference to Figure 9. Hereinafter, the information processing apparatus, information processing method, and recording medium according to the fifth embodiment will be described using the information processing apparatus 10. In the fifth embodiment, a specific example of the low-noise label generation means 11 will be described, similar to the third and fourth embodiments described above. For the fifth embodiment, explanations that overlap with the first to fourth embodiments described above will be omitted as appropriate, and common parts in the drawings will be denoted by the same reference numerals.

[0073] Figure 9 shows another example of the configuration of the low-noise label generation means 11. In Figure 9, the low-noise label generation means 11 includes a foreground candidate region extraction means 250 and a label determination means 260.

[0074] As shown in Figure 9, a manual label map (e.g., manual label map ML) is input to the foreground candidate region extraction means 250. The foreground candidate region extraction means 250 extracts foreground candidate regions from the manual label map. As described in the third embodiment, the "foreground region" means a region in the manual label map consisting of pixels having a first pixel value (i.e., pixels corresponding to a part of the detection target). For example, region FA in Figure 6 corresponds to an example of a "foreground region". The "foreground candidate region" means at least a part of the periphery of the foreground region in the manual label map. As described in the third embodiment, the manual label map (e.g., manual label map ML) is a label map generated by a human determining pixels in the input image (e.g., input image Img1) that correspond to a part of the detection target. Furthermore, in digital images, it can be difficult to clearly determine the boundary between an object (i.e., the detection target) and the background. For this reason, some of the pixels corresponding to the background in the manual label map may be pixels corresponding to a part of the detection target. An example of the "foreground candidate region" described above is a region consisting of pixels that, among the pixels corresponding to the background in a manual label map, may be some of the pixels to be detected.

[0075] The foreground candidate region extraction means 250 may extract foreground candidate regions as follows. For example, the foreground candidate region extraction means 250 may define a bounding box surrounding the foreground region of the manual label map. The foreground candidate region extraction means 250 may extract the portion of the region enclosed by the bounding box that is not the foreground region as a foreground candidate region. For example, region FC shown in Figure 6 corresponds to an example of a foreground candidate region.

[0076] For example, the foreground candidate region extraction means 250 may perform an expansion process on the foreground region of the manually labeled map. The foreground candidate region extraction means 250 may extract the difference between the foreground region expanded by the expansion process and the original foreground region as a foreground candidate region. For example, region FC shown in Figure 7 corresponds to another example of a foreground candidate region. The foreground candidate region extraction means 250 may perform the expansion process multiple times.

[0077] The label determination means 260 generates a low-noise label map by assigning a pixel value between a first pixel value and a second pixel value to each pixel in the foreground candidate region. In other words, the label determination means 260 generates a low-noise label map by changing the pixel value of each pixel in the foreground candidate region to a pixel value between a first pixel value and a second pixel value. Here, the pixel value between the first pixel value and the second pixel value corresponds to the class likelihood. That is, the pixel value between the first pixel value and the second pixel value may indicate the likelihood that the pixel belongs to the foreground (i.e., the detected object).

[0078] For example, the label determination means 260 may determine the pixel value of each pixel in the foreground candidate region such that the pixel value of pixels close to the foreground region in the manual label map approaches the first pixel value (in other words, the pixel value of pixels far from the foreground region approaches the second pixel value). For example, if the foreground region in the manual label map has been expanded once, the label determination means 260 may assign the average value of the first pixel value and the second pixel value to each pixel in the foreground candidate region.

[0079] For example, if the foreground region in a manual label map is subjected to two expansion processes, the label determination means 260 may assign a pixel value to each pixel in the portion increased by the first expansion process that is 2 / 3 of the sum of the first pixel value and the second pixel value. The label determination means 260 may assign a pixel value to each pixel in the portion increased by the second expansion process that is 1 / 3 of the sum of the first pixel value and the second pixel value.

[0080] For example, the label determination means 260 may determine the pixel value of each pixel in the foreground candidate region such that the pixel value of a pixel close to the centroid of the foreground region in the manual label map approaches the first pixel value (in other words, the pixel value of a pixel farther from the centroid of the foreground region approaches the second pixel value).

[0081] For example, the label determination means 260 may determine the pixel value of each pixel in the foreground candidate region based on the pixel value (e.g., brightness value) of the input image pixel corresponding to each pixel in the foreground candidate region. Specifically, the label determination means 260 may determine the foreground brightness value as the average value of the brightness values ​​of multiple pixels in the input image that correspond to each of the multiple pixels constituting the foreground candidate region. The label determination means 260 may determine the background brightness value as the average value of the brightness values ​​of n pixels surrounding each of the multiple pixels in the input image that correspond to each of the multiple pixels constituting the foreground candidate region. The label determination means 260 may change the brightness value of each of the multiple pixels in the input image that correspond to each of the multiple pixels constituting the foreground candidate region so that the foreground brightness value becomes a first pixel value (e.g., "1") and the background brightness value becomes a second pixel value (e.g., "0"). The label determination means 260 may use the changed brightness value as the pixel value of each pixel in the foreground candidate region.

[0082] (Technical Effects) In the third and fourth embodiments described above, a low-noise label map is generated using a pre-trained second model related to segmentation. In contrast, in the fifth embodiment, a low-noise label map is generated without using a machine learning model. Therefore, according to the fifth embodiment, the low-noise label generation means 11 can be realized at a lower cost compared to the third and fourth embodiments.

[0083] <Sixth Embodiment> The information processing apparatus, information processing method, and recording medium according to the sixth embodiment will be described with reference to Figure 10. In the following, the information processing apparatus, information processing method, and recording medium according to the sixth embodiment will be described using the information processing apparatus 10. In the sixth embodiment, explanations that overlap with the first to fifth embodiments described above will be omitted as appropriate, and common parts in the drawings will be denoted by the same reference numerals.

[0084] In the sixth embodiment, a specific example of the learning means 13 will be described. Figure 10 is a diagram showing an example of the configuration of the learning means 13. In Figure 10, the learning means 13 includes a loss function calculation means 310, a fusion weight calculation means 320, a loss function calculation means 330, a loss fusion means 340, and a parameter update means 350.

[0085] The low-noise label generation means 11 according to the sixth embodiment generates a low-noise label map (for example, low-noise label map LL) and outputs a confidence score indicating the degree of confidence related to the low-noise label map. Here, the low-noise label generation means 11 according to the sixth embodiment may have the configuration described in any of the third to fifth embodiments as the configuration for generating the low-noise label map.

[0086] The low-noise label generation means 11 may determine the confidence score as follows. For example, the low-noise label generation means 11 may have the configuration described in the third or fourth embodiment. In this case, the low-noise label generation means may obtain the pixel value of one image corresponding to each of the multiple second label maps based on the multiple second label maps generated by the multiple segmentation means 211, 212, ..., 21n. The low-noise label generation means may determine the variance of the pixel value of the one pixel. The low-noise label generation means may calculate the average value of the variance of all pixel values ​​as the uncertainty of the low-noise label map. The low-noise label generation means 11 may use the value obtained by subtracting the above uncertainty from 1 as the confidence score. Alternatively, the low-noise label generation means 11 may determine the entropy of the pixel value instead of the variance of the pixel value. In this case, the low-noise label generation means may calculate the average value of the entropy of all pixel values ​​as the uncertainty of the low-noise label map. Furthermore, the low-noise label generation means 11 may determine the confidence score by estimating a scalar for the low-noise label map. Alternatively, the low-noise label generation means 11 may determine the confidence score by estimating a map score of the same size as the low-noise label map.

[0087] The loss function calculation means 310 receives a low-noise label map (for example, low-noise label map LL) and a first label map (for example, first label map EL1) as input. The loss function calculation means 310 uses the loss function to calculate the difference (i.e., loss) between the low-noise label map and the first label map.

[0088] The loss function calculation means 330 receives a manual label map (for example, manual label map ML) and a first label map (for example, first label map EL1) as input. The loss function calculation means 310 uses a loss function to calculate the difference (i.e., loss) between the manual label map and the first label map. The loss function used by the loss function calculation means 330 may be the same as the loss function used by the loss function calculation means 310. However, the loss function used by the loss function calculation means 330 may be different from the loss function used by the loss function calculation means 310.

[0089] The fusion weight calculation means 320 receives a confidence score as input. Based on the confidence score, the fusion weight calculation means 320 calculates fusion weights for fusing the loss calculated by the loss function calculation means 310 with the loss calculated by the loss function calculation means 330. The larger the confidence score, the larger the weight of the loss calculated by the loss function calculation means 310 may be compared to the weight of the loss calculated by the loss function calculation means 330. In other words, the smaller the confidence score, the smaller the weight of the loss calculated by the loss function calculation means 310 may be compared to the weight of the loss calculated by the loss function calculation means 330. The fusion weights may be values ​​between 0 and 1.

[0090] The loss fusion means 340 calculates the fusion loss by fusing the loss calculated by the loss function calculation means 310 and the loss calculated by the loss function calculation means 330 using the fusion weights calculated by the fusion weight calculation means 320. Here, the fusion loss may be a weighted average of the loss calculated by the loss function calculation means 310 and the loss calculated by the loss function calculation means 330.

[0091] The parameter update means 350 updates (or modifies) the parameters related to the first model based on the fusion loss as part of the learning of the first model. For example, the parameter update means 350 may update (or modify) the parameters related to the first model so that the fusion loss is reduced.

[0092] (Technical Effects) The certainty of each low-noise label map and manual label map as ground truth data is relative. The certainty of the low-noise label map is not always greater than the certainty of the manual label map. Therefore, the learning means 13 according to the sixth embodiment fuses the loss calculated by the loss function calculation means 310 and the loss calculated by the loss function calculation means 330 using a fused weight based on the confidence score output from the low-noise label generation means 11. Then, the learning means 13 learns the first model based on the fused loss (i.e., the fused loss). By configuring it in this way, the first model can be learned based on the label map that is more likely among the low-noise label map and the manual label map. As a result, the information processing device 10 according to the sixth embodiment can improve the estimation accuracy of the first model.

[0093] <Seventh Embodiment> The information processing device, information processing method, and recording medium according to the seventh embodiment will be described with reference to Figure 11. Hereinafter, the information processing device, information processing method, and recording medium according to the seventh embodiment will be described using the information processing device 10. In the seventh embodiment, a specific example of the learning means 13 will be described, similar to the sixth embodiment described above. For the seventh embodiment, explanations that overlap with the first to sixth embodiments described above will be omitted as appropriate, and common parts in the drawings will be denoted by the same reference numerals.

[0094] Figure 11 shows another example of the configuration of the learning means 13. In Figure 11, the learning means 13 includes a loss function calculation means 310, a loss function calculation means 330, a loss fusion means 340, and a parameter update means 350.

[0095] The low-noise label generation means 11 according to the seventh embodiment includes a segmentation means 210 and a background noise improvement means 230. The segmentation means 210 uses a second model related to segmentation to estimate pixels in the input image that correspond to a part of the target to be detected. In the seventh embodiment, the second model may be an untrained model or a trained model. The operation of the background noise improvement means 230 is the same as the operation of the background noise improvement means 230 according to the third embodiment described above. Therefore, a description of the background noise improvement means 230 is omitted.

[0096] The segmentation means 210 (specifically, the second model) may determine the likelihood that each pixel in the input image corresponds to a part of the target to be detected. In other words, the segmentation means 210 (specifically, the second model) may determine the likelihood that each pixel belongs to the target to be detected. For example, the likelihood may be expressed as a value in the range of 0 to 1. The segmentation means 210 may generate a fourth label map that includes pixels having pixel values ​​corresponding to the likelihood that they belong to the target to be detected. Here, the fourth label map is a label map with three or more grayscale levels.

[0097] The background noise improvement means may generate a low-noise label map (e.g., low-noise label map LL) based on the fourth label map and a manual label map (e.g., manual label map ML).

[0098] The information processing device 10 according to the seventh embodiment learns a first model (i.e., the model used by the segmentation means 12) and a second model (i.e., the model used by segmentation 210). Here, the model structure of the first model and the model structure of the second model may be the same.

[0099] The loss function calculation means 310 receives a low-noise label map (for example, low-noise label map LL) and a first label map (for example, first label map EL1) as input. The loss function calculation means 310 uses the loss function to calculate the difference (i.e., loss) between the low-noise label map and the first label map.

[0100] The loss function calculation means 330 receives a manual label map (for example, manual label map ML) and a first label map (for example, first label map EL1) as input. The loss function calculation means 310 uses the loss function to calculate the difference (i.e., loss) between the manual label map and the first label map.

[0101] The loss fusion means 340 calculates a fusion loss by fusing the loss calculated by the loss function calculation means 310 with the loss calculated by the loss function calculation means 330. Here, the fusion loss may be the average value of the loss calculated by the loss function calculation means 310 and the loss calculated by the loss function calculation means 330.

[0102] The parameter update means 350 updates (or modifies) the parameters of the first model based on the fusion loss as part of learning the first model. For example, the parameter update means 350 may update (or modify) the parameters of the first model so that the fusion loss is reduced. The parameter update means 350 further updates (or modifies) the parameters of the second model based on the updated parameters of the first model.

[0103] For example, each time an input image is input to the segmentation means 12, the loss fusion means 340 calculates the fusion loss. Then, the parameter update means 350 updates the parameters related to the first model based on the fusion loss. As a result, the parameters related to the first model change over time. In other words, the parameters related to the first model can be said to be time-series data. For example, the parameter update means 350 may update the parameters related to the second model based on the moving average value of the parameters related to the first model, which are time-series data. For example, the parameter update means 350 may update the parameters related to the second model based on the exponentially smoothed moving average value of the parameters related to the first model, which are time-series data.

[0104] As described above, updating the parameters of the second model based on the moving average or exponentially smoothed moving average of the parameters of the first model produces the following effect. As a premise, averaging the inference results of multiple models reduces noise compared to the inference result of a single model. The configuration of the low-noise label generation means 11 according to the third and fourth embodiments takes this point into consideration. By updating the parameters of the second model based on the moving average or exponentially smoothed moving average of the parameters of the first model (i.e., a single model), the inference result of the second model approaches the average of the inference result of the first model at the first time point, the inference result of the first model at the second time point, ..., the inference result of the first model at the nth time point. In other words, the inference result of the second model approaches the average of n different inference results of the first model. Therefore, the inference result of the second model has less noise than the inference result of the first model. Accordingly, the low-noise label generation means 11 according to the seventh embodiment can generate a low-noise label map (for example, a low-noise label map LL).

[0105] Furthermore, the fusion loss may be a weighted average of the loss calculated by the loss function calculation means 310 and the loss calculated by the loss function calculation means 330. In this case, in the initial stages of training the first model, the loss fusion means 340 may make the weight of the loss calculated by the loss function calculation means 330 greater than the weight of the loss calculated by the loss function calculation means 310. In the middle and later stages of training the first model, the loss fusion means 340 may make the weight of the loss calculated by the loss function calculation means 310 greater than the weight of the loss calculated by the loss function calculation means 330. Furthermore, instead of the low-noise label generation means 11 shown in Figure 11, the low-noise label generation means shown in Figure 8 (in other words, the one according to the fourth embodiment) may be used.

[0106] (Technical Effects) In the configurations according to the third and fourth embodiments, it is necessary to pre-train the second model (i.e., the second model used by each of the multiple segmentation means 211, 212, ..., 21n). However, training the second model takes a considerable amount of time. For this reason, in the configurations according to the third and fourth embodiments, the period from when the training of the second model begins until the training of the first model is completed is relatively long. In contrast, in the configuration according to the seventh embodiment, the training of the first model and the training of the second model can be performed in parallel. For this reason, the information processing device 10 according to the seventh embodiment can shorten the training period of the first model.

[0107] <Eighth Embodiment> The information processing apparatus, information processing method, and recording medium according to the eighth embodiment will be described with reference to Figures 12 and 13. Hereinafter, the information processing apparatus, information processing method, and recording medium according to the eighth embodiment will be described using the information processing apparatus 20. In the eighth embodiment, explanations that overlap with the first to seventh embodiments described above will be omitted as appropriate, and common parts in the drawings will be denoted by the same reference numerals.

[0108] Figure 12 is a conceptual diagram showing the concept of the information processing device 20. In Figure 12, the information processing device 20 comprises a foreground region extraction means 21, a segmentation means 12, and a learning means 13. The information processing device 20, like the information processing device 10, also learns the first model (i.e., the model related to segmentation) used by the segmentation means 12.

[0109] The information processing device 20 will be described with reference to Figure 13 in addition to Figure 12. Figure 13 is a conceptual diagram showing the operation of the information processing device 20. In the eighth embodiment, an input image Img2 (see Figure 13) containing multiple detection targets is input to the segmentation means 12. The input image Img2 shown in Figure 13 includes detection targets Ob1, Ob2, and Ob3.

[0110] The segmentation means 12 uses a first model relating to segmentation to estimate pixels in the input image Img2 that correspond to a portion of any of the multiple detection targets Ob1, Ob2, and Ob3. The segmentation means 12 may generate an estimated label map EL2 that includes pixels having a first pixel value indicating that they belong to any of the multiple detection targets Ob1, Ob2, and Ob3, and pixels having a second pixel value indicating that they do not belong to any of the multiple detection targets Ob1, Ob2, and Ob3. For example, the first pixel value may be "1" and the second pixel value may be "0". In addition, the estimated label map EL2 may include pixels having a pixel value between the first and second pixel values, as well as pixels having the first and second pixel values. In other words, the estimated label map EL2 may be a label map with three or more levels of grayscale. In this case, the segmentation means 12 (specifically, the first model) may determine the likelihood that each pixel in the input image Img2 belongs to one of the multiple detection targets Ob1, Ob2, and Ob3. The segmentation means 12 may generate an estimated label map EL2, which includes pixels having pixel values ​​corresponding to the likelihood that the pixel belongs to one of the multiple detection targets Ob1, Ob2, and Ob3.

[0111] For example, in the estimated label map EL2, a collection of pixels having a first pixel value (in other words, a region consisting of pixels having a first pixel value) corresponds to a detection target detected by the segmentation means 12. The estimated label map EL2 includes a plurality of first regions, each corresponding to a plurality of detection targets Ob1, Ob2, and Ob3. Each of these plurality of first regions may be a region consisting of pixels having a first pixel value. Furthermore, each of the plurality of first regions is separated from the other first regions.

[0112] The foreground region extraction means 21 receives a ground truth label map as input. The ground truth label map includes multiple second regions, each corresponding to one of the multiple detection targets Ob1, Ob2, and Ob3. Each of the above multiple second regions is separated from the other second regions. The ground truth label map may be a low-noise label map or a manual label map. If the ground truth label map is a low-noise label map, the information processing device 20 may be equipped with means equivalent to the low-noise level generation means 11 described above.

[0113] The foreground region extraction means 21 extracts foreground regions from the ground truth label map. For example, the foreground region extraction means 21 may use image segmentation techniques to determine whether each pixel in the ground truth label map is a foreground region. For example, if the ground truth label map is a low-noise label map, the foreground region extraction means 21 may determine that pixels with a pixel value greater than or equal to a threshold are foreground regions. For example, if the ground truth label map is a manual label map, the foreground region extraction means 21 may determine that pixels having a first pixel value are foreground regions. If the foreground region extraction means 21 determines that both one pixel and other pixels adjacent to that pixel are foreground regions, it may extract the one pixel and the other pixels as a single foreground region. The foreground region extraction means 21 may also assign identification numbers to multiple foreground regions that are separated from each other to identify each foreground region. Note that the multiple foreground regions correspond to the multiple second regions described above.

[0114] The foreground region extraction means 21 may set a bounding box surrounding each foreground region to which an identification number has been assigned. The foreground region extraction means 21 may obtain the center coordinates of the bounding box, as well as the width and height of the bounding box. Here, the width and height of the bounding box may be expressed in terms of the number of pixels.

[0115] Alternatively, the foreground region extraction means 21 may extract foreground regions assigned identification numbers from the ground truth label map and generate multiple label maps, each containing a single foreground region. For example, the foreground region extraction means 21 may extract a foreground region corresponding to the detection target Ob1 from the ground truth label map and generate a label map containing only the foreground region corresponding to the detection target Ob1. The foreground region extraction means 21 may extract a foreground region corresponding to the detection target Ob2 from the ground truth label map and generate a label map containing only the foreground region corresponding to the detection target Ob2. The foreground region extraction means 21 may extract a foreground region corresponding to the detection target Ob3 from the ground truth label map and generate a label map containing only the foreground region corresponding to the detection target Ob3.

[0116] The learning means 13 includes a parameter update means 350 and a region loss calculation means 360. The region loss calculation means 360 calculates the loss between one of the multiple first regions included in the estimated label map EL2 and the foreground region relating to one of the multiple second regions included in the ground truth label map that corresponds to the first region. In other words, the region loss calculation means 360 calculates the loss for each region. As a result, the region loss calculation means 360 calculates multiple losses corresponding to each of the multiple first regions.

[0117] If the foreground region extraction means 21 sets the bounding box described above, the region loss calculation means 360 may calculate the loss as follows. First, the region loss calculation means 360 may extract rectangular regions containing each first region from the estimated label map EL2. For example, the region loss calculation means 360 may extract rectangular regions containing the first region by setting a bounding box that surrounds the first region. Next, the region loss calculation means 360 may compare the rectangular region containing one first region with the bounding box that surrounds the foreground region to which the identification number corresponding to the first region is assigned. For example, the region loss calculation means 360 may calculate the loss related to one first region based on the center coordinates, width, and height of the rectangular region containing one first region and the center coordinates, width, and height of the bounding box.

[0118] When the foreground region extraction means 21 generates the above-mentioned plurality of label maps, the region loss calculation means 360 may calculate the loss as follows: The region loss calculation means 21 may calculate the loss related to the first region by comparing the first region included in the estimated label map EL2 with a label map that includes only the foreground regions to which the identification number corresponding to the first region has been assigned.

[0119] Furthermore, the region loss calculation means 360 may calculate the above loss using a cross-entropy loss function, an IoU loss function, a Dice loss function, an L1 loss, or an L2 loss. Furthermore, the region loss calculation means 360 may calculate the above loss as the difference between the centroid of one first region and the centroid of the foreground region relating to one second region corresponding to the first region.

[0120] The parameter update means 350 may update (or change) the parameters related to the first model based on a plurality of losses calculated by the domain loss calculation means 360 during the training of the first model. For example, the parameter update means 350 may update (or change) the parameters related to the first model so that the sum or average value of the plurality of losses becomes smaller.

[0121] Alternatively, the region loss calculation means 360 may calculate the loss based on the following formula. In the following formula, "L" represents the loss, and "w" represents the loss. i The "hat" represents weight, and "A P " represents the region consisting of pixels that have the first pixel value in the estimated label map (for example, estimated label map EL2) (in other words, the region consisting of pixels that belong to the detection target), and "A GT " represents the region consisting of pixels that have the first pixel value in the ground truth label map (in other words, the region consisting of pixels that belong to the detection target). Note that "w i In the "hat" symbol, "i" is a variable. The range of "i" changes depending on the number of detection targets included in the input image. For example, in the case of input image Img2, "i = 1, 2, 3".

[0122]

[0123] The above “w iThe "hat" is determined based on the following formula. In the following formula, "num target " represents the number of objects to be detected in the input image. For example, in the case of input image Img2, "num target = 3".

[0124]

[0125] In this case, the parameter update means 350 may update (or change) the parameters related to the first model based on the loss L calculated by the domain loss calculation means 360 during the training of the first model. For example, the parameter update means 350 may update (or change) the parameters related to the first model so that the loss L is reduced.

[0126] (Technical Effects) When an input image contains multiple detection targets, the loss calculated by comparing the entire estimated label map with the entire ground truth label map is susceptible to the influence of detection accuracy for detection targets with a relatively large number of pixels in the input image (i.e., detection targets that are relatively large in the input image). As a result, it becomes difficult to improve the detection accuracy for detection targets with a relatively small number of pixels in the input image (i.e., detection targets that are relatively small in the input image). In contrast, in the eighth embodiment, the region loss calculation means 360 calculates the loss for each of the multiple foreground regions corresponding to each of the multiple detection targets included in the input image (for example, input image Img2). The parameter update means 350 updates the parameters related to the first model based on the loss calculated by the region loss calculation means 360. This makes it possible to learn to improve detection accuracy regardless of the size of the detection target. As a result, according to the information processing device 20 of the eighth embodiment, it is possible to improve the detection accuracy for detection targets with a relatively small number of pixels in the input image in particular.

[0127] <9th Embodiment>An information processing apparatus, an information processing method, and a recording medium according to the 9th embodiment will be described. Hereinafter, the information processing apparatus, the information processing method, and the recording medium according to the 9th embodiment will be described using the information processing apparatus 10. In the 9th embodiment, a specific example of the loss calculation method in the learning means 13 according to the above-described 1st to 7th embodiments will be described. Regarding the 9th embodiment, descriptions overlapping with the above-described 1st to 7th embodiments will be omitted as appropriate.

[0128] The learning means 13 according to the 9th embodiment may generate a weight map having the same size as the correct label map (for example, a low-noise label map) (that is, the same number of pixels as the correct label map), and the pixel value of each pixel is "a (0 ≤ a ≤ 1)". Here, "a" corresponds to the initial value of the weight. The learning means 13 adds "γ" calculated based on the following formula to the pixel of the weight map corresponding to the pixel belonging to the detection target in the correct label map. i Alternatively, the learning means 13 adds "γ" calculated based on the following formula to the pixel of the weight map corresponding to the pixel within the bounding box surrounding the detection target in the correct label map. i In the following formula, "d" represents the distance from the center of gravity of the detection target to pixel i, and "A" represents the area (for example, the number of pixels) of the detection target. As is clear from the following formula, the value of "γ" increases as the distance from the center of gravity of the detection target increases. i i

[0129]

[0130] The learning means 13 may calculate the loss for each pixel using the first label map generated by the segmentation means 12 and the correct label map. The learning means 13 may calculate the loss of the entire first label map using the loss calculated for each pixel and the above weight map. For example, the learning means 13 may calculate the weighted average of the losses with the pixel value (for example, "a + γ") of each pixel of the above weight map as the weight. i

[0131] (Technical Effects) When loss is calculated in units of the region corresponding to the object to be detected, the loss may be relatively small even if the displacement near the edge of the object to be detected is relatively large. As a result, the detection accuracy of pixels corresponding to the vicinity of the edge of the object to be detected may be relatively low. In contrast, in the ninth embodiment, the weight of pixels relatively far from the centroid of the object to be detected (for example, pixels corresponding to part of the edge of the object to be detected) becomes larger than the weight of pixels relatively close to the centroid of the object to be detected. Therefore, in the ninth embodiment, the displacement near the edge of the object to be detected is emphasized. As a result, according to the ninth embodiment, the detection accuracy of pixels corresponding to the vicinity of the edge of the object to be detected can be improved.

[0132] <Tenth Embodiment> The information processing apparatus according to the tenth embodiment will be described with reference to Figure 14. In the following, the information processing apparatus according to the tenth embodiment will be described using the information processing apparatus 30. Figure 14 is a conceptual diagram showing the concept of the information processing apparatus 30.

[0133] In Figure 14, the information processing device 30 includes a segmentation means 31 and a learning means 32. The segmentation means 31 estimates pixels in the input image that correspond to a part of the target to be detected, using a third model related to segmentation. Here, the third model is a model learned by the information processing device 30. The third model may be an unlearned model or a learned model. The third model may have a neural network model structure.

[0134] For example, the segmentation means 31 (specifically, the third model) may determine the likelihood that each pixel in the input image corresponds to a part of the target to be detected. In other words, the segmentation means 31 (specifically, the third model) may determine the likelihood that each pixel belongs to the target to be detected. For example, the likelihood may be expressed as a value in the range of 0 to 1. The segmentation means 31 generates an estimated label map having pixel values ​​corresponding to the likelihood that each pixel belongs to the target to be detected. The estimated label map is a map with three or more levels of grayscale.

[0135] The learning means 32 trains the third model using the estimated label map and the ground truth label map. The ground truth label map according to the 10th embodiment is a label map that includes pixels having a first pixel value indicating that it belongs to the detection target and pixels having a second pixel value indicating that it does not belong to the detection target. In other words, the ground truth label map may be a binary image of the first pixel value and the second pixel value. For example, the ground truth label map may be a map corresponding to the manual label map ML described above.

[0136] For example, the learning means 32 may calculate the loss related to the estimated label map based on the following formula. In the following formula, “L focal " represents a loss, and "A IoU This represents the IoU score (IoU loss) calculated using the estimated label map and the ground truth label map.

[0137]

[0138] In the above formula, "γ" is determined based on the following formula. In the following formula, "B contrast " represents the contrast score of the input image, and "C Area The symbol " represents the area score of the object to be detected. The contrast score of the input image may be a value calculated based on the variance or entropy of the luminance values ​​of the input image. The area score may be the number of pixels belonging to the object to be detected.

[0139]

[0140] Furthermore, “γ” is not limited to the above formula, but can also be expressed as “γ = 1 / (1 + exp(B) contrast )) / (1+exp(C Area ))" or "γ = exp(-B contrast ) × exp(-C Area It may be determined based on the formula ")".

[0141] Alternatively, the learning means 32 may generate a weight map of the same size as the ground truth label map (i.e., the same number of pixels as the ground truth label map), where the pixel value of each pixel is "a (0 ≤ a ≤ 1)". "a" corresponds to the initial value of the weight. The learning means 32 may add the aforementioned "γ" to the pixels in the weight map that correspond to the pixels belonging to the detection target in the ground truth label map. Alternatively, the learning means 32 may add the aforementioned "γ" to the pixels in the weight map that correspond to the pixels within the bounding box surrounding the detection target in the ground truth label map.

[0142] The learning means 32 may calculate a loss for each pixel using the estimated label map generated by the segmentation means 31 and the ground truth label map. This loss may be a BCE (Binary Cross-Entropy) loss. The learning means 32 may calculate the loss for the entire estimated label map using the loss calculated for each pixel and the weight map. For example, the learning means 32 may calculate a weighted average of the losses using the pixel value of each pixel in the weight map (e.g., "a + γ") as the weight. The learning means 32 may also calculate an IoU loss or a Dice loss instead of a BCE loss.

[0143] As mentioned above, "γ" increases as the contrast score of the input image decreases (i.e., the contrast of the input image is low). Also, "γ" increases as the area of ​​the detection target decreases (i.e., the number of pixels belonging to the detection target decreases). For this reason, when the contrast of the input image is low, and / or when the size of the detection target in the input image is small, the loss calculated by the learning means 32 tends to be large. Therefore, by learning the third model based on the loss calculated as described above, it is possible to improve the detection accuracy of detection targets from low-contrast input images, and / or the detection accuracy of relatively small detection targets in the input image. In other words, the information processing device 30 according to the 10th embodiment can improve the estimation accuracy of the third model.

[0144] <Note> Some or all of the above embodiments may also be described as follows, but are not limited to the following.

[0145] (Note 1) An information processing device comprising: label generation means for generating a low-noise label map relating to a detection target included in an input image; first segmentation means for estimating pixels in the input image that correspond to a part of the detection target using a first segmentation model, and generating a first label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the detection means; and learning means for learning the first model using the low-noise label map and the first label map.

[0146] (Note 2) The information processing apparatus described in Note 1, wherein the object to be detected is a minute object composed of a predetermined number of pixels or less in the input image.

[0147] (Note 3) The information processing apparatus according to Note 1 or 2, wherein the learning means performs, as learning of the first model, calculation of the loss between the low-noise label map and the first label map using a loss function, and updating the parameters related to the first model based on the calculated loss.

[0148] (Note 4) The information processing apparatus according to any one of Notes 1 to 3, wherein the label generation means comprises a plurality of second segmentation means that each use a second model relating to segmentation to estimate pixels in the input image that correspond to a part of the object to be detected, and generate a second label map that includes pixels having a pixel value corresponding to the likelihood of belonging to the object to be detected; and a fusion means that fuses the plurality of second label maps generated by the plurality of second segmentation means to generate a fused label map.

[0149] (Note 5) The information processing apparatus according to any one of Notes 1 to 3, wherein the label generation means comprises: a processed image generation means that generates a plurality of processed images in which the input image has been artificially modified by applying predetermined image processing to the input image; a plurality of second segmentation means that each use a second model relating to segmentation to estimate pixels in corresponding processed images among the plurality of processed images that correspond to a part of the object to be detected, and generate a second label map that includes pixels having a pixel value corresponding to the likelihood of belonging to the object to be detected; and a fusion means that fuses the plurality of second label maps generated by the plurality of second segmentation means to generate a fused label map.

[0150] (Note 6) The information processing apparatus according to any one of claims 1 to 3, wherein the label generation means obtains a manual label map which includes pixels having a first pixel value indicating that they belong to the detection target and pixels having a second pixel value indicating that they do not belong to the detection target, generated by a human determining pixels in the input image that correspond to a part of the detection target, and in the manual label map, the pixel values ​​of at least some pixels around a region consisting of a plurality of pixels having the first pixel value are changed to a value between the first pixel value and the second pixel value.

[0151] (Note 7) The information processing apparatus according to any one of Notes 1 to 6, wherein the label generation means generates the low-noise label map and outputs a confidence score indicating the degree of confidence regarding the low-noise label map, the learning means calculates a first loss of the low-noise label map and the first label map using a loss function, calculates a second loss of the first label map and the manual label map using the loss function, calculates a fusion weight for fusing the first loss and the second loss based on the confidence score, fusing the first loss and the second loss using the fusion weight to calculate a fusion loss, and updates the parameters of the first model based on the fusion loss as learning of the first model, and the manual label map is a label map generated by a human determining pixels corresponding to a part of the object to be detected in the input image, and includes pixels having a first pixel value indicating that they belong to the object to be detected and pixels having a second pixel value indicating that they do not belong to the object to be detected.

[0152] (Note 8) The information processing apparatus according to any one of Notes 1 to 3, wherein the label generation means generates the low-noise label map using a second model relating to segmentation, the learning means calculates a first loss between the low-noise label map and the first label map using a loss function, calculates a second loss between the first label map and the manual label map using the loss function, updates the parameters relating to the first model based on a fused loss obtained by fusing the first loss and the second loss as learning of the first model, updates the parameters relating to the second model based on the updated parameters relating to the first model, and the manual label map is a label map generated by a human determining pixels corresponding to a part of the object to be detected in the input image, and includes pixels having a first pixel value indicating that they belong to the object to be detected and pixels having a second pixel value indicating that they do not belong to the object to be detected.

[0153] (Note 9) The information processing apparatus according to Note 4, wherein the fusion means generates an image as the fused label map, which includes pixels having a weighted average value of pixel values ​​calculated based on the plurality of second label maps.

[0154] (Note 10) The information processing apparatus according to Note 4, wherein the label generation means obtains a manual label map which includes pixels having a first pixel value indicating that they belong to the detection target and pixels having a second pixel value indicating that they do not belong to the detection target, which are generated by a human determining pixels in the input image that correspond to a part of the detection target; generates a third label map by changing the pixel values ​​of at least some pixels around a region consisting of a plurality of pixels having the first pixel value from the second pixel value to the first pixel value in the manual label map; and generates the low-noise label map by calculating the product of the pixel value of one pixel in the fused label map and the pixel value of the pixel in the third label map corresponding to the one pixel.

[0155] (Note 11) The information processing apparatus according to Note 5, wherein the fusion means generates an image as the fused label map, which includes pixels having a weighted average of pixel values ​​calculated based on the plurality of second label maps.

[0156] (Note 12) The information processing apparatus according to Note 5, wherein the label generation means obtains a manual label map which includes pixels having a first pixel value indicating that they belong to the detection target and pixels having a second pixel value indicating that they do not belong to the detection target, generated by a human determining pixels in the input image that correspond to a part of the detection target; generates a third label map by changing the pixel values ​​of at least some pixels around a region consisting of a plurality of pixels having the first pixel value from the second pixel value to the first pixel value in the manual label map; and generates the low-noise label map by calculating the product of the pixel value of one pixel in the fused label map and the pixel value of the pixel in the third label map corresponding to the one pixel.

[0157] (Note 13) The weights relating to the weighted average are set according to the accuracy of each of the plurality of second models, as described in Note 9.

[0158] (Note 14) The weights relating to the weighted average value are set according to the intensity relating to the predetermined image processing, as described in Note 11.

[0159] (Note 15) An information processing method that generates a low-noise label map relating to the object to be detected in the input image, estimates pixels in the input image that correspond to a part of the object to be detected using a first model relating to segmentation, generates a first label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the object to be detected, and trains the first model using the low-noise label map and the first label map.

[0160] (Note 16) A recording medium on which a computer program is recorded that causes a computer to execute an information processing method that generates a low-noise label map relating to a detection target included in an input image, estimates pixels in the input image that correspond to a part of the detection target using a first segmentation model, generates a first label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the detection target, and learns the first model using the low-noise label map and the first label map.

[0161] (Note 17) Information processing device comprising: a first segmentation means that uses a first segmentation model to estimate pixels corresponding to any part of any of the multiple detection targets in an input image containing a plurality of detection targets, and generates an estimated label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to any of the plurality of detection targets; and a learning means that uses the estimated label map and the ground truth label map to train the first model, wherein the estimated label map includes a plurality of first regions corresponding to each of the plurality of detection targets, each of the plurality of first regions is separated from the other first regions, the ground truth label map includes a plurality of second regions corresponding to each of the plurality of detection targets, and the learning means calculates the loss between one of the plurality of first regions and one of the plurality of second regions corresponding to the one first region, and trains the first model based on the plurality of losses corresponding to each of the plurality of first regions.

[0162] Furthermore, some or all of the configurations described in Appendices 2 to 14, which are subordinate to Appendice 1 above, may also be subordinate to Appendices 15 and 16, respectively, in the same manner as the subordinate relationships described in Appendices 2 to 14. Moreover, not limited to Appendices 1, 15, and 16, some or all of the configurations described as appendices may also be subordinate to various hardware, software, various recording means for recording software, or systems, without departing from the embodiments described above.

[0163] This disclosure is not limited to the embodiments described above and may be modified as appropriate, provided that it does not contradict the gist or idea of ​​the invention as can be inferred from the claims and the specification as a whole. Information processing devices, information processing methods, and recording media that involve such modifications are also included in the technical scope of this disclosure.

[0164] 10, 20, 30 Information processing device 11 Low-noise label generation means 12, 31, 210, 211, 212, 21n Segmentation means 13, 32 Learning means

Claims

1. An information processing device comprising: label generation means for generating a low-noise label map relating to a detection target included in an input image; first segmentation means for estimating pixels in the input image that correspond to a part of the detection target and generating a first label map that includes pixels having a pixel value corresponding to the likelihood of belonging to the detection target using a first segmentation model; and learning means for learning the first model using the low-noise label map and the first label map.

2. The information processing apparatus according to claim 1, wherein the object to be detected is a minute object composed of a predetermined number of pixels or less in the input image.

3. The information processing apparatus according to claim 1, wherein the learning means performs, as learning of the first model, calculation of the loss between the low-noise label map and the first label map using a loss function, and updating the parameters related to the first model based on the calculated loss.

4. The information processing apparatus according to claim 1, wherein the label generation means comprises: a plurality of second segmentation means each using a second model relating to segmentation to estimate pixels in the input image that correspond to a part of the object to be detected, and generate a second label map that includes pixels having a pixel value corresponding to the likelihood of belonging to the object to be detected; and a fusion means that fuses the plurality of second label maps generated by the plurality of second segmentation means to generate a fused label map.

5. The information processing apparatus according to claim 1, wherein the label generation means comprises: a processed image generation means that generates a plurality of processed images in which the input image has been artificially modified by applying predetermined image processing to the input image; a plurality of second segmentation means that each use a second model relating to segmentation to estimate pixels in corresponding processed images among the plurality of processed images that correspond to a part of the detection target, and generate a second label map that includes pixels having a pixel value corresponding to the likelihood of belonging to the detection target; and a fusion means that fuses the plurality of second label maps generated by the plurality of second segmentation means to generate a fused label map.

6. The information processing apparatus according to claim 1, wherein the label generation means obtains a manual label map which includes pixels having a first pixel value indicating that they belong to the detection target and pixels having a second pixel value indicating that they do not belong to the detection target, generated by a human determining pixels in the input image that correspond to a part of the detection target, and in the manual label map, the pixel values ​​of at least some pixels around a region consisting of a plurality of pixels having the first pixel value are changed to a value between the first pixel value and the second pixel value.

7. The information processing apparatus according to claim 1, wherein the label generation means generates the low-noise label map and outputs a confidence score indicating the degree of confidence regarding the low-noise label map; the learning means calculates a first loss of the low-noise label map and the first label map using a loss function; calculates a second loss of the first label map and the manual label map using the loss function; calculates a fusion weight for fusing the first loss and the second loss based on the confidence score; fusing the first loss and the second loss using the fusion weight to calculate a fusion loss; and updates the parameters of the first model based on the fusion loss as learning of the first model; and the manual label map is a label map generated by a human determining pixels corresponding to a part of the object to be detected in the input image, and includes pixels having a first pixel value indicating that they belong to the object to be detected and pixels having a second pixel value indicating that they do not belong to the object to be detected.

8. The information processing apparatus according to claim 1, wherein the label generation means generates the low-noise label map using a second model relating to segmentation; the learning means calculates a first loss between the low-noise label map and the first label map using a loss function; calculates a second loss between the first label map and the manual label map using the loss function; updates the parameters relating to the first model based on a fused loss obtained by fusing the first loss and the second loss as learning of the first model; updates the parameters relating to the second model based on the updated parameters relating to the first model; and the manual label map is a label map generated by a human determining pixels corresponding to a part of the object to be detected in the input image, and includes pixels having a first pixel value indicating that they belong to the object to be detected and pixels having a second pixel value indicating that they do not belong to the object to be detected.

9. The information processing apparatus according to claim 4, wherein the fusion means generates a map as the fused label map, which includes pixels having a weighted average value of pixel values ​​calculated based on the plurality of second label maps.

10. The information processing apparatus according to claim 4, wherein the label generation means obtains a manual label map which includes pixels having a first pixel value indicating that they belong to the detection target and pixels having a second pixel value indicating that they do not belong to the detection target, generated by a human determining pixels in the input image that correspond to a part of the detection target; generates a third label map by changing the pixel values ​​of at least some pixels around a region consisting of a plurality of pixels having the first pixel value in the manual label map from the second pixel value to the first pixel value; and generates the low-noise label map by calculating the product of the pixel value of one pixel in the fused label map and the pixel value of the pixel in the third label map corresponding to the one pixel.

11. The information processing apparatus according to claim 5, wherein the fusion means generates a map as the fused label map, which includes pixels having a weighted average value of pixel values ​​calculated based on the plurality of second label maps.

12. The information processing apparatus according to claim 5, wherein the label generation means obtains a manual label map which includes pixels having a first pixel value indicating that they belong to the detection means and pixels having a second pixel value indicating that they do not belong to the detection means, which are generated by a human determining pixels in the input image that correspond to a part of the object to be detected; generates a third label map by changing the pixel values ​​of at least some pixels around a region consisting of a plurality of pixels having the first pixel value in the manual label map from the second pixel value to the first pixel value; and generates the low-noise label map by calculating the product of the pixel value of one pixel in the fused label map and the pixel value of the pixel in the third label map corresponding to the one pixel.

13. The information processing apparatus according to claim 9, wherein the weights relating to the weighted mean are set according to the accuracy of each of the plurality of second models.

14. The information processing apparatus according to claim 11, wherein the weights relating to the weighted average value are set according to the intensity relating to the predetermined image processing.

15. An information processing method that generates a low-noise label map relating to a detection target contained in an input image, estimates pixels in the input image that correspond to a part of the detection target using a first segmentation model, generates a first label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the detection target, and trains the first model using the low-noise label map and the first label map.

16. A recording medium on which a computer program is stored that causes a computer to execute an information processing method that generates a low-noise label map relating to a detection target contained in an input image, estimates pixels in the input image that correspond to a part of the detection target using a first segmentation model, generates a first label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to the detection target, and trains the first model using the low-noise label map and the first label map.

17. Information processing device comprising: a first segmentation means that uses a first segmentation model to estimate pixels corresponding to any part of any of the multiple detection targets in an input image containing a plurality of detection targets, and generates an estimated label map that includes pixels having pixel values ​​corresponding to the likelihood of belonging to any of the plurality of detection targets; and a learning means that uses the estimated label map and the ground truth label map to train the first model, wherein the estimated label map includes a plurality of first regions corresponding to each of the plurality of detection targets, each of the plurality of first regions is separated from the other first regions, the ground truth label map includes a plurality of second regions corresponding to each of the plurality of detection targets, and the learning means calculates the loss between one of the plurality of first regions and one of the plurality of second regions corresponding to the one first region, and trains the first model based on the plurality of losses corresponding to each of the plurality of first regions.