Information processing device, information processing method, and program

The information processing apparatus addresses the accuracy decrease in CNNs by adding plausible images to the periphery of input images and removing edge regions from feature maps, thereby improving the performance of CNNs in image classification and detection tasks.

WO2025127032A1PCT designated stage expired Publication Date: 2025-06-19CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/043623
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-05
Filing Date
2024-12-10
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Convolutional neural networks (CNNs) face a decrease in accuracy for classification or detection tasks when targets are located on the edges of images, due to the padding process generating non-contributory data that affects the convolution process.

Method used

The proposed solution involves an information processing apparatus that acquires an image, adds a plausible image to its peripheral regions using a machine learning model, and then inputs the modified image into another machine learning model for analysis. This process generates a feature map that removes a predetermined edge region, improving analysis accuracy.

Benefits of technology

This approach effectively reduces the decrease in accuracy of analysis processing by minimizing the impact of padding-generated data, especially at image edges, thereby enhancing the overall performance of CNNs in image classification and detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024043623_19062025_PF_FP_ABST
    Figure JP2024043623_19062025_PF_FP_ABST
Patent Text Reader

Abstract

An image processing device according to the present disclosure, which is an information processing device, comprises: an image acquisition means for acquiring a first image to be analyzed; an image addition means for inputting the first image to a first machine learning model and acquiring a second image by adding an image to a peripheral region, of the first image, which is not depicted in the first image; and an analysis means for inputting the second image to a second machine learning model and acquiring an analysis result of the second image.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present invention relates to an information processing device, an information processing method, and a program.

[0002] In recent years, convolutional neural networks (CNNs) have been adopted in image processing applications that perform image classification, object detection, semantic segmentation, etc. CNNs are a type of deep learning technology that repeatedly executes convolution processing, and are capable of performing image processing with high accuracy, as described in Non-Patent Documents 1 and 2, for example.

[0003] Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun, "Deep Residual Learning for Image Recognition", <https: / / arxiv.org / abs / 1512.03385> (Published on December 10, 2015) Jiahui Yu, Zhe Lin, Jimei Yang, et al., "Generative Image Inpainting With Contextual Attention", <https: / / arxiv.org / abs / 1801.07892> (released January 24, 2018)

[0004] However, many CNNs have issues due to padding (a process that adjusts the size of an output image by filling in the periphery of an input image with a predetermined value). For example, a CNN can accurately process classification or detection targets that exist in the center of an image, but the accuracy decreases for classification or detection targets that exist on the periphery of the image. This is because data that does not contribute to improving accuracy is generated during padding and other processes, and the generated data is incorporated into the results of the convolution process.

[0005] The technology of the present disclosure has been made in consideration of the above, and aims to reduce a decrease in accuracy of analysis processing in a machine learning model to which an image to be analyzed is input.

[0006] The information processing device according to the present disclosure includes an information processing device comprising: image acquisition means for acquiring a first image of an object to be analyzed; image addition means for inputting the first image into a first machine learning model and acquiring a second image in which an image is added to a peripheral region of the first image that is not depicted in the first image; and analysis means for inputting the second image into a second machine learning model and acquiring an analysis result of the second image. The information processing device according to the present disclosure also includes an information processing device comprising: image acquisition means for acquiring an image of an object to be analyzed; and analysis means for inputting the image into a machine learning model and acquiring an analysis result of the image, wherein the machine learning model generates a feature map indicating feature amounts related to an object to be analyzed included in the image, removes a predetermined peripheral region from the feature map, and outputs the analysis result using the feature map from which the predetermined peripheral region has been removed. The information processing device according to the present disclosure also includes an information processing device comprising: an image acquisition means for acquiring a first image of an object to be analyzed; an image addition means for acquiring a second image in which an image is added to a peripheral area of ​​the first image that is not depicted in the first image; and an analysis means for inputting the second image into a machine learning model to acquire an analysis result of the second image, wherein the machine learning model generates a feature map indicating features related to the object to be analyzed included in the second image, removes a predetermined peripheral area from the feature map, and outputs the analysis result using the feature map from which the predetermined peripheral area has been removed.

[0007] The present disclosure also relates to an information processing method comprising: an image acquisition step of acquiring a first image of an analysis target; an image addition step of inputting the first image into a first machine learning model to acquire a second image in which an image is added to a peripheral region of the first image that is not depicted in the first image; and an analysis step of inputting the second image into a second machine learning model to acquire an analysis result of the second image. The present disclosure also relates to an information processing method comprising: an image acquisition step of acquiring an image of an analysis target; and an analysis step of inputting the image into a machine learning model to acquire an analysis result of the image, the machine learning model generating a feature map indicating feature amounts related to an analysis target included in the image, removing a predetermined peripheral region from the feature map, and outputting the analysis result using the feature map from which the predetermined peripheral region has been removed. The information processing method according to the present disclosure also includes an information processing method comprising: an image acquisition step of acquiring a first image of an object to be analyzed; an image addition step of acquiring a second image in which an image is added to a peripheral area of ​​the first image that is not depicted in the first image; and an analysis step of inputting the second image into a machine learning model to acquire an analysis result of the second image, wherein the machine learning model generates a feature map indicating features related to the object to be analyzed included in the second image, removes a predetermined peripheral area from the feature map, and outputs the analysis result using the feature map from which the predetermined peripheral area has been removed.

[0008] According to the technology of the present disclosure, it is possible to reduce the decrease in accuracy of the analysis process in a machine learning model to which an image to be analyzed is input.

[0009] FIG. 1 is a diagram illustrating the configuration of an information processing device according to an embodiment. FIG. 2 is a flowchart of processing executed by an information processing device according to an embodiment. FIGS. 3A to 3D are diagrams illustrating the relationship between an image to be analyzed and a restored image according to an embodiment. FIGS. 4A to 4D are diagrams illustrating restored images according to an embodiment. FIGS. 5A and 5B are diagrams illustrating an example of padding processing according to an embodiment. FIG. 6 is a diagram illustrating another example of padding processing according to an embodiment. FIGS. 7A to 7D are diagrams illustrating the relationship between an image to be analyzed and a restored image according to an embodiment. FIG. 8 is a diagram illustrating an example of processing by a CNN classifier according to an embodiment. FIG. 9 is a diagram illustrating another example of processing by a CNN classifier according to an embodiment. FIGS. 10A to 10F are diagrams illustrating trimming processing of a restored image according to an embodiment.

[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the present disclosure is not limited to the following embodiments and can be modified as appropriate without departing from the spirit of the present disclosure. In the drawings described below, components having the same functions are denoted by the same reference numerals, and their description may be omitted or simplified.

[0011] Hereinafter, with reference to the drawings, embodiments and variations of the information processing device, information processing method, and program will be described in detail. The embodiments can be combined with conventional technology, other embodiments, or variations to the extent that no contradiction occurs. Similarly, the variations can be combined with conventional technology, embodiments, or other variations to the extent that no contradiction occurs. In the following description, similar components will be assigned common reference numerals, and duplicate descriptions may be omitted.

[0012] First Embodiment FIG. 1 is a diagram illustrating an example of the configuration of an information processing device 100 according to a first embodiment. The information processing device 100 acquires image data including an image to be analyzed and performs information processing on the image data. The information processing device 100 may display the results of the information processing, store the results of the information processing in association with the image data, or output the results of the information processing to an external device. For example, the information processing device 100 may generate data indicating an analysis result regarding whether an object of interest (analysis target) is depicted in the image data, and display the analysis result on a display device (not shown) connected to the information processing device 100. The information processing device 100 is realized by a computer device such as a server or a workstation.

[0013] The information processing device 100 may be communicably connected to a data management device (not shown) via a network 200 or a communication cable or communication circuit (not shown) in order to acquire image data to be processed and store analysis results of the image data. Such a data management device is a device that stores various data such as image data and analysis results of the image data, and can send and receive image data to other devices such as the information processing device 100 that can communicate with the data management device. The data management device may be configured integrally with the information processing device 100 as one of the components of the information processing device 100.

[0014] There are no particular limitations on the type of image data processed by the information processing device 100, and medical images captured by various devices may be used as image data. Such devices may include, for example, a fundus camera, an OCT (Optical Coherence Tomography) imaging device, and a CT (Computed Tomography) device. Furthermore, such devices may include an ultrasound diagnostic device and a magnetic resonance imaging (MRI) device. Furthermore, such devices may include a PET (Positron Emission Tomography) device and a SPECT (Single Photon Emission Computed Tomography) device. Furthermore, the device may include a whole slide scanner or the like. For example, photographic images of people, animals, man-made objects, landscapes, celestial bodies, etc. taken with a digital camera may also be included in the image data of this embodiment. For example, images of documents, forms, paintings, etc. scanned with a document scanner may also be included in the image data of this embodiment. Note that this image data is not limited to two-dimensional image data, and may be multidimensional image data such as three-dimensional image data.

[0015] 1, the information processing device 100 includes a communication interface 101, a memory circuit 102, a processing circuit 103, an input interface 104, and a display 105. The information processing device 100 can be communicatively connected to a network 200 via the communication interface 101.

[0016] The communication interface 101 is an interface for communicating image data, analysis results, etc. with other devices. The communication interface 101 is realized by a network communication interface such as a network adapter or a NIC (Network Interface Controller). The communication interface 101 may also be realized by a device connection interface such as USB (Universal Serial Bus), PCI Express, SATA (Serial ATA), or M.2.

[0017] The memory circuitry 102 stores various data and programs used in the processes executed by the information processing device 100 according to this embodiment. Specifically, the memory circuitry 102 is connected to the processing circuitry 103 and stores image data and analysis results under the control of the processing circuitry 103. The memory circuitry 102 also functions as a work memory that temporarily stores various data used in the processes executed by the processing circuitry 103. The memory circuitry 102 is realized, for example, by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, a hard disk, an optical disk, or the like.

[0018] The processing circuitry 103 controls the operation of each of the above-mentioned components of the information processing device 100. For example, the processing circuitry 103 performs various processes in response to instructions received from a user via an input interface 104 connected to the information processing device 100. Alternatively, for example, the processing circuitry 103 may perform various processes in response to instructions received from a user via the communication interface 101. Alternatively, for example, the processing circuitry 103 may perform various processes after detecting that image data has been stored in the memory circuitry 102. The processing circuitry 103 is realized, for example, by a CPU (Central Processing Unit).

[0019] The processing circuit 103 includes, for example, an image acquisition circuit 103a that realizes an image acquisition means for acquiring an image to be analyzed, a restoration circuit 103b that realizes an acquisition means for acquiring a restored image of the acquired image, and an analysis circuit 103c that realizes an analysis means for performing analysis using the restored image.

[0020] Here, for example, each processing function realized by the image acquisition circuit 103a, the restoration circuit 103b, and the analysis circuit 103c, which are components of the processing circuit 103 shown in FIG. 1, is stored in the storage circuit 102 in the form of a computer-executable program. The processing circuit 103 reads each program from the storage circuit 102 and executes the read program to realize the function corresponding to each program. That is, the processing circuit 103 that reads each program has the functions realized by the image acquisition circuit 103a, the restoration circuit 103b, and the analysis circuit 103c. As a result, the image acquisition circuit 103a functions as an image acquisition means that acquires an image to be analyzed (first image). Furthermore, the restoration circuit 103b functions as an image restoration means (image addition means) that inputs the image (first image) into a machine learning model (restoration model) and acquires a restored image (second image) in which an image is added to a peripheral area of ​​the image that was not depicted in the image (first image). The analysis circuit 103c also functions as an analysis unit that inputs the restored image (second image) into a machine learning model (analysis model) and acquires an analysis result of the restored image (second image). Note that adding a plausible image to a peripheral area that was not originally depicted as an image will hereinafter be referred to as restoration.

[0021] The input interface 104 accepts various instructions and information input operations from a user of the information processing device 100. Specifically, the input interface 104 is connected to the processing circuit 103 and converts the input operations received from the user into electrical signals and transmits them to the processing circuit 103. For example, the input interface 104 may be realized by a trackball, a switch button, a mouse, a keyboard, or a touchpad that performs input operations by touching the operation surface. Alternatively, the input interface 104 may be realized by a touchscreen that integrates a display screen and a touchpad, a non-contact input interface using an optical sensor, a voice input interface, or the like. Note that the input interface 104 is not limited to those that include physical operation components such as a mouse and a keyboard. For example, an electrical signal processing circuit that receives electrical signals corresponding to input operations from an external input device provided separately from the information processing device 100 and transmits these electrical signals to the processing circuit 103 is also included as an example of the input interface 104.

[0022] The display 105 displays various data such as image data processed by the information processing device 100 and data based on analysis results. Specifically, the display 105 is connected to the processing circuitry 103 and displays various data received from the processing circuitry 103. For example, the display 105 displays medical images based on image data. For example, the display 105 is realized by a liquid crystal monitor, a CRT (Cathode Ray Tube) monitor, a touch panel, or the like.

[0023] The above is a description of an example of the configuration of the information processing device 100 according to this embodiment. In this embodiment, the information processing device 100 performs various processes described below to reduce the impact of performance degradation of a convolutional neural network (CNN) in the peripheral portions of an image to be analyzed. Below, examples of various processes performed on image data by the information processing device 100 are described. In the following description, it is assumed that the information processing device 100 performs image classification processing on two-dimensional image data as an example of analysis processing. However, the information processing device 100 can also be configured to perform similar processing on multidimensional image data such as three-dimensional image data.

[0024] The image classification process performed by the information processing device 100 is a process of classifying images according to whether or not an object of interest (analysis target) is depicted in the image data. Specifically, there is a binary classification process in which an image is classified as positive if the analysis target is depicted in the image data, and as negative if the analysis target is not depicted. Another example is a multi-class classification process in which, if two or more types of objects are depicted in the image data, the image is classified according to the types. In the description of this embodiment, for ease of understanding, a case in which binary classification processing is performed will be described as an example, but the following description may also be interpreted as a case in which multi-class classification processing is performed.

[0025] The restoration model realized by the restoration circuit 103b and the analysis model realized by the analysis circuit 103c in the processing executed by the information processing device 100 according to this embodiment will be described below.

[0026] The restoration model realized by the restoration circuit 103b according to this embodiment is an image processing model that applies machine learning technology and restores an input image by adding (rendering) an image to a peripheral area not rendered in the input image. The restoration model realized by the restoration circuit 103b is the first machine learning model in this embodiment. Specifically, as shown in FIGS. 3A and 3B , the restoration model restores an input image Im100 (first image) by adding (rendering) an image to at least a portion of the peripheral area of ​​the image, and outputs a restored image Im101 (second image). In this case, as shown in FIG. 3C , the restored peripheral area is the portion indicated by region Re101, relative to region Re100 corresponding to the entire area of ​​image Im100.

[0027] It is preferable that the restored image of the peripheral region Re101 plausibly and continuously depicts the features depicted in the image data input to the restoration model. Specifically, the restored peripheral region (corresponding to region Re101) in restored image Im101 corresponding to image Im100 is taken as an example. It is preferable that restored image Im101 is an image in which features such as patterns and contours related to the figures and background depicted in image Im100 are also plausibly and continuously depicted in the peripheral region. For example, as shown in FIG. 3D , consider image Im102, in which a partial image filled with a predetermined pixel value is added to image Im100 as a peripheral region. Compared to image Im102, restored image Im101 is preferable because the features depicted in image Im100 are plausibly and continuously depicted in the peripheral region of image Im100.

[0028] 4A to 4D, an example of a case in which image data other than image Im100 is input to the restoration model is described. As shown in FIG. 4A, the information processing device 100 generates restored image data from image Im103, which captures the fundus retinal layer acquired by an OCT imaging device. When image Im103 is input to the restoration model implemented by the restoration circuit 103b, a restored image Im104 is output. As shown in FIG. 4B, the restored image Im104 plausibly depicts the continuation of the fundus retinal layer that was not depicted in image Im103 in the peripheral region of image Im103.

[0029] As an image processing algorithm included in a restoration model for outputting restored image data, there is a technology related to inpainting processing using machine learning, disclosed in Non-Patent Document 2. Inpainting processing is a process for restoring a partially filled-in image by replacing the filled-in portion with a plausible partial image. For example, when inpainting processing is performed on an image of an old photograph with a damaged portion, the damaged portion is replaced with a plausible image, thereby restoring the image to an undamaged state. In this embodiment, as an example, it is assumed that inpainting processing is performed on image Im100 shown in FIG. 3A. In this case, image Im102 is generated by adding a partial image filled with a predetermined pixel value around image Im100, which corresponds to the filled-in portion in the inpainting processing. The generated image Im102 is then input to a restoration model that performs the inpainting processing.

[0030] In this case, the machine learning model that performs the inpainting process provided in the restoration model is assumed to have been trained using a dataset composed of images of the same type as or simulating images expected to be input to the restoration model. For example, in this embodiment, a certain original image data is used as ground truth data, and image data created by filling in the periphery of the original image data with a predetermined pixel value is used as input to train the machine learning model that performs the inpainting process.

[0031] Furthermore, even if the image data input to the restoration model is originally negative, if a feature classified as positive is depicted in the restored peripheral region, the image data may be erroneously classified as positive, resulting in an erroneous analysis result. Therefore, the restoration model may be trained by being guided to depict only negative-like features in the peripheral region. For example, the images constituting the training dataset may be composed entirely of negative images, or a penalty may be imposed when the restoration model outputs an image depicting a positive feature, thereby training the restoration model not to depict positive features. In this way, the restoration model is trained to restore the peripheral region so that features that may cause an erroneous analysis result are not depicted in the peripheral region, thereby preventing positive features from being depicted in the peripheral region restored by the restoration model.

[0032] In addition, when the image data restored by the restoration model is multidimensional image data of three or more dimensions, the peripheral area to be restored is determined according to the number of dimensions of the image data. For example, assume that three-dimensional image data is the object of analysis. In this case, compared to when two-dimensional image data is the object of analysis, an image in which the peripheral area corresponding to the front side or the back side of the image data is restored taking into account the dimension in the depth direction is output from the restoration model.

[0033] Furthermore, the restoration model does not need to be implemented in the information processing device 100; for example, the restoration circuit 103b may be implemented in a restoration device (not shown). In this case, the information processing device 100 may be connected to the restoration device via the network 200 and transmit image data including the acquired image to be analyzed to the restoration device to realize the processing of step S102 (described later). Alternatively, for example, it is assumed that the information processing device 100 is connected via the network 200 to a data management device (not shown) that stores configuration data and parameter data related to the restoration model. In this case, the processing circuit 103 may acquire the above data from the data management device and reproduce the restoration model within the information processing device 100 to realize the processing of step S102 (described later).

[0034] The analysis model realized by the analysis circuit 103c according to this embodiment is a classifier that applies machine learning technology and performs classification processing using a CNN, and is the second machine learning model according to this embodiment. Furthermore, the analysis model performs padding processing before at least one convolution processing so that the spatial size (width, height) of the input tensor and output tensor of the convolution processing does not change. Specifically, an image processing algorithm for performing classification processing included in the analysis model is a classifier technology using machine learning disclosed in Non-Patent Document 1.

[0035] In the analytical model having a classifier using machine learning for binary classification processing in this embodiment, for example, the number of output nodes in the fully connected layer, which is the final layer in the classifier's network model, is defined as one or two. When the number of output nodes is defined as one, the analytical model, for example, normalizes the output value to a range of 0 to 1 using a sigmoid function, and classifies the normalized value as negative if it is less than a predetermined threshold, and as positive if it is equal to or greater than the threshold. In this case, the threshold can be set to 0.5, which is the median of the range after normalization, or a threshold that can accurately classify validation data in the dataset on which the classifier was trained. Furthermore, when the number of output nodes is defined as two, the analytical model classifies the first output value as negative if it is greater than the second output value, and classifies the second output value as positive if it is greater than the first output value.

[0036] The analytical model is assumed to be trained using a dataset composed of a group of images of the same type as the images expected to be input to the analytical model or a group of images simulating such images. Furthermore, the analytical model does not need to be implemented in the information processing device 100; for example, the analysis circuit 103c may be implemented in an analysis device (not shown). In this case, the information processing device 100 may be connected to the analysis device via the network 200 and transmit image data including the acquired restored image to the analysis device to realize the processing of step S103 (described below). Alternatively, for example, it is assumed that the information processing device 100 is connected to a data management device (not shown) via the network 200 that stores configuration data and parameter data related to the analytical model. In this case, the processing circuit 103 may acquire the above data from the data management device and reproduce the analytical model within the information processing device 100 to realize the processing of step S103 (described below).

[0037] Generally, CNN classifiers are often designed to perform padding before performing convolution processing to prevent the spatial size of the input tensor and output tensor of the convolution processing from changing. One reason for this is that in a network model requiring residual feature extraction, as disclosed in Non-Patent Document 1, the input tensor and output tensor of a processing block including convolution processing are added. Another reason is that if the spatial size of the tensor changes each time a convolution processing is performed, an interpolation process that adversely affects accuracy is required for pooling processing or tensor combining processing.

[0038] In addition, in general, a fixed value such as 0 is often set as the pixel value of the pixel in the padding region in the padding process. This will be specifically explained using FIGS. 5A, 5B, and 6. As shown in FIG. 5A, it is assumed that a convolution process with a kernel size of 3×3 (width×height) is performed on a tensor Te100, which is a feature amount map indicating feature amounts related to an analysis object included in a restored image. The feature amounts related to the analysis object are values ​​output from each module constituting a network model, and as an example, these values ​​are expressed as tensors.

[0039] Here, tensor Te101 is generated by padding tensor Te100 by 1×1 (one pixel above, below, left, and right) so that the spatial size of tensor Te102 after convolution is the same as that of tensor Te100. The generated tensor Te101 is then used as the input for convolution. At this time, fixed values ​​are often set to pixels in a one-pixel peripheral region of tensor Te101, and these fixed values ​​are set regardless of the state of the original tensor Te100.

[0040] Furthermore, compared to the central region, the peripheral region of the tensor Te102 affected by this fixed value may not have features calculated that are useful for improving accuracy. Furthermore, when multiple convolution processes are performed, pixels of the tensor affected by the fixed value affect adjacent pixels through the convolution process. Therefore, as the convolution process is repeated, the influence of the fixed value spreads from the peripheral region of the tensor toward the center.

[0041] Here, tensor Te102 is compared with tensor Te105, which is obtained by convolution processing of tensor Te103, which has a spatial size larger than tensor Te100. In this case, the influence of fixed values ​​due to padding processing is relatively smaller for tensor Te105, which has a larger spatial size, than for tensor Te102. Specifically, the degree of influence of fixed values ​​due to padding processing on spatial size can be calculated by dividing the number of pixels affected by fixed values ​​by the spatial size. For example, the degree of influence of fixed values ​​in tensor Te102 of FIG. 5A is 8 / 9 = approximately 89%, and the degree of influence of fixed values ​​in tensor Te105 of FIG. 5B is 16 / 25 = 64%. Therefore, the larger the spatial size of the tensor, the relatively smaller the degree of influence of fixed values ​​due to padding processing.

[0042] In the example shown in FIG. 6 , a tensor Te110, which is an image to be analyzed, is input to a network model Ne110, which is a feature extractor that repeatedly performs convolution and pooling processes, to obtain an output tensor Te111. In this example, the central region of the tensor Te111 obtained as the output of the network model Ne110 contains high-quality features (dark areas in the figure) that are not affected by fixed values ​​due to padding or are relatively little affected by them and contribute to improving the accuracy of the analysis. On the other hand, the peripheral regions of the tensor Te111 are regions that are repeatedly affected by fixed values ​​due to padding and contain low-quality features (light areas in the figure) that are unlikely to contribute to improving the accuracy of the analysis. Therefore, one method for obtaining highly accurate analysis results in a machine learning model equipped with a network model that repeatedly performs convolution and pooling processes is to utilize features that are little affected by fixed values ​​due to padding.

[0043] The size of the peripheral region in the restored image restored by the restoration model can be determined exploratoryally, for example, so as to increase the accuracy of the analytical model. The size of the peripheral region in the restored image restored by the restoration model can also be determined by other methods. For example, the size of the peripheral region may be determined so that the influence of a fixed value set in the padding region in the padding process performed by the network model related to the analytical model does not reach a predetermined range of the tensor output by the network model related to the analytical model. Here, the predetermined range may be, for example, a range that includes the approximate center coordinates of the tensor, or a range that corresponds to the size of the image to be analyzed when the spatial size of the restored image is transformed to the spatial size of the tensor at the same ratio.

[0044] Furthermore, the size of the peripheral area in the restored image restored by the restoration model can be determined, for example, according to the conditions of the computational cost allowed in the execution environment of this embodiment. The computational costs that are affected and change depending on the size of the peripheral area restored by the restoration model are, in particular, the computational cost of the process in which the restoration model restores the peripheral area and the computational cost of the process in which the analytical model analyzes the input restored image data. Generally, as the peripheral area restored by the restoration model becomes larger, both the computational cost of the process in which the restoration model restores the peripheral area and the computational cost of the process in which the analytical model analyzes the input restored image data increase, but the accuracy of the analytical model tends to increase. Conversely, as the peripheral area restored by the restoration model becomes smaller, both the computational cost of the process in which the restoration model restores the peripheral area and the computational cost of the process in which the analytical model analyzes the input restored image data decrease, but the accuracy of the analytical model tends to decrease. That is, the accuracy of the analytical model approaches the accuracy of analytical models according to conventional technology.

[0045] 2 is a flowchart showing an example of the flow of processing executed by the information processing device 100 according to the first embodiment. In this embodiment, as an example, when image data to be analyzed is stored in the memory circuitry 102 and the user operates the input interface 104 to issue an instruction to start processing, the processing circuitry 103 starts the processing shown in FIG. 2. The following describes the steps of the flowchart shown in FIG. 2.

[0046] In step S101, the image acquisition circuitry 103a acquires image data to be analyzed that is stored in the memory circuitry 102.

[0047] In step S102, the restoration circuit 103b inputs the image data of the analysis target acquired in step S101 into the restoration model, and acquires restored image data in which the peripheral area of ​​the image of the analysis target is restored.

[0048] In step S103, the analysis circuit 103c inputs the restored image data acquired in step S102 into the analysis model, and acquires the result of the classification process output from the analysis model as the analysis result.

[0049] In step S104, the analysis circuit 103c outputs the analysis results of the restored image data output from the analysis model in step S103 to the display 105. As a result, the analysis results using the image to be analyzed are displayed on the display 105. Note that the analysis results may be output as data from the information processing device 100 to a data management device (not shown) via the network 200 or a communication cable or communication circuit (not shown). Furthermore, when the analysis results of the restored image data are output as data, the analysis results do not need to be output to the display 105.

[0050] As described above, the information processing device 100 according to this embodiment can generate analysis results for restored image data in which the peripheral region of the image to be analyzed is restored. By restoring the peripheral region of the image to be analyzed, the restored image data becomes a larger image compared to the image data to be analyzed. This reduces the influence of fixed values ​​in the padding process on image data indicating the feature quantities of the image data to be analyzed, output from the network model of the analysis model. As a result, it is possible to suppress a decrease in classification performance in the peripheral portions of the image due to the padding process of the CNN. Furthermore, by using a restored image in which the peripheral region has been restored, the phenomenon of the edges of the image data being cut off is suppressed, which is expected to improve the accuracy of the analysis results.

[0051] Next, modified examples (modifications 1 and 2) of the processing of the information processing device 100 in the first embodiment will be described. In the following description, configurations and processing similar to those of the information processing device 100 described above will be assigned the same reference numerals, and detailed description thereof will be omitted.

[0052] (Variation 1) In the above embodiment, the restoration circuit 103b of the information processing device 100 generates a restored image using the top, bottom, left, and right regions of the image Im100 to be analyzed as the restoration region, such as region Re101 illustrated in Fig. 3C. On the other hand, in this modification, as shown in Fig. 7C and 7D, only a portion of the peripheral region of the image to be analyzed may be restored, such as regions Re103 and Re104 on the left and right sides excluding the top and bottom of image Im107. As a result, in this modification, the data size of the restored image can be reduced compared to the restored image generated in the above embodiment.

[0053] According to this modification, the information processing device 100 can reduce the calculation cost compared to the analysis process in the above embodiment by reducing the size of the restored image data. Furthermore, for example, it may be known that the analysis target is not depicted in an area that is not the restoration target (e.g., the upper and lower peripheral areas of image Im107 shown in FIG. 7A). In this case, by not including some of the peripheral areas in the restoration target, the calculation cost associated with the analysis process can be reduced, and it is expected that classification results with the same accuracy as the analysis process in the above embodiment can be obtained.

[0054] (Variation 2) In this variation, the size of the peripheral region restored by the restoration model realized by the restoration circuit 103b is changed. This makes it possible to adjust the calculation cost related to the restoration process of the restoration model and the calculation cost related to the analysis process of the analysis model realized by the subsequent analysis circuit 103c. Furthermore, by changing the size of the peripheral region restored by the restoration model, it is possible to adjust the accuracy of the analysis results output from the analysis model.

[0055] For example, an administrator of the information processing device 100 according to this modification may store in the storage circuitry 102 a default value for the size of the peripheral region of an image restored by the restoration model, and may read it out when executing the process shown in Fig. 2 to adjust the calculation cost and the accuracy of the analysis results. Also, for example, a user of the information processing device 100 according to this modification may operate the input interface 104 to change the size of the peripheral region restored by the restoration model before starting the process shown in Fig. 2.

[0056] According to this modified example, the balance between the computational cost and the accuracy of the analysis results is adjusted compared to the processing of the information processing device 100 according to the above embodiment, and a more suitable form of execution of the analysis processing by the information processing device 100 can be realized.

[0057] Second Embodiment Next, an example of the configuration and processing of an information processing device according to a second embodiment will be described. In the following description, the configuration, processing, etc. of the information processing device similar to those of the first embodiment will be denoted by the same reference numerals, and detailed description thereof will be omitted.

[0058] The analysis model realized by the analysis circuit 103c of the information processing device 100 according to this embodiment is a classifier that performs classification processing using a CNN using machine learning technology, similar to the first embodiment. However, the information processing device 100 according to this embodiment differs from the first embodiment in that it performs a trimming process on a feature tensor that indicates the feature quantities of an image output from a network model provided by the classifier. Specifically, the trimming process removes a predetermined peripheral region of the feature tensor before it is input to a fully connected layer, which is the final layer in the network model provided by the classifier, or a pooling layer corresponding to that fully connected layer. This reduces the influence of fixed values ​​in the padding process performed by the network model.

[0059] First, a typical CNN-based classifier will be described with reference to FIG. 8 . In a typical CNN-based classifier, a tensor Te200 corresponding to an image to be classified is input to a network model Ne200, which is a feature extractor. The tensor Te200 then passes through a convolutional layer, a pooling layer, and the like included in the network model Ne200, and is output from the network model Ne200 as a tensor Te201. The tensor Te201 then passes through the pooling layer Ne201 and is transformed into a tensor Te202 that can be input to the fully connected layer Ne202. The fully connected layer Ne202 then performs linear regression processing on the tensor Te202 so that the number of values ​​corresponding to the number of output nodes defined in the fully connected layer Ne202 is output. At this time, the tensor Te201 is affected by fixed values ​​due to padding processing associated with the convolutional layer included in the network model Ne200, and there is a possibility that the tensor Te201 will not be classified as positive even if the object to be analyzed is depicted in the peripheral portion of the image data. This can cause degradation of the analytical model's performance.

[0060] Next, a network model, which is a classifier included in the analysis model realized by the analysis circuit 103c according to this embodiment, will be described with reference to FIG. 9 . In the network model included in the analysis model, a tensor Te203 corresponding to a restored image in which the peripheral region of the image to be classified is restored is input to a network model Ne203, which is a feature extractor. The tensor Te203 then passes through a convolutional layer, a pooling layer, and the like included in the network model Ne203, and is output from the network model Ne203 as a tensor Te204. The tensor Te204 is then trimmed to fit a predetermined image shape based on an image region affected by a fixed value due to padding caused by the network structure of the network model Ne203. This trimming process generates a tensor Te205. The tensor Te205 then passes through the pooling layer Ne205 and is transformed into a tensor Te206, which has a shape that can be input to a fully connected layer Ne206. Then, the tensor Te 206 is subjected to linear regression processing by the fully connected layer Ne 206 so that values ​​equal to the number of output nodes defined in the fully connected layer Ne 206 are output.

[0061] In this case, the area affected by the fixed value in the tensor Te204 can be determined based on the network structure of the network model Ne203. Specifically, the area affected by the fixed value can be determined based on, for example, the padding size in the padding process by the network model Ne203, the kernel size related to the convolution process, and the number of times the convolution process is executed. Alternatively or in addition to this, the area affected by the fixed value can also be determined based on the process content and order in which the padding area caused by the padding process is propagated to the tensor, such as a pooling process or a tensor scaling process. In other words, the peripheral area affected by the fixed value caused by the padding process is removed from the tensor, and an image having features calculated from the pixel values ​​of the trimmed image is input to the pooling layer Ne205.

[0062] Depending on the spatial size of the tensor Te203 input to the network model Ne203 and the number of convolution and pooling processes performed by the network model Ne203, the fixed value of the padding process may affect all pixels of the tensor Te205. In this case, for example, when training a classifier included in the analysis model, a region other than a region corresponding to a predetermined spatial size determined exploratoryally to achieve the highest accuracy may be trimmed as a peripheral region. Alternatively, for example, the number of times that each pixel constituting the tensor Te205 is affected by the fixed value of the padding process may be calculated, and a region that is affected by the fixed value a predetermined number of times or more may be trimmed as a peripheral region. Alternatively, for example, by increasing the size of the peripheral region restored by the restoration model realized by the restoration circuit 103b according to this embodiment, it is possible to avoid a situation in which all pixels of the tensor Te205 are affected by the fixed value of the padding process.

[0063] The flow of an example of processing performed by the information processing device 100 according to this embodiment is the same as that of the first embodiment, and therefore a detailed description thereof will be omitted here. As described above, the information processing device 100 according to this embodiment can acquire analysis results for an image to be analyzed. In particular, a feature of this embodiment is that a trimming process is performed on a feature tensor calculated by a network model included in the analysis model based on an area affected by a fixed value due to padding. As a result, compared to the first embodiment, the area affected by a fixed value due to padding is removed as a peripheral area, thereby further suppressing performance degradation of the analysis process in the peripheral area of ​​the image caused by the CNN padding process.

[0064] Next, modified examples (modifications 3, 4, and 5) of the processing of the information processing device 100 in the second embodiment will be described. In the following description, configurations and processing similar to those of the information processing device 100 described above will be assigned the same reference numerals, and detailed description thereof will be omitted.

[0065] In this modification, the size of the edge region of the tensor to be trimmed is changed in the tensor trimming process executed in the network model included in the analysis model realized by the analysis circuit 103 c. This makes it possible to adjust the computational cost and accuracy of the network model included in the analysis model without completely eliminating the influence of fixed values ​​due to the padding process associated with the convolutional layer of the network model.

[0066] According to this modified example, in the information processing device 100, the balance between the calculation cost and the accuracy of the analysis results is adjusted compared to the information processing device 100 according to the second embodiment described above, and an execution form suitable for the execution environment can be achieved.

[0067] (Modification 4) In this modification, acquisition of restored image data by the restoration model realized by the restoration circuit 103b is skipped, and image data to be analyzed is directly input to the analysis model realized by the analysis circuit 103c.

[0068] The series of processes executed by the information processing device 100 according to this modification differs from the processes described in the second embodiment in the following process. That is, the process of step S102 is omitted, and the information processing device 100 proceeds to step S103 after completing the process of step S101. In step S103, the analysis circuit 103c inputs the image data to be analyzed acquired in step S101 into the analysis model instead of the restored image data, and acquires the result of the classification process using the analysis model as the analysis result.

[0069] According to this modification, the information processing device 100 does not perform output processing of a restored image using a restoration model, and instead inputs an analysis target image that is smaller in size than the restored image to the analysis model. This reduces the computational cost compared to the processing performed by the information processing device 100 according to the second embodiment. Furthermore, according to this modification, by trimming the tensors output from the network model included in the analysis model, the influence of fixed values ​​due to padding processing associated with the convolutional layer can be further reduced.

[0070] In this modification, the image processing algorithm included in the restoration model realized by the restoration circuit 103 b may be, for example, a rule-based image processing algorithm, such as known reflection padding, symmetric padding, or replication padding.

[0071] Specifically, in the case of the image Im103 shown in FIG. 4A, a restored image Im105 can be generated by reflection padding processing, and a restored image Im106 can be generated by replication padding processing.

[0072] When these restored image data are compared with restored image Im104, which is an example output of the restoration model described above, the image of the peripheral region is restored so that the fundus retinal layer depicted in image Im103 suddenly bends, as shown in the figure. Thus, the image of the peripheral region is not depicted contiguously, resulting in an unnatural OCT image. This raises concerns about a decrease in the accuracy of classification processing using the analysis model, but depending on the characteristics of the analysis model, this decrease in accuracy may be minor or may not occur at all.

[0073] For example, assume that an analytical model is trained to classify the presence or absence of edema in the fundus retina layer, and an OCT image of the fundus retina layer, such as image Im103 shown in FIG. 4A, is input to the restoration model. In this case, if edema is not depicted in the OCT image, the rule-based image processing algorithm described above will not depict any new edema in the peripheral region. Therefore, even if an image with a restored peripheral region is input to the analytical model, the analytical model will not output a classification result indicating the presence of edema. On the other hand, if edema is depicted in the OCT image, the rule-based image processing algorithm may depict edema in the peripheral region. However, even if an image with a restored peripheral region is input to the analytical model, the analytical model will still output a classification result indicating the presence of edema. Therefore, there is no problem because the classification results by the analytical model are the same for the OCT image and the restored image.

[0074] However, due to the portion of the fundus retina layer depicted as bent in the peripheral region of the restored image, an analytical model with insufficient accuracy may output an incorrect classification result. Therefore, to reduce the possibility of outputting an incorrect classification result, the analytical model may be trained using image data generated using the above-mentioned image processing algorithm. That is, a dataset may be constructed and the analytical model may be trained so as to output a correct classification result for image data in which the fundus retina layer is depicted as bent in the peripheral region by the image processing algorithm and is not depicted plausibly continuously.

[0075] This makes it possible to adjust the computational cost by changing the image processing algorithm for processing the restoration model, and to eliminate the need to prepare a restoration model using machine learning, compared to the processing performed by the information processing device 100 according to the second embodiment described above.

[0076] Third Embodiment Next, an example of the configuration and processing of an information processing device according to a second embodiment will be described. In the following description, the configuration, processing, etc. of the information processing device similar to those of the first embodiment will be denoted by the same reference numerals, and detailed description thereof will be omitted.

[0077] An example of image generation processing for two-dimensional image data executed by the information processing device 100 according to this embodiment will be described below. Note that processing similar to this embodiment can be easily extended to multidimensional image data such as three-dimensional image data, and may be interpreted differently as necessary.

[0078] In this embodiment, the image generation process performed by the information processing device 100 is a process for generating a new image of an analysis target based on the image acquired in step S101. Specifically, the image generation process is a segmentation process that divides regions of an analysis target, such as a shape or a lesion, depicted in the image data into different colors. The image generation process is, for example, a domain conversion process that converts image data captured with a certain modality into an image appearance that appears as if it were captured with another modality. The image generation process is, for example, an image quality improvement process that removes noise and artifacts depicted in the image data. In this embodiment, in the various image generation processes described above, a CNN padding process may cause a decrease in the accuracy of classification processing by an analytical model in the peripheral regions of the generated image.

[0079] The analysis model realized by the analysis circuit 103c according to this embodiment is an image generator that applies machine learning technology and executes image generation processing using a CNN. This analysis model performs padding processing before convolution processing so that the spatial size (width, height) of the input tensor and the output tensor does not change in at least one convolution processing. Note that the image processing algorithm that performs the image generation processing included in the analysis model may be technology related to an image generator using machine learning.

[0080] In addition, when the image generation process of the analytical model is a segmentation process, the analytical model processes the image or tensor output by the CNN according to the conditions of the segmentation process. Specifically, for example, assume that a binary segmentation process is performed in which a figure or lesion depicted in the image input to the analytical model is the foreground and the remaining parts are the background. In this case, it is assumed that the output data, which is the image or tensor output by the CNN related to the analytical model, is data with one channel of information. In this case, for example, the pixel value of each pixel in the output image can be normalized to a range of 0 to 1 using a sigmoid function, and if the normalized value is less than a predetermined threshold, it can be processed as a background pixel, and if it is greater than or equal to the threshold, it can be processed as a foreground pixel. In addition, the threshold can be set to 0.5, which is the median of the range after normalization, or a threshold that allows the validation data in the dataset used to train the image generator to generate an image with the highest accuracy. Alternatively, it is assumed that the output data, which is the image or tensor output by the CNN related to the analytical model, is data with two channels of information. In this case, for example, if the output value of the first channel is greater than the output value of the second channel, it can be determined to be negative, and if the output value of the second channel is equal to or greater than the output value of the first channel, it can be determined to be positive.

[0081] Furthermore, when the image generation process of the analytical model is a high-quality image process, the analytical model can normalize the pixel values ​​of the image or tensor output by the CNN using the number of channels and pixel value range of the image input to the analytical model.

[0082] The analytical model is assumed to be trained using a dataset consisting of a group of images of the same type as the input images or a group of images that mimic the input images. Furthermore, when the analytical model is trained by supervised learning, the dataset may include ground truth data that is expected to be output by the analytical model and corresponds to each of the images that make up the dataset.

[0083] The process executed by the information processing device 100 according to this embodiment is the same as that of the first embodiment. The process executed by the information processing device 100 according to this embodiment will be described below in accordance with the steps of the flowchart shown in Fig. 2. Note that steps that are not particularly described are the same as those of the first embodiment.

[0084] In step S103, the analysis circuit 103c inputs the restored image data into the analysis model and acquires the image generation result by the analysis model as the analysis result. This allows the information processing device 100 according to this embodiment to acquire the analysis result of a restored image in which the peripheral region of the image to be analyzed is restored. By using a restored image in which the peripheral region of the image to be analyzed is restored, the restored image becomes larger than the image to be analyzed. Therefore, the analysis result by the analysis model acquires an image larger than the image to be analyzed. Furthermore, the accuracy of the analysis process is improved in areas where there is concern about problems such as a decrease in the accuracy of the analysis process due to the influence of fixed values ​​caused by padding processing in the network model included in the analysis model (i.e., areas corresponding to the peripheral regions of the image to be analyzed).

[0085] Next, modified examples (modifications 6 and 7) of the processing of the information processing device 100 in the above-described third embodiment will be described. In the following description, configurations and processing similar to those of the above-described information processing device 100 will be assigned the same reference numerals, and detailed description thereof will be omitted.

[0086] (Variation 6) In this variation, the analysis model realized by the analysis circuit 103c performs a cropping process on the image obtained by the image generation process described above to obtain a smaller-sized image. This allows the analysis model to perform classification processing based on an image obtained by cropping the image generation result to the same size as the image to be analyzed. In other words, the predetermined peripheral area removed by the analysis model matches the restored peripheral area in the restored image. Performing classification processing using a cropped image in this way makes it easier to understand the correspondence between pixel coordinates between the image to be analyzed and the image of the analysis result.

[0087] (Variation 7) In this variation, the size of the peripheral area restored by the restoration model realized by the restoration circuit 103 b is adjusted, which makes it possible to adjust the computational cost for processing the restoration model and the computational cost for processing the analysis model realized by the subsequent analysis circuit 103 c, as well as the accuracy of the analysis result by the analysis model and the image size.

[0088] Therefore, compared to the information processing device 100 according to the third embodiment described above, the information processing device 100 according to this modified example can more appropriately adjust the balance between computational cost and accuracy of analysis results, and can perform analysis that is suitable for the execution environment.

[0089] Fourth Embodiment Next, an example of the configuration and processing of an information processing device according to a second embodiment will be described. In the following description, the configuration, processing, etc. of the information processing device similar to those of the first embodiment will be denoted by the same reference numerals, and detailed description thereof will be omitted.

[0090] An example of image generation processing for two-dimensional image data executed by the information processing device 100 according to this embodiment will be described below. Note that processing similar to this embodiment can be easily extended to multidimensional image data such as three-dimensional image data, and may be interpreted differently as necessary.

[0091] The object detection process in this embodiment is a process for identifying the location of an object to be analyzed depicted in an image. In the object detection process, each object to be analyzed identified in the image is surrounded by a bounding box, an identifier identifying the type of object is added, and the area occupied by each object is converted into a mask image. Mask imaging is also called instance segmentation processing. This makes it easier for users to identify the location of an object by including indices indicating the object, such as a bounding box, identifier, or mask image, in the image of the analysis result. However, in the object detection process, as in the classification process, there is a possibility that analysis accuracy will decrease in the peripheral portions of the image to be analyzed due to the padding process of the CNN.

[0092] The analysis model realized by the analysis circuit 103c according to this embodiment is an object detector that applies machine learning technology and performs object detection processing using a CNN. This analysis model performs padding processing before at least one convolution process so that the spatial size (width, height) of the input tensor and output tensor of the convolution process does not change. Note that the image processing algorithm that performs the object detection process included in the analysis model can be a technology related to an object detector using machine learning.

[0093] The process executed by the information processing device 100 according to this embodiment is the same as that of the first embodiment. The process executed by the information processing device 100 according to this embodiment will be described below in accordance with the steps of the flowchart shown in Fig. 2. Note that steps that are not particularly described are the same as those of the first embodiment.

[0094] In step S103, the analysis circuit 103c inputs the restored image data into the analysis model and acquires the results of the object detection process performed by the analysis model as the analysis result. This allows the information processing device 100 according to this embodiment to acquire the analysis result of a restored image in which the peripheral area of ​​the image to be analyzed is restored. By using a restored image in which the peripheral area of ​​the image to be analyzed is restored, the restored image becomes larger than the image to be analyzed. Therefore, the analysis result obtained by the analysis model is an object detection result based on an image with a wider range than the image to be analyzed. Furthermore, the accuracy of the analysis process is improved in areas where there is concern about problems such as a decrease in the accuracy of the analysis process due to the influence of fixed values ​​caused by padding processing in the network model included in the analysis model (i.e., areas corresponding to the peripheral areas of the image to be analyzed).

[0095] Next, a description will be given of a modified example (modification 8) of the processing of the information processing device 100 in the above-described fourth embodiment. In the following description, configurations and processes similar to those of the information processing device 100 described above will be assigned the same reference numerals, and detailed description thereof will be omitted.

[0096] (Modification 8) In this modification, the analysis model realized by the analysis circuit 103c may perform a cropping process on the image obtained as the object detection result, and output an image with a smaller image size as the analysis result.

[0097] Specifically, the analysis model changes the cropping process depending on the content of the object detection process. As a result, the analysis model performs analysis on an image that has been modified so that indices do not overlap with a predetermined edge region. For example, assume that the object detection process is a process of enclosing an object in an image with a bounding box. In this case, the analysis model reduces the size of a bounding box whose portion is outside the image obtained by the cropping process so that it fits within the image. Furthermore, the analysis model may remove the bounding box if its area or volume is below a predetermined size. For example, assume that the object detection process is a process of adding an identifier that identifies the type of object depicted in the image. In this case, the analysis model may remove the identifier if the object corresponding to the identifier does not fit within the image obtained by the cropping process. For example, assume that the object detection process is a process of adding a mask image to an object depicted in the image by instance segmentation. In this case, the analysis model may crop the mask image whose portion is outside the image obtained by the cropping process. In addition, if the area or volume of a mask image falls below a predetermined size, the analysis model may remove the corresponding mask image.

[0098] 10A to 10F show examples of the above-described trimming process. FIG. 10A shows image Im301 before the trimming process. Image Im301 depicts a bounding box 310 that surrounds an object depicted in the image. For convenience of explanation, a boundary 320 of the peripheral region that is trimmed by the trimming process in image Im301 is indicated by a dotted line, although it is not shown in image Im301. The region outside boundary 320 is an example of the peripheral region. By performing the above-described trimming process on image Im301, image Im302 shown in FIG. 10B is generated. As shown in the figure, a portion of the object depicted in image Im301 is depicted in image Im302. In this case, in image Im302, the bounding box 310 in image Im301 is modified to a bounding box 311 that surrounds the object.

[0099] FIG. 10C shows image Im303 before the trimming process. Two objects 312 and 313 are depicted in image Im303. Identifiers "P001" and "P002" are assigned to objects 312 and 313, respectively. For ease of explanation, a boundary 320 of the peripheral region to be trimmed in image Im303 by the trimming process is indicated by a dotted line, although it is not displayed in image Im303. The region outside boundary 320 is an example of the peripheral region. By performing the trimming process on image Im303, image Im304 shown in FIG. 10D is generated. As shown in the figure, in image Im304, object 312 does not fit within the image, and only a portion of it is depicted. In this case, the identifier "P001" corresponding to object 312 that does not fit within the image is removed from image Im304.

[0100] FIG. 10E shows image Im305 before the trimming process. A mask image 314 that covers the object is applied to image Im305. For convenience of explanation, a boundary 320 of the peripheral region that is trimmed by the trimming process in image Im305 is indicated by a dotted line, although it is not displayed in image Im305. The region outside boundary 320 is an example of a peripheral region. By performing the trimming process on image Im305, image Im306 shown in FIG. 10E is generated. As shown in the figure, a portion of the object depicted in image Im305 is depicted in image Im306. In this case, in image Im306, mask image 314 in image Im305 is corrected to mask image 315 that covers the object.

[0101] In addition, in the above-described trimming process, the trimming process may be adjusted so that the size of the image obtained as the object detection result is the same as the size of the image to be analyzed, which makes it easier to understand the correspondence between pixel coordinates between the image to be analyzed and the image of the analysis result.

[0102] (Other Embodiments) The processing of the information processing device 100 may be realized in various different forms, without being limited to the above-described embodiments and modifications. For example, the information processing device 100 may perform, on multidimensional image data, processing similar to the various processes performed on the above-described two-dimensional image data.

[0103] Furthermore, the components of each device shown in the figure are functional concepts and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions performed by each device can be realized by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic.

[0104] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method.In addition, the information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified.

[0105] The methods described in the embodiments can be realized by executing a prepared program on a computer such as a personal computer or a workstation. This program can be distributed via a network such as the Internet. The control program can be recorded on a non-transitory computer-readable recording medium such as a hard disk, a floppy disk (FD), a CD-ROM, a magneto-optical disk (MO), or a DVD, and can be read from the recording medium and executed by a computer.

[0106] Although several embodiments have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, modifications, and combinations of embodiments can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as within the scope of the invention and its equivalents as defined in the claims.

[0107] Furthermore, the disclosed technology can be embodied as, for example, a system, a device, a method, a program, or a recording medium (storage medium), etc. Specifically, the technology may be applied to a system consisting of multiple devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or may be applied to an apparatus consisting of a single device.

[0108] The present invention provides a program for causing a computer to execute each step of the image processing method according to the above-described embodiment, which is supplied to a system or device via a network or a storage medium. The program is configured to be read and executed by one or more processors in the computer of the system or device. The program can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions.

[0109] The present invention is not limited to the above-described embodiments, and various modifications and variations can be made without departing from the spirit and scope of the present invention. Therefore, the following claims are appended to apprise the public of the scope of the present invention.

[0110] This application claims priority based on Japanese Patent Application No. 2023-210815 filed on December 14, 2023 and Japanese Patent Application No. 2024-193479 filed on November 5, 2024, the entire contents of which are incorporated herein by reference.

[0111] 100 Information processing device, 103 Processing circuit, 103a Image acquisition circuit, 103b Restoration circuit, 103c Analysis circuit

Claims

1. An information processing device comprising: an image acquisition means for acquiring a first image to be analyzed; an image addition means for inputting the first image into a first machine learning model to acquire a second image in which an image is added to a peripheral area of ​​the first image that is not depicted in the first image; and an analysis means for inputting the second image into a second machine learning model to acquire an analysis result of the second image.

2. The information processing device according to claim 1, characterized in that the first machine learning model adds an image to a portion of the surrounding area of ​​the first image.

3. The information processing device according to claim 1 or 2, characterized in that the second machine learning model generates a feature map indicating features related to the object to be analyzed contained in the second image, removes a predetermined peripheral area of ​​the feature map, and outputs the analysis result using the feature map from which the predetermined peripheral area has been removed.

4. The information processing device described in claim 3, characterized in that the second machine learning model has a network model that outputs the feature map indicating the feature relating to the object to be analyzed contained in the second image.

5. The information processing device described in claim 4, characterized in that the second machine learning model inputs the feature map from which the specified peripheral region has been removed to a fully connected layer or a pooling layer corresponding to the fully connected layer, and outputs the analysis result.

6. An information processing device as described in any one of claims 1 to 5, characterized in that the second machine learning model removes a predetermined peripheral area corresponding to the first image acquired by the image acquisition means from the image showing the analysis result.

7. The information processing device of claim 6, wherein the predetermined peripheral region removed by the second machine learning model coincides with the surrounding region to which an image is added in the second image.

8. An information processing device as described in any one of claims 1 to 7, characterized in that when the analysis result includes an index indicating an object located in a specified peripheral region of the second image, the second machine learning model corrects the index so that it does not overlap with the specified peripheral region and outputs the analysis result.

9. An information processing apparatus according to claim 8, characterized in that the predetermined peripheral region of the second image coincides with the peripheral region in the second image to which an image has been added.

10. An information processing device according to claim 8 or 9, characterized in that the indicator is a bounding box surrounding the object, an identifier of the object, or a mask image overlapping the object.

11. An information processing device as described in any one of claims 1 to 10, characterized in that the first machine learning model is trained to add images to the surrounding area so that features that may cause the analysis result to be erroneous are not depicted in the surrounding area.

12. An information processing device comprising: an image acquisition means for acquiring an image of an object to be analyzed; and an analysis means for inputting the image into a machine learning model and acquiring an analysis result of the image, wherein the machine learning model generates a feature map indicating features relating to the object to be analyzed contained in the image, removes a predetermined peripheral region of the feature map, and outputs the analysis result using the feature map from which the predetermined peripheral region has been removed.

13. An information processing device comprising: an image acquisition means for acquiring a first image of an object to be analyzed; an image addition means for acquiring a second image in which an image is added to a peripheral area of ​​the first image that is not depicted in the first image; and an analysis means for inputting the second image into a machine learning model and acquiring an analysis result of the second image, wherein the machine learning model generates a feature map indicating features relating to the object to be analyzed contained in the second image, removes a predetermined peripheral area of ​​the feature map, and outputs the analysis result using the feature map from which the predetermined peripheral area has been removed.

14. An information processing device according to claim 1 or 13, characterized in that the image added to the peripheral region is an image in which the manner in which the image is depicted in the first image is depicted plausibly and continuously.

15. An information processing method comprising: an image acquisition step of acquiring a first image to be analyzed; an image addition step of inputting the first image into a first machine learning model to acquire a second image in which an image is added to a peripheral area of ​​the first image that is not depicted in the first image; and an analysis step of inputting the second image into a second machine learning model to acquire an analysis result of the second image.

16. An information processing method comprising: an image acquisition step of acquiring an image of an object to be analyzed; and an analysis step of inputting the image into a machine learning model to acquire an analysis result of the image, wherein the machine learning model generates a feature map indicating features relating to the object to be analyzed contained in the image, removes a predetermined peripheral region of the feature map, and outputs the analysis result using the feature map from which the predetermined peripheral region has been removed.

17. An information processing method comprising: an image acquisition step of acquiring a first image of an object to be analyzed; an image addition step of acquiring a second image in which an image is added to a peripheral area of ​​the first image that is not depicted in the first image; and an analysis step of inputting the second image into a machine learning model and acquiring an analysis result of the second image, wherein the machine learning model generates a feature map indicating features related to the object to be analyzed contained in the second image, removes a predetermined peripheral area of ​​the feature map, and outputs the analysis result using the feature map from which the predetermined peripheral area has been removed.

18. A program for causing a computer to execute each step of the information processing method according to any one of claims 15 to 17.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, and program

    JP2025096159A

  • Area conversion apparatus, area conversion method, and area conversion system

    JP2022129792A

  • Image processing method and device, electronic device, and storage medium

    JP2022517571A

  • Deep multi-scale networks for multi-class image segmentation

    JP2022550413A

  • JP2023210815A