Information processing apparatus, information processing method, and program
The information processing apparatus addresses the accuracy decline in CNNs due to padding by adding plausible images to the periphery of input images and removing edge regions from feature maps, thereby improving analysis accuracy.
Patent Information
- Application Number
- JP2024193479
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-11-05
- Publication Date
- 2025-06-26
AI Technical Summary
Convolutional neural networks (CNNs) face accuracy issues in image classification and object detection due to the padding process, which generates non-contributory data affecting the edges of images.
An information processing apparatus that acquires an image, adds a plausible image to its peripheral regions using a machine learning model, and then inputs the enhanced image into another machine learning model to generate a feature map from which a predetermined edge region is removed, improving analysis accuracy.
This approach reduces the decrease in analysis processing accuracy by minimizing the impact of padding-generated data, particularly at image edges, thereby enhancing the overall performance of machine learning models in image analysis tasks.
Smart Images

Figure 2025096159000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] In recent years, in applications in the field of image processing such as image classification, object detection, and semantic segmentation, a convolutional neural network (CNN) has been adopted. The CNN is a type of deep learning technology that repeatedly executes convolutional processing, and for example, can perform image processing with high accuracy in Non-Patent Documents 1 and 2.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, problems arise due to the padding process (a process of adjusting the size of the output image by supplementing the periphery of the input image with a predetermined value) included in many CNNs. For example, in CNNs, there is a problem that while classification targets or detection targets existing in the center of an image can be processed accurately, the accuracy decreases for classification targets or detection targets existing on the edges of the image. This is because data that does not contribute to the improvement of accuracy is generated in the padding process or the like, and the generated data is incorporated into the result of the convolution process.
[0005] The technology of the present disclosure has been made in view of the above, and aims to reduce a decrease in the accuracy of analysis processing in a machine learning model into which an image to be analyzed is input.
Means for Solving the Problem
[0006] The information processing apparatus according to the present disclosure includes an image acquisition unit that acquires a first image to be analyzed, an image addition unit that inputs the first image into a first machine learning model and acquires a second image in which an image is added to a peripheral region of the first image that is not depicted in the first image, and an analysis unit that inputs the second image into a second machine learning model and acquires an analysis result of the second image. Further, the information processing apparatus according to the present disclosure includes an image acquisition unit that acquires an image to be analyzed, and an analysis unit that inputs the image into a machine learning model and acquires an analysis result of the image, wherein the machine learning model generates a feature map indicating a feature amount related to an analysis target object included in the image, removes a predetermined edge region of the feature map, and outputs the analysis result using the feature map from which the predetermined edge region has been removed. Further, the information processing apparatus according to the present disclosure includes an image acquisition unit that acquires a first image to be analyzed, an image addition unit that acquires a second image in which an image is added to a peripheral region of the first image that is not depicted in the first image, and inputs the second image into a machine learning model to acquire an analysis result of the second image It includes analysis means, and the machine learning model generates a feature map indicating feature quantities related to the object to be analyzed included in the second image, removes a predetermined edge region of the feature map, and outputs the analysis result using the feature map from which the predetermined edge region has been removed. It is characterized by including an information processing apparatus.
[0007] Also, the information processing method according to the present disclosure includes an image acquisition step of acquiring a first image of an object to be analyzed, an image addition step of inputting the first image into a first machine learning model to acquire a second image in which an image is added to a peripheral region of the first image that is not depicted in the first image, and an analysis step of inputting the second image into a second machine learning model to acquire an analysis result of the second image. Further, the information processing method according to the present disclosure includes an image acquisition step of acquiring an image of an object to be analyzed, and an analysis step of inputting the image into a machine learning model to acquire an analysis result of the image, wherein the machine learning model generates a feature map indicating feature quantities related to the object to be analyzed included in the image, removes a predetermined edge region of the feature map, and outputs the analysis result using the feature map from which the predetermined edge region has been removed. Additionally, the information processing method according to the present disclosure includes an image acquisition step of acquiring a first image of an object to be analyzed, an image addition step of acquiring a second image in which an image is added to a peripheral region of the first image that is not depicted in the first image, and an analysis step of inputting the second image into a machine learning model to acquire an analysis result of the second image, wherein the machine learning model generates a feature map indicating feature quantities related to the object to be analyzed included in the second image, removes a predetermined edge region of the feature map, and outputs the analysis result using the feature map from which the predetermined edge region has been removed.
Advantages of the Invention
[0008] According to the technology of the present disclosure, in a machine learning model into which an image of an object to be analyzed is input, it is possible to reduce a decrease in the accuracy of analysis processing.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the present disclosure is not limited to the following embodiments and can be appropriately changed without departing from the gist thereof. Also, in the drawings described below, components having the same function may be denoted by the same reference numerals, and the description thereof may be omitted or simplified.
[0011] Hereinafter, each embodiment and each modification example of the information processing apparatus, the information processing method, and the program will be described in detail with reference to the drawings. Note that the embodiments can be combined with the prior art, other embodiments, or modification examples as long as there is no contradiction in the content. Similarly, the modification examples can be combined with the prior art, the embodiments, or other modification examples as long as there is no contradiction in the content. Also, in the following description, the same components may be given common reference numerals, and redundant descriptions may be omitted. This is possible.
[0012] (First Embodiment) FIG. 1 is a diagram showing an example of the configuration of an information processing apparatus 100 according to the first embodiment. The information processing apparatus 100 acquires image data including an image to be analyzed and executes information processing on the image data. Note that the information processing apparatus 100 may display the result of the information processing, save the result of the information processing in association with the image data, or output the result of the information processing to an external device. For example, the information processing apparatus 100 may generate data indicating an analysis result regarding whether an object (object to be analyzed) of interest is depicted in the image data, and display the analysis result on a display device (not shown) connected to the information processing apparatus 100. The information processing apparatus 100 is realized by, for example, a computer device such as a server or a workstation.
[0013] Note that the information processing apparatus 100 may be communicably connected to a data management device (not shown) via a network 200 or a communication cable or communication circuit (not shown) in order to acquire the image data to be processed or save the analysis result of the image data. Such a data management device is a device that stores various data such as image data and the analysis result of the image data, and can transmit and receive the image data to and from other devices such as the information processing apparatus 100 that can communicate with the data management device. Note that the data management device may be configured integrally with the information processing apparatus 100 as one component of the information processing apparatus 100.
[0014] There is no special restriction on the type of image data processed by the information processing apparatus 100, and medical images captured by various apparatuses may be adopted as the image data. The apparatuses may include, for example, fundus cameras, OCT (Optical Coherence Tomography) imaging apparatuses, CT (Computed Tomography) apparatuses. Further, the apparatuses may include ultrasonic diagnostic apparatuses, magnetic resonance imaging (MRI) apparatuses. Further, the apparatuses may include PET (Positron Emission Tomography) apparatuses, SPECT (Single Photon Emission Computed Tomography) apparatuses. Furthermore, the apparatuses may include Whole Slide Scanners and the like. Also, for example, images of photographs of people, animals, artifacts, landscapes, celestial bodies, etc. taken with a digital camera may also be included in the image data of the present embodiment. Also, for example, images of documents, forms, paintings, etc. read by a document scanner may also be included in the image data of the present embodiment. Note that these image data are not limited to 2D image data, and may be multi-dimensional image data such as 3D image data.
[0015] As shown in FIG. 1, the information processing apparatus 100 includes a communication interface 101, a storage circuit 102, a processing circuit 103, an input interface 104, and a display 105. Further, the information processing apparatus 100 can be communicably connected to the network 200 via the communication interface 101.
[0016] The communication interface 101 is an interface for communicating image data, analysis results, etc. with other devices. The communication interface 101 is realized, for example, by a network communication interface such as a network adapter or a NIC (Network Interface Controller). Also, the communication interface 101 may be realized by a device connection interface such as USB (Universal Serial Bus), PCI Express, SATA (Serial ATA), M.2, etc. It may also be realized.
[0017] The storage circuit 102 stores various data and various programs used in the processing executed by the information processing apparatus 100 according to this embodiment. Specifically, the storage circuit 102 is connected to the processing circuit 103 and stores image data and analysis results under the control of the processing circuit 103. Also, the storage circuit 102 has a function as a work memory for temporarily storing various data used in the processing executed by the processing circuit 103. The storage circuit 102 is realized, for example, by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a hard disk, an optical disk, etc.
[0018] The processing circuit 103 controls the operations of the above-described respective parts of the information processing apparatus 100. For example, the processing circuit 103 performs various processes in response to an instruction received from a user via the input interface 104 connected to the information processing apparatus 100. Alternatively, for example, the processing circuit 103 may perform various processes in response to an instruction received from a user via the communication interface 101. Alternatively, for example, the processing circuit 103 may perform various processes after detecting that image data has been stored in the storage circuit 102. The processing circuit 103 is realized, for example, by a CPU (Central Processing Unit).
[0019] The processing circuit 103 includes, for example, an image acquisition circuit 103a that realizes an image acquisition unit for acquiring an image to be analyzed, a restoration circuit 103b that realizes an acquisition unit for acquiring a restored image of the acquired image, and an analysis circuit 103c that realizes an analysis unit for performing analysis using the restored image.
[0020] Here, for example, each processing function realized by the image acquisition circuit 103a, the restoration circuit 103b, and the analysis circuit 103c, which are components of the processing circuit 103 shown in FIG. 1, is stored in the storage circuit 102 in the form of a program executable by a computer. The processing circuit 103 reads each program from the storage circuit 102 and executes each read program to realize the function corresponding to each program. That is, the processing circuit 103 that has read each program has the functions realized by the image acquisition circuit 103a, the restoration circuit 103b, and the analysis circuit 103c. As a result, the image acquisition circuit 103a functions as an image acquisition unit for acquiring an image (first image) to be analyzed. Further, the restoration circuit 103b functions as an image restoration unit (image addition unit) that inputs an image (first image) into a machine learning model (restoration model) and acquires a restored image (second image) in which an image is added to a peripheral region of the image that is not depicted in the image (first image). Further, the analysis circuit 103c functions as an analysis unit that inputs the restored image (second image) into a machine learning model (analysis model) and acquires an analysis result of the restored image (second image). Hereinafter, adding a plausible image to a peripheral region that was not originally depicted as an image is referred to as restoration.
[0021] The input interface 104 receives input operations of various instructions and various information from the user of the information processing apparatus 100. Specifically, the input interface 104 is connected to the processing circuit 103, converts the input operation received from the user into an electrical signal, and transmits it to the processing circuit 103. For example, the input interface 104 is realized by a trackball, a switch button, a mouse, a keyboard, a touch pad that performs an input operation by touching an operation surface. Alternatively, the input interface 104 may be realized by a touch screen in which a display screen and a touch pad are integrated, a non-contact input interface using an optical sensor, a voice input interface, and the like. Note that the input interface 104 is not limited to those having physical operation components such as a mouse and a keyboard. For example, a processing circuit for receiving an electrical signal corresponding to an input operation from an external input device provided separately from the information processing apparatus 100 and transmitting this electrical signal to the processing circuit 103 is also included in the example of the input interface 104.
[0022] The display 105 displays various data such as image data processed by the information processing apparatus 100 and data based on analysis results. Specifically, the display 105 is connected to the processing circuit 103 and displays various data received from the processing circuit 103. For example, the dis play 105 displays a medical image based on the image data. For example, the display 105 is realized by a liquid crystal monitor, a CRT (Cathode Ray Tube) monitor, a touch panel, or the like.
[0023] The above is an explanation of an example of the configuration of the information processing apparatus 100 according to the present embodiment. In the present embodiment, the information processing apparatus 100 executes various processes described below so as to reduce the influence of performance degradation of a convolutional neural network (CNN) in the edge portion of an image to be analyzed. Hereinafter, examples of various processes for image data executed by the information processing apparatus 100 will be described. In the following description, it is assumed that the information processing apparatus 100 executes image classification processing of two-dimensional image data as an example of analysis processing. However, the information processing apparatus 100 can also be configured to perform similar processing on multi-dimensional image data such as three-dimensional image data.
[0024] The image classification process executed by the information processing apparatus 100 is a process of classifying an image according to whether an object of interest (object to be analyzed) is depicted in the image data. Specifically, there is a binary classification process in which when an object to be analyzed is depicted in the image data, it is classified as positive, and when it is not depicted, the image is classified as negative. In addition, there is a multi-class classification process in which when two or more types of objects are depicted in the image data, the image is classified according to the types. In the description of the present embodiment, for the sake of clarity, the case of executing the binary classification process will be described as an example, but the following description may be read as if the multi-class classification process is executed.
[0025] Hereinafter, in the process executed by the information processing apparatus 100 according to the present embodiment, the restoration model realized by the restoration circuit 103b and the analysis model realized by the analysis circuit 103c will be described.
[0026] The restoration model realized by the restoration circuit 103b according to this embodiment is an image processing model that applies machine learning technology to restore an image by adding (drawing) an image to a peripheral area not depicted in the input image. The restoration model realized by the restoration circuit 103b is the first machine learning model in this embodiment. Specifically, as shown in FIGS. 3A and 3B, the restoration model restores the input image Im100 (the first image) by adding (drawing) an image to at least a part of the peripheral area of the image, and outputs a restored image Im101 (the second image). At this time, as shown in FIG. 3C, for the area Re100 corresponding to the entire area of the image Im100, the restored peripheral area is the part indicated by the area Re101.
[0027] The image of the restored peripheral area Re101 is preferably drawn in a manner that is likely and continuous with the manner depicted in the image data input to the restoration model. Specifically, taking the restored peripheral area (corresponding to the area Re101) in the restored image Im101 corresponding to the image Im100 as an example. The restored image Im101 is preferably an image in which the manner such as the figure, pattern, and contour related to the background depicted in the image Im100 is also likely and continuously drawn in the peripheral area. For example, as shown in FIG. 3D, assume an image Im102 obtained by adding a partial image filled with a predetermined pixel value as a peripheral area to the image Im100. Compared with the image Im102, in the restored image Im101, the manner depicted in the image Im100 is likely and continuously drawn in the peripheral area of the image Im100, which is preferable.
[0028] Also, regarding the process of restoring the image of the peripheral area by continuously drawing with the edge part of the image to be analyzed, an example of the case where image data other than the image Im100 is input to the restoration model will be described with reference to FIGS. 4A to 4D. As shown in FIG. 4A, the information processing apparatus 100 is OC Assume a case where restored image data is generated from an image Im103 that captures the fundus retinal layer acquired by a T imaging device. When the image Im103 is input into a restoration model realized by a restoration circuit 103b, a restored image Im104 is output. As shown in FIG. 4B, in the restored image Im104, in the peripheral region of the image Im103, the continuation of the fundus retinal layer that was not depicted in the image Im103 is plausibly and continuously depicted.
[0029] As an image processing algorithm included in a restoration model for outputting restored image data, there is a technique related to inpainting processing using machine learning disclosed in Non-Patent Document 2. Inpainting processing is a process of restoring a damaged image in which a part of the image is filled in by replacing the filled-in part with a plausible partial image. For example, when inpainting processing is executed on an image in which a part of an old photo is damaged, the damaged part can be replaced with a plausible image and restored to an image without damage. In the present embodiment, it is assumed that inpainting processing is executed on the image Im100 shown in FIG. 3A as an example. At this time, an image Im102 is generated by adding a partial image filled with a predetermined pixel value corresponding to the filled-in part in the inpainting processing to the periphery of the image Im100. Then, the generated image Im102 is input into a restoration model that executes inpainting processing.
[0030] At this time, assume that the machine learning model that executes the inpainting processing included in the restoration model is trained with a data set composed of an image of the same type as or simulating the image assumed to be input into the restoration model. For example, in the present embodiment, a certain original image data is used as correct answer data, and an image data created by filling in the periphery of the original image data with a predetermined pixel value is used as input to train a machine learning model that executes inpainting processing.
[0031] Even if the image data input to the restoration model is originally negative, if a mode in which the restored peripheral region is classified as positive is depicted, there is a possibility that the image data is erroneously classified as positive and the analysis result becomes incorrect. Therefore, the restoration model may be induced and trained to depict only a negative-like mode in the peripheral region. For example, the images constituting the dataset used for training may be composed only of negative images, or a penalty may be given when the restoration model outputs an image in which a positive mode is depicted, and the restoration model may be trained not to depict a positive mode. In this way, by learning the restoration model to restore the peripheral region so that a mode that causes the analysis result to be incorrect is not depicted in the peripheral region, it is possible to suppress the depiction of a positive mode in the peripheral region restored by the restoration model.
[0032] When the image data restored by the restoration model is multi-dimensional image data of three or more dimensions, the peripheral region to be restored is determined according to the number of dimensions of the image data. For example, assume a case where three-dimensional image data is the analysis target. In this case, compared with the case where two-dimensional image data is the analysis target, an image in which the peripheral region corresponding to the front side or the back side of the image data is restored is output from the restoration model in consideration of the dimension in the depth direction.
[0033] Also, the restoration model does not necessarily have to be implemented in the information processing apparatus 100. For example, the restoration circuit 103b may be implemented in a restoration apparatus (not shown). In this case, the information processing apparatus 100 may be connected to the restoration apparatus via the network 200, and transmit the image data including the acquired image to be analyzed to the restoration apparatus to realize the process of step S102 described later. Alternatively, for example, assume a case where the information processing apparatus 100 is connected to a data management apparatus (not shown) that stores configuration data and parameter data related to the restoration model via the network 200. At this time, the processing circuit 103 may acquire the above data from the data management apparatus, reproduce the restoration model in the information processing apparatus 100, and realize the process of step S102 described later.
[0034] In addition, the analysis model realized by the analysis circuit 103c according to the present embodiment is a classifier that applies machine learning technology to perform classification processing using a CNN, and is the second machine learning model in the present embodiment. Further, in at least one or more convolutional processes, the analysis model performs padding processing before performing the convolutional process so that the spatial sizes (width, height) of the input tensor and the output tensor of the convolutional process do not change. As an image processing algorithm for performing classification processing included in the analysis model, specifically, there is a technique related to a classifier using machine learning disclosed in Non-Patent Document 1.
[0035] The analysis model having a classifier using machine learning for the binary classification process in the present embodiment defines, for example, the number of output nodes of the fully connected layer, which is the final layer in the network model of the classifier, as one or two. When the number of output nodes is defined as one, the analysis model normalizes the output value to a value range of 0 or more and 1 or less by a sigmoid function, and classifies it as negative if the normalized value is less than a preset threshold value, and positive if it is equal to or more than the threshold value. At this time, as the threshold value, 0.5, which is the median value of the value range after normalization, can be set, or a threshold value at which the verification data in the dataset used to train the classifier can be classified with high accuracy can be set. Further, when the number of output nodes is defined as two, the analysis model determines that it is negative if the first output value is greater than the second output value, and positive if the second output value is equal to or greater than the first output value.
[0036] Note that the analysis model is assumed to be trained using a dataset composed of an image group of the same type as the images assumed to be input to the analysis model or an image group simulating the same. Further, the analysis model does not necessarily have to be implemented in the information processing apparatus 100. For example, the analysis circuit 103c may be implemented in an analysis apparatus (not shown). In this case, the information processing apparatus 100 may be connected to the analysis apparatus via the network 200, and transmit the image data including the acquired restored image to the analysis apparatus to realize the process of step S103 described later. Alternatively, for example, assume a case where the information processing apparatus 100 is connected to a data management apparatus (not shown) that stores configuration data and parameter data related to the analysis model via the network 200. At this time, the processing circuit 103 may acquire the above data from the data management apparatus, reproduce the analysis model in the information processing apparatus 100, and realize the process of step S103 described later.
[0037] Generally, in a classifier using a CNN, padding processing is often performed before convolution processing so that the spatial sizes of the input tensor and output tensor of the convolution processing do not change. This is one of the reasons because in a network model that requires residual feature extraction disclosed in Non-Patent Document 1, the input tensor and output tensor of a processing block including convolution processing are added together. Further, as another reason, if the spatial size of the tensor changes every time convolution processing is performed, interpolation processing that adversely affects accuracy becomes necessary for pooling processing or tensor combination processing, etc.
[0038] Also, generally, fixed values such as 0 are often set for the pixel values of the pixels in the padding area in the padding process. This will be specifically described with reference to FIGS. 5A, 5B, and 6. As shown in FIG. 5A, it is assumed that a convolution process with a kernel size of 3×3 (width×height) is performed on a tensor Te100, which is a feature map showing the feature amounts related to the object to be analyzed included in the restored image. The feature amounts related to the object to be analyzed are values output from each module constituting the network model, and as an example, this value is represented by a tensor.
[0039] Here, a tensor Te101 is generated by performing a padding process of 1×1 (one pixel in the vertical and horizontal directions) on the tensor Te100 so that the spatial size of the tensor Te102 after the convolution process is the same as that of the tensor Te100. Then, the generated tensor Te101 is used as the input of the convolution process. At this time, fixed values are often set for the pixels corresponding to the one-pixel peripheral area of the tensor Te101, and this fixed value is a value set regardless of the aspect of the original tensor Te100.
[0040] In addition, in the edge area of the tensor Te102 affected by this fixed value, feature amounts useful for improving the accuracy may not be calculated as compared with the central area. Furthermore, when multiple convolution processes are performed, the pixels of the tensor affected by the fixed value affect the adjacent pixels by the convolution process. Therefore, by repeating the convolution process, the influence of the fixed value spreads from the edge to the center of the tensor.
[0041] Here, a comparison is made between the tensor Te102 and the tensor Te105 obtained by performing a convolution process on the tensor Te103 having a spatial size larger than that of the tensor Te100. At this time, the tensor Te105 with a larger spatial size is relatively less affected by the fixed value due to the padding process than the tensor Te102. Specifically, the degree of influence of the fixed value due to the padding process on the spatial size can be calculated by "(the number of pixels affected by the fixed value)÷(spatial size)". For example, the degree of influence of the fixed value in the tensor Te102 in FIG. 5A is 8÷9≈89%, and the degree of influence of the fixed value in the tensor Te105 in FIG. 5B is 16÷25 = 64%. Therefore, the larger the spatial size of the tensor, the relatively smaller the degree of influence of the fixed value due to the padding process.
[0042] In the example shown in FIG. 6, the tensor Te110, which is the image to be analyzed, is input into the network model Ne110, which is a feature extractor that repeats the convolution process and the pooling process, to obtain the output tensor Te111. In this example, the central region of the tensor Te111 obtained as the output of the network model Ne110 contains high-quality feature amounts (the dark parts in the figure) that contribute to improving the accuracy of the analysis and are not affected or are relatively less affected by the fixed value due to the padding process. On the other hand, the edge region of the tensor Te111 is a region that has repeatedly been affected by the fixed value due to the padding process and contains low-quality feature amounts (the light parts in the figure) that are less likely to contribute to improving the accuracy of the analysis. Therefore, as one method for obtaining highly accurate analysis results in a machine learning model equipped with a network model that repeats the convolution process and the pooling process, there is a method of utilizing feature amounts that are less affected by the fixed value due to the padding process.
[0043] Note that the size of the peripheral region in the restored image restored by the restoration model can be determined heuristically, for example, so that the accuracy of the analysis model is increased. Also, the size of the peripheral region in the restored image restored by the restoration model can be determined by other methods. For example, the size of the peripheral region may be determined so that the influence of the fixed value set in the padding region in the padding process executed by the network model related to the analysis model does not reach a predetermined range of the tensor output by the network model related to the analysis model. Here, the predetermined range may be, for example, a range including the approximate center coordinates of the tensor, or a range corresponding to the size of the image to be analyzed deformed at the same ratio when the spatial size of the restored image is deformed to the spatial size of the tensor.
[0044] Also, the size of the peripheral region in the restored image restored by the restoration model can be determined according to, for example, the conditions of the computational cost allowed in the execution environment of the present embodiment. Note that the computational cost that is affected and changed by the size of the peripheral region restored by the restoration model is particularly the computational cost of the process in which the restoration model restores the peripheral region and the computational cost of the process in which the analysis model analyzes the restored image data input. Generally, if the peripheral region restored by the restoration model becomes larger, the computational cost of the process in which the restoration model restores the peripheral region and the computational cost of the process in which the analysis model analyzes the restored image data input both become higher, but the accuracy of the analysis model tends to be higher. Conversely, if the peripheral region restored by the restoration model becomes smaller, the computational cost of the process in which the restoration model restores the peripheral region and the computational cost of the process in which the analysis model analyzes the restored image data input both become lower, but the accuracy of the analysis model tends to be lower. That is, the accuracy of the analysis model approaches the accuracy in the analysis model according to the prior art. However, the accuracy of the analysis model is likely to increase. Conversely, if the peripheral region restored by the restoration model becomes smaller, the computational cost of the process in which the restoration model restores the peripheral region and the computational cost of the process in which the analysis model analyzes the restored image data input both become lower, but the accuracy of the analysis model is likely to decrease. That is, the accuracy of the analysis model approaches the accuracy in the analysis model according to the prior art.
[0045] FIG. 2 is a flowchart showing an example of the flow of processing executed by the information processing apparatus 100 according to the first embodiment. In the present embodiment, as an example, when the user operates the input interface 104 to give an instruction to start the processing in a state where the image data to be analyzed is stored in the storage circuit 102, the processing circuit 103 starts the processing shown in FIG. 2. Hereinafter, the description will be made according to the steps of the flowchart shown in FIG. 2.
[0046] In step S101, the image acquisition circuit 103a acquires the image data to be analyzed stored in the storage circuit 102.
[0047] In step S102, the restoration circuit 103b inputs the image data to be analyzed acquired in step S101 into the restoration model, and acquires the restored image data in which the peripheral area of the image to be analyzed is restored.
[0048] In step S103, the analysis circuit 103c inputs the restored image data acquired in step S102 into the analysis model, and acquires the result of the classification process output from the analysis model as the analysis result.
[0049] In step S104, the analysis circuit 103c outputs the analysis result of the restored image data output from the analysis model in step S103 to the display 105. As a result, the analysis result using the image to be analyzed is displayed on the display 105. Note that the analysis result may be output from the information processing apparatus 100 as data to a data management apparatus (not shown) via the network 200 or a communication cable and a communication circuit (not shown). Further, when outputting the analysis result of the restored image data as data, the analysis result may not be output to the display 105.
[0050] As described above, the information processing apparatus 100 according to the present embodiment can generate an analysis result of restored image data obtained by restoring a peripheral region of an image to be analyzed. By restoring the peripheral region of the image to be analyzed, the restored image data becomes a larger image compared to the image data to be analyzed. As a result, in the image data indicating the feature amount of the image data to be analyzed output from the network model included in the analysis model, the degree of influence of the fixed value in the padding process becomes relatively small. As a result, it is possible to suppress a decrease in classification performance at the edge portion of the image caused by the padding process of the CNN. In addition, by using the restored image in which the peripheral region is restored, it is possible to suppress the phenomenon that the aspect of the edge portion of the image data is cut off, and thus it is also expected that the accuracy of the analysis result will be improved.
[0051] Next, modification examples (modification examples 1 and 2) of the processing of the information processing apparatus 100 in the first embodiment described above will be described. In the following description, the same components and processes as those of the information processing apparatus 100 described above will be denoted by the same reference numerals, and detailed description thereof will be omitted.
[0052] (Modification Example 1) In the above embodiment, the restoration circuit 103b of the information processing apparatus 100 generates a restored image using the upper, lower, left, and right regions of the image Im100 to be analyzed as restoration regions, such as the region Re101 illustrated in FIG. 3C. On the other hand, in this modification example, as shown in FIGS. 7C and 7D, only a part of the peripheral region of the image to be analyzed, such as the left and right regions Re103 and Re104 excluding the upper and lower portions of the image Im107 to be analyzed, is restored. Thereby, in this modification example, the data size of the restored image can be made smaller than the restored image generated in the above embodiment. As a result, in this modification example, the data size of the restored image can be made smaller than the restored image generated in the above embodiment.
[0053] According to this modification example, since the size of the restored image data is small, the computational cost can be reduced compared to the analysis process in the above-described embodiment. Further, for example, there may be a case where it is known that the analysis object is not drawn in an area that is not a restoration target (for example, the upper and lower peripheral areas of the image Im107 shown in FIG. 7A). In this case, by not setting a part of the peripheral area as a restoration target, the computational cost associated with the analysis process can be reduced, and it can be expected that a classification result with the same level of accuracy as the analysis process in the above-described embodiment can be obtained.
[0054] (Modification Example 2) In this modification example, the size of the peripheral area restored by the restoration model realized by the restoration circuit 103b is changed. Thereby, the computational cost related to the restoration process of the restoration model and the computational cost related to the analysis process of the analysis model realized by the subsequent analysis circuit 103c can be adjusted. Further, by changing the size of the peripheral area restored by the restoration model, the accuracy of the analysis result output from the analysis model can be adjusted.
[0055] For example, the administrator of the information processing apparatus 100 according to this modification example may store the default value of the size of the peripheral area of the image restored by the restoration model in the storage circuit 102, read it out when executing the process shown in FIG. 2, and adjust the computational cost and the accuracy of the analysis result. Also, for example, the user of the information processing apparatus 100 according to this modification example may operate the input interface 104 to change the size of the peripheral area restored by the restoration model before starting the process shown in FIG. 2.
[0056] According to this modification example, compared with the process of the information processing apparatus 100 according to the above-described embodiment, the balance between the computational cost and the accuracy of the analysis result is adjusted, and a more suitable execution form of the analysis process by the information processing apparatus 100 can be realized.
[0057] (Second Embodiment) Next, an example of the configuration and processing of an information processing device according to the second embodiment will be described. In the following description, the configuration and processing of the information processing device similar to those of the first embodiment will be denoted by the same reference numerals, and detailed description thereof will be omitted.
[0058] The analysis model realized by the analysis circuit 103c of the information processing device 100 according to this embodiment is a classifier that performs classification processing by CNN using machine learning technology, as in the first embodiment. However, the information processing device 100 according to this embodiment is different from the first embodiment in that a feature tensor indicating the feature of an image output from a network model included in the classifier is trimmed. Specifically, a trimming process is performed to remove a predetermined peripheral region of the feature tensor before it is input to a fully connected layer, which is the final layer in the network model included in the classifier, or a pooling layer corresponding to that fully connected layer. This makes it possible to reduce the influence of a fixed value in the padding process performed by the network model.
[0059] First, a general CNN classifier will be described with reference to FIG. 8. In a general CNN classifier, a tensor Te200 corresponding to an image to be classified is input to a network model Ne200, which is a feature extractor. The tensor Te200 is then extracted as a tensor by passing through a convolution layer, a pooling layer, and the like included in the network model Ne200. The tensor Te201 is output from the network model Ne200 as a tensor Te201. The tensor Te201 is transformed into a tensor Te202 that can be input to the fully connected layer Ne202 through the pooling layer Ne201. The tensor Te202 is then subjected to linear regression processing by the fully connected layer Ne202 so that the tensor Te202 outputs a value equal to the number of output nodes defined in the fully connected layer Ne202. At this time, the tensor Te201 includes the influence of a fixed value due to padding processing associated with the convolution layer included in the network model Ne200, and even if the analysis target is depicted in the peripheral portion of the image data, it may not be classified as positive. This may cause a decrease in the performance of the analysis model.
[0060] Next, a network model, which is a classifier included in the analysis model realized by the analysis circuit 103c according to this embodiment, will be described with reference to FIG. 9. In the network model included in the analysis model, a tensor Te203 corresponding to a restored image in which a peripheral region of an image to be classified is restored is input to a network model Ne203 that is a feature extractor. Then, the tensor Te203 passes through a convolutional layer, a pooling layer, etc. included in the network model Ne203 and is output from the network model Ne203 as a tensor Te204. Further, the tensor Te204 is trimmed according to a preset image shape based on an image region affected by a fixed value due to padding processing resulting from the network structure of the network model Ne203. By this trimming process, a tensor Te205 is generated. Furthermore, the tensor Te205 is transformed into a tensor Te206 having a shape that can be input to a fully connected layer Ne206 through a pooling layer Ne205. Then, the tensor Te206 is linearly regression-processed by the fully connected layer Ne206 so that values corresponding to the number of output nodes defined in the fully connected layer Ne206 are output.
[0061] At this time, the region affected by the fixed value in the tensor Te204 can be determined based on the network structure of the network model Ne203. Specifically, the region affected by the fixed value can be determined based on, for example, the padding size in the padding process by the network model Ne203, the kernel size related to the convolution process, and the number of executions of the convolution process. Instead of or in addition to this, the region affected by the fixed value can also be determined based on the processing content and order of processes that spread the padding region by the padding process, such as the pooling process and the tensor scaling process. That is, an edge region affected by the fixed value due to the padding process is removed from the tensor, and an image having a feature amount calculated from the pixel values of the trimmed image is input to the pooling layer Ne205.
[0062] Note that depending on the spatial size of the tensor Te203 input to the network model Ne203 and the number of convolution processes and pooling processes of the network model Ne203, the fixed value of the padding process affects all pixels of the tensor Te205. In this case, for example, when training the classifier included in the analysis model, an area other than the area corresponding to a predetermined spatial size determined heuristically so as to achieve the highest accuracy may be trimmed as an edge area. Alternatively, for example, for each pixel constituting the tensor Te205, the number of times the fixed value by the padding process affects may be calculated, and an area that receives the influence of the fixed value a predetermined number of times or more may be trimmed as an edge area. Alternatively, for example, by increasing the size of the peripheral area restored by the restoration model realized by the restoration circuit 103b according to the present embodiment, it is possible to avoid a situation where all pixels of the tensor Te205 are affected by the fixed value of the padding process.
[0063] An example of the flow of the process executed by the information processing apparatus 100 according to the present embodiment is the same as that of the first embodiment, and thus detailed description thereof is omitted here. As described above, the information processing apparatus 100 according to the present embodiment can acquire an analysis result for an image to be analyzed. In particular, for the feature amount tensor calculated by the network model included in the analysis model, padding It is characterized in that trimming processing is executed based on the area affected by the fixed value by the processing. As a result, compared with the first embodiment, the area affected by the fixed value by the padding process is removed as the edge area, so that the performance degradation of the analysis process in the edge area of the image due to the padding process of the CNN can be further suppressed.
[0064] Next, modification examples (modification examples 3, 4, and 5) of the process of the information processing apparatus 100 in the second embodiment described above will be described. In the following description, the same components and processes as those of the information processing apparatus 100 described above are denoted by the same reference numerals, and detailed description thereof is omitted.
[0065] (Modification Example 3) In this modification example, for the trimming process of the tensor executed in the network model included in the analysis model realized by the analysis circuit 103c, the size of the edge region of the tensor to be trimmed is changed. As a result, without completely removing the influence of the fixed value due to the padding process associated with the convolutional layer of the network model, the calculation cost and accuracy of the network model included in the analysis model can be adjusted.
[0066] According to this modification example, in the information processing apparatus 100, compared with the information processing apparatus 100 according to the second embodiment described above, the balance between the calculation cost and the accuracy of the analysis result is adjusted, and an execution form suitable for the execution environment can be achieved.
[0067] (Modification Example 4) In this modification example, the acquisition of the restored image data by the restoration model realized by the restoration circuit 103b is skipped, and the image data to be analyzed is directly input to the analysis model realized by the analysis circuit 103c.
[0068] A series of processes executed by the information processing apparatus 100 according to this modification example are different from the processes described in the second embodiment above. That is, the process of step S102 is omitted, and when the process of step S101 is completed, the information processing apparatus 100 proceeds to step S103. In step S103, the analysis circuit 103c inputs the image data to be analyzed acquired in step S101 instead of the restored image data to the analysis model, and acquires the result of the classification process by the analysis model as the analysis result.
[0069] According to this modification example, the information processing apparatus 100 does not execute the output process of the restored image by the restoration model, and an image to be analyzed that is smaller in size than the restored image is input to the analysis model. Thereby, compared with the process executed by the information processing apparatus 100 according to the second embodiment described above, the calculation cost can be reduced. Further, according to this modification example, by trimming the tensor output from the network model included in the analysis model, the influence of the fixed value due to the padding process associated with the convolutional layer can be further reduced.
[0070] (Modification Example 5) In this modification example, as the image processing algorithm included in the restoration model realized by the restoration circuit 103b, for example, a rule-based image processing algorithm may be adopted. Examples of the rule-based image processing algorithm include algorithms such as known reflection padding processing, symmetric padding processing, and replication padding processing.
[0071] Specifically, in the case of the image Im103 shown in FIG. 4A, the restored image Im105 can be generated by reflection padding processing, or the restored image Im106 can be generated by replication padding processing.
[0072] Note that, compared with the restored image Im104, which is an output example of the above restoration model, the restored image data restores the image in the peripheral region such that the fundus retinal layer depicted in the image Im103 bends suddenly as shown in the figure. In this way, the image in the peripheral region is not plausibly and continuously depicted, resulting in an OCT image in an unnatural form. For this reason, there is a concern that the accuracy of the classification process by the analysis model may decrease, but depending on the characteristics of the analysis model, this accuracy decrease may be minor or may not occur.
[0073] For example, assume that the analysis model is trained to classify the presence or absence of edema in the fundus retinal layer, and an OCT image of the fundus retinal layer such as the image Im103 shown in FIG. 4A is input to the restoration model. In this case, if no edema is depicted in the OCT image, no new edema pattern will be depicted in the peripheral region by the above rule-based image processing algorithm. Therefore, even if the restored image of the peripheral region is input to the analysis model, the analysis model will not output a classification result indicating the presence of edema. On the other hand, if edema is depicted in the OCT image, there is a possibility that an edema pattern will be depicted in the peripheral region by the above rule-based image processing algorithm. However, even if the restored image of the peripheral region is input to the analysis model in this way, the analysis model will output a classification result indicating the presence of edema. Therefore, there is no problem because the classification results by the analysis model are the same between the OCT image and the restored image.
[0074] However, due to a portion depicted as if the fundus retinal layer is bent in the peripheral region of the restored image, an analysis model with insufficient accuracy may output an incorrect classification result. Therefore, in order to reduce the possibility of outputting an incorrect classification result, the analysis model may be trained using the image data generated by the above image processing algorithm. That is, a dataset may be constructed and the analysis model may be trained so that a correct classification result is output for image data that is depicted as if the fundus retinal layer is bent in the peripheral region by the image processing algorithm and is not plausibly continuous.
[0075] Thereby, compared with the processing executed by the information processing apparatus 100 according to the second embodiment described above, the image processing algorithm related to the processing of the restoration model can be changed to adjust the calculation cost, or the labor of preparing the restoration model by machine learning can be omitted.
[0076] (Third Embodiment) Next, an example of the configuration and processing of the information processing apparatus according to the second embodiment will be described. In the following description, the same reference numerals are given to the configurations, processing, etc. of the information processing apparatus similar to those in the first embodiment, and detailed descriptions thereof are omitted.
[0077] An example of the image generation process of two-dimensional image data executed by the information processing apparatus 100 according to the present embodiment will be described. Note that the same processing as in the present embodiment can be easily extended to multi-dimensional image data such as three-dimensional image data, and may be readjusted as necessary.
[0078] The image generation process executed by the information processing apparatus 100 in the present embodiment is a process of generating a new image to be analyzed based on the image acquired in step S101. As a specific example, the image generation process is a segmentation process of regionally differentiating by coloring different colors the regions of analysis objects such as figures and lesions drawn in the image data. Further, the image generation process is, for example, a domain conversion process of converting the image data captured in a certain modality into an image in a manner as if it were captured in another different modality. Further, the image generation process is, for example, a high-quality processing of removing noise and artifacts drawn in the image data. In the present embodiment, in each of the above various image generation processes, there may be a decrease in the accuracy of the classification process by the analysis model in the edge region of the generated image due to the padding process of the CNN.
[0079] The analysis model realized by the analysis circuit 103c according to the present embodiment is an image generator that executes an image generation process by a CNN applying machine learning technology. This analysis model performs a padding process before performing the convolution process so that the spatial sizes (width, height) of the input tensor and the output tensor do not change in at least one or more convolution processes. Note that, as an image processing algorithm for performing the image generation process included in the analysis model, a technique related to an image generator using machine learning may be used.
[0080] In addition, when the image generation process of the analysis model is a segmentation process, the analysis model processes the image or tensor output by the CNN according to the conditions of the segmentation process. Specifically, for example, assume a case where a binary segmentation process is performed with the figures or lesions depicted in the input image to the analysis model as the foreground and the other parts as the background. Also, at this time, assume that the output data, which is the image or tensor output by the CNN related to the analysis model, is data with information of one channel. In this case, for example, the pixel value of each pixel in the output image is normalized to a value range of 0 or more and 1 or less by the Sigmoid function. If the normalized value is less than a preset threshold, it can be processed as a background pixel, and if it is equal to or greater than the threshold, it can be processed as a foreground pixel. In addition to this, 0.5, which is the median value of the value range after normalization to the threshold, may be set, or the threshold at which the verification data in the dataset used to train the image generator can generate images with the highest accuracy may be set. Alternatively, assume that the output data, which is the image or tensor output by the CNN related to the analysis model, is data with information of two channels. In this case, for example, if the output value of the first channel is larger than the output value of the second channel, it can be considered negative, and if the output value of the second channel is equal to or greater than the output value of the first channel, it can be considered positive.
[0081] In addition, when the image generation process of the analysis model is a high-quality processing, the analysis model can normalize the pixel values of the image or tensor output by the CNN using the number of channels and the pixel value range of the image input to the analysis model.
[0082] Note that it is assumed that the analysis model is trained using a dataset composed of an image group of the same type as the input image or an image group imitating the input image. Furthermore, when the analysis model is trained by supervised learning, correct data that the analysis model is expected to output, corresponding to each of the images constituting the dataset, may be included in the dataset.
[0083] The processing executed by the information processing apparatus 100 according to this embodiment is the same as that of the first embodiment. Hereinafter, the processing executed by the information processing apparatus 100 according to this embodiment will be described according to the steps of the flowchart shown in FIG. 2. For steps not particularly described, they are the same as those of the first embodiment.
[0084] In step S103, the analysis circuit 103c inputs the restored image data into the analysis model and obtains the image generation result by the analysis model as the analysis result. Thereby, the information processing apparatus 100 according to this embodiment can obtain the analysis result of the restored image in which the peripheral region of the image to be analyzed is restored. By using the restored image in which the peripheral region of the image to be analyzed is restored, the restored image becomes a larger image compared to the image to be analyzed. Therefore, an image larger than the image to be analyzed can be obtained as the analysis result by the analysis model. In addition, in a region where there is a concern about problems such as a decrease in the accuracy of the analysis process due to the influence of fixed values in the padding process in the network model included in the analysis model (that is, a region corresponding to the edge region of the image to be analyzed), the accuracy of the analysis process is improved.
[0085] Next, modification examples (modification examples 6 and 7) of the processing of the information processing apparatus 100 in the above-described third embodiment will be described. In the following description, the same configurations and processes as those of the above-described information processing apparatus 100 are denoted by the same reference numerals and detailed descriptions thereof are omitted.
[0086] (Modification Example 6) In this modification example, the analysis model realized by the analysis circuit 103c performs a trimming process on the image obtained by the above-described image generation process to obtain an image of a smaller size. As a result, the analysis model can perform a classification process based on, for example, an image obtained by trimming the image generation result to the same size as the image to be analyzed. That is, a predetermined edge region removed by the analysis model coincides with a restored peripheral region in the restored image. By performing the classification process using the image subjected to the trimming process in this way, the correspondence relationship between pixel coordinates becomes clearer between the image to be analyzed and the image of the analysis result.
[0087] (Modification Example 7) In this modification example, the size of the peripheral region restored by the restoration model realized by the restoration circuit 103b is adjusted. As a result, it is possible to adjust the calculation cost related to the process of the restoration model and the calculation cost of the process of the analysis model realized by the subsequent analysis circuit 103c, or to adjust the accuracy of the analysis result and the image size by the analysis model.
[0088] Therefore, the information processing apparatus 100 according to this modification example can more suitably adjust the balance between the calculation cost and the accuracy of the analysis result as compared with the information processing apparatus 100 according to the above-described third embodiment, and can perform an analysis suitable for the execution environment.
[0089] (Fourth Embodiment) Next, an example of the configuration and process of the information processing apparatus according to the second embodiment will be described. In the following description, the same reference numerals are given to the configuration and process of the information processing apparatus similar to those in the first embodiment, and detailed description thereof will be omitted.
[0090] An example of the image generation process of two-dimensional image data executed by the information processing apparatus 100 according to this embodiment will be described. Note that the same process as in this embodiment can be easily extended to multi-dimensional image data such as three-dimensional image data, and may be readjusted as necessary.
[0091] The object detection process in this embodiment is a process of identifying the location of the object to be analyzed depicted in the image. In the object detection process, each object of the object to be analyzed identified in the image is surrounded by a bounding box, an identifier for identifying the type of the object is added, or the area occupied by each object is masked and imaged. The masked imaging is also called instance segmentation processing. As a result, indicators indicating objects such as bounding boxes, identifiers, and mask images are included in the image of the analysis result, making it easier for the user to identify the position of the object. However, also in the object detection process, as in the classification process, there is a possibility that the analysis accuracy may decrease at the edge portion of the image to be analyzed due to the padding process of the CNN.
[0092] The analysis model realized by the analysis circuit 103c according to this embodiment is an object detector that executes object detection processing by CNN applying machine learning technology. This analysis model performs padding processing before performing the convolution process so that the spatial sizes (width, height) of the input tensor and the output tensor of the convolution process do not change in at least one or more convolution processes. Note that examples of the image processing algorithm for performing object detection processing provided in the analysis model include technologies related to object detectors using machine learning.
[0093] The processing executed by the information processing apparatus 100 according to this embodiment is the same as that in the first embodiment. Hereinafter, the processing executed by the information processing apparatus 100 according to this embodiment will be described according to the steps of the flowchart shown in FIG. 2. Note that steps not particularly described are the same as those in the first embodiment.
[0094] In step S103, the analysis circuit 103c inputs the restored image data into the analysis model, and acquires, as an analysis result, the result of the above-described object detection process by the analysis model. Thereby, the information processing apparatus 100 according to the present embodiment can acquire an analysis result of a restored image obtained by restoring a peripheral region of the image to be analyzed. By using the restored image in which the peripheral region of the image to be analyzed is restored, the restored image becomes a larger image compared to the image to be analyzed. Therefore, an object detection result based on an image in a wider range than the image to be analyzed can be obtained as an analysis result by the analysis model. In addition, in a region where there is a concern about problems such as a decrease in the accuracy of the analysis process due to the influence of fixed values by the padding process in the network model included in the analysis model (that is, a region corresponding to the edge region of the image to be analyzed), the accuracy of the analysis process is improved.
[0095] Next, a modification example (modification example 8) of the process of the information processing apparatus 100 in the above-described fourth embodiment will be described. In the following description, the same components and processes as those of the information processing apparatus 100 described above are denoted by the same reference numerals, and detailed descriptions thereof are omitted.
[0096] (Modification example 8) In this modification example, the analysis model realized by the analysis circuit 103c may perform trimming processing on the image obtained as the object detection result, and output an image with a smaller image size as the analysis result.
[0097] Specifically, the analysis model changes the trimming process according to the content of the object detection process. As a result, the analysis model performs analysis on an image modified so that the index does not overlap with a predetermined edge area. For example, assume that the object detection process is a process of surrounding an object in an image with a bounding box. In this case, the analysis model reduces the size of the bounding box so that it fits within the image for a bounding box that has a part outside the image obtained by the trimming process. Also, when the area or volume of the bounding box is below a predetermined size, the analysis model may remove the corresponding bounding box. Further, for example, assume that the object detection process is a process of adding an identifier that identifies the type of object drawn in the image. In this case, when the object corresponding to the identifier does not fit within the image obtained by the trimming process, the analysis model may remove the identifier. Further, for example, assume that the object detection process is a process of adding a mask image to an object drawn in the image by instance segmentation processing. In this case, the analysis model may trim a mask image that has a part outside the image obtained by the trimming process. Note that when the area or volume of the mask image is below a predetermined size, the analysis model may remove the corresponding mask image.
[0098] Examples of the above trimming process are shown in FIGS. 10A to 10F. FIG. 10A shows an image Im301 before the trimming process. In the image Im301, a bounding box 310 that surrounds an object drawn in the image is drawn. For convenience of explanation, although not shown in the image Im301, the boundary 320 of the edge area trimmed by the trimming process in the image Im301 is shown by a dotted line. The area outside the boundary 320 is an example of the edge area. By the above trimming process on the image Im301, an image Im302 shown in FIG. 10B is generated. As shown in the figure, in the image Im302, a part of the object drawn in the image Im301 is drawn. In this case, in the image Im302, the bounding box 310 in the image Im301 is modified to a bounding box 311 that surrounds the object.
[0099] In addition, FIG. 10C shows the image Im303 before the trimming process. In the image Im303, two objects 312 and 313 are depicted. In addition, identifiers "P001" and "P002" are added to the objects 312 and 313, respectively. For convenience of explanation, although not shown in the image Im303, the boundary 320 of the edge region trimmed by the trimming process in the image Im303 is shown by a dotted line. The region outside the boundary 320 is an example of the edge region. By the above-described trimming process for the image Im303, the image Im304 shown in FIG. 10D is generated. As shown in the figure, in the image Im304, the object 312 does not fit within the image, and only a part of it is depicted. In this case, in the image Im304, the identifier "P001" corresponding to the object 312 that does not fit within the image is removed.
[0100] In addition, FIG. 10E shows the image Im305 before the trimming process. A mask image 314 covering the object is applied to the image Im305. For convenience of explanation, although not shown in the image Im305, the boundary 320 of the edge region trimmed by the trimming process in the image Im305 is shown by a dotted line. The region outside the boundary 320 is an example of the edge region. By the above-described trimming process for the image Im305, the image Im306 shown in FIG. 10E is generated. As shown in the figure, in the image Im306, a part of the object depicted in the image Im305 is depicted. In this case, in the image Im306, the mask image 314 in the image Im305 is corrected to a mask image 315 covering the object.
[0101] In addition, in the above-described trimming process, the trimming process may be adjusted so that the size of the image obtained as the object detection result is the same as the size of the image to be analyzed. Thereby, the correspondence relationship between the pixel coordinates becomes clearer between the image to be analyzed and the image of the analysis result.
[0102] (Other Embodiments) The processing of the information processing apparatus 100 may be realized in various different forms, not limited to the above-described embodiments and modifications. For example, the information processing apparatus 100 may execute the same processing as that for executing various processes on the above-described two-dimensional image data on multi-dimensional image data.
[0103] In addition, each component of each illustrated apparatus is a functional concept, and it is not necessarily physically configured as illustrated. That is, the specific form of the distribution and integration of each apparatus is not limited to that illustrated, and all or a part of it can be functionally or physically distributed and integrated in any unit according to various loads, usage situations, and the like. Furthermore, each processing function performed by each apparatus can be realized in whole or in any part by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware by wired logic.
[0104] Also, among the processes described in the embodiments, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be arbitrarily changed unless otherwise specified.
[0105] Also, the method described in the embodiments can be realized by executing a program prepared in advance on a computer such as a personal computer or a workstation. This program can be distributed via a network such as the Internet. In addition, this control program is recorded on a non-transitory recording medium readable by a computer such as a hard disk, a floppy disk (FD), a CD-ROM, a magneto-optical disk (MO), or a DVD, and can be read from the recording medium by the computer and executed.
[0106] Although some embodiments have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, changes, and combinations of embodiments can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, as well as in the invention described in the claims and the equivalent scope thereof.
[0107] Also, the disclosed technology can be implemented, for example, as an embodiment such as a system, device, method, program, or recording medium (storage medium). Specifically, it may be applied to a system composed of a plurality of devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or it may also be applied to a device consisting of one device.
[0108] The present invention supplies a program that causes a computer to execute each step of the image processing method according to the above-described embodiment to a system or device via a network or a storage medium. Such a program is configured to be read and executed by one or more processors in the computer of the system or device. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0109] The disclosure of this embodiment includes the following configurations, methods, and programs. (Configuration 1) Image acquisition means for acquiring a first image to be analyzed, Image addition means for inputting the first image into a first machine learning model and acquiring a second image in which an image is added to a peripheral area of the first image that is not depicted in the first image, Analysis means for inputting the second image into a second machine learning model and acquiring an analysis result of the second image, An information processing apparatus comprising the above. (Configuration 2) The information processing apparatus according to Configuration 1, wherein the first machine learning model adds an image to a part of the peripheral region of the first image. (Configuration 3) The second machine learning model generates a feature map indicating a feature amount related to an analysis target object included in the second image, removes a predetermined edge region of the feature map, and outputs the analysis result using the feature map from which the predetermined edge region has been removed. The information processing apparatus according to Configuration 1 or 2, characterized in that. (Configuration 4) The information processing apparatus according to Configuration 3, wherein the second machine learning model has a network model that outputs a feature map indicating the feature amount related to the analysis target object included in the second image. (Configuration 5) The information processing apparatus according to Configuration 4, wherein the second machine learning model inputs the feature map from which the predetermined edge region has been removed into a fully connected layer or a pooling layer corresponding to the fully connected layer to output the analysis result. (Configuration 6) The information processing apparatus according to any one of Configurations 1 to 5, wherein the second machine learning model removes a predetermined edge region corresponding to the first image acquired by the image acquisition unit from the image indicating the analysis result. (Configuration 7) The predetermined edge region removed by the second machine learning model coincides with the peripheral region to which an image is added in the second image. Information processing apparatus. (Configuration 8) When the analysis result includes an index indicating an object located in a predetermined edge region of the second image, the second machine learning model corrects the index so as not to overlap with the predetermined edge region and outputs the analysis result. (Configuration 9) The information processing apparatus according to configuration 8, wherein the predetermined edge region of the second image coincides with the peripheral region where the image is added in the second image. (Configuration 10) The information processing apparatus according to configuration 8 or 9, wherein the index is a bounding box surrounding the object, an identifier of the object, or a mask image overlapping the object. (Configuration 11) The information processing apparatus according to any one of configurations 1 to 10, wherein the first machine learning model is trained to add an image to the peripheral region so that an aspect that causes the analysis result to be incorrect is not depicted in the peripheral region. (Configuration 12) Image acquisition means for acquiring an image to be analyzed; Analysis means for inputting the image into a machine learning model and acquiring an analysis result of the image; comprising The machine learning model generates a feature map showing feature amounts related to an object to be analyzed included in the image, removes a predetermined edge region of the feature map, and outputs the analysis result using the feature map from which the predetermined edge region has been removed. The information processing apparatus is characterized in that. (Configuration 13) Image acquisition means for acquiring a first image to be analyzed; Image addition means for acquiring a second image in which an image is added to a peripheral region of the first image that is not depicted in the first image; Analysis means for inputting the second image into a machine learning model and acquiring an analysis result of the second image; comprising The machine learning model generates a feature map showing feature amounts related to an object to be analyzed included in the second image, removes a predetermined edge region of the feature map, and outputs the analysis result using the feature map from which the predetermined edge region has been removed. The information processing apparatus is characterized in that. (Configuration 14) The information processing apparatus according to Configuration 1 or 13, wherein the image added to the peripheral area is an image in which the mode depicted in the first image is plausibly and continuously depicted. (Method 1) An image acquisition step of acquiring a first image to be analyzed; An image addition step of inputting the first image into a first machine learning model to acquire a second image in which an image is added to a peripheral area of the first image that is not depicted in the first image; An analysis step of inputting the second image into a second machine learning model to acquire an analysis result of the second image; An information processing method, characterized by comprising: (Method 2) An image acquisition step of acquiring an image to be analyzed; An analysis step of inputting the image into a machine learning model to acquire an analysis result of the image; Comprising: The machine learning model generates a feature map indicating a feature amount related to an object to be analyzed included in the image, removes a predetermined edge area of the feature map, and outputs the analysis result using the feature map from which the predetermined edge area has been removed. An information processing method, characterized by this. (Method 3) An image acquisition step of acquiring a first image to be analyzed; An image addition step of acquiring a second image in which an image is added to a peripheral area of the first image that is not depicted in the first image; An analysis step of inputting the second image into a machine learning model to acquire an analysis result of the second image; Comprising: The machine learning model generates a feature map indicating a feature amount related to an object to be analyzed included in the second image, removes a predetermined edge area of the feature map, and outputs the analysis result using the feature map from which the predetermined edge area has been removed. An information processing method, characterized by this. (Program) A program for causing a computer to execute each step of the information processing method described in any one of Methods 1 to 3.
Explanation of Signs
[0110] 100 Information processing apparatus, 103 Processing circuit, 103a Image acquisition circuit, 103b Restoration circuit, 103c Analysis circuit
Claims
1. image acquisition means for acquiring a first image of an object to be analyzed; an image adding means for inputting the first image into a first machine learning model to obtain a second image in which an image is added to a peripheral area of the first image that is not depicted in the first image; an analysis means for inputting the second image into a second machine learning model to obtain an analysis result of the second image; An information processing device comprising:
2. The information processing device according to claim 1 , wherein the first machine learning model adds an image to a portion of the surrounding area of the first image.
3. the second machine learning model generates a feature amount map indicating features related to an analysis target included in the second image, removes a predetermined peripheral region of the feature amount map, and outputs the analysis result using the feature amount map from which the predetermined peripheral region has been removed.
3. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
4. The information processing apparatus according to claim 3 , wherein the second machine learning model has a network model that outputs the feature amount map indicating the feature amount related to the analysis target object included in the second image.
5. The information processing device according to claim 4 , wherein the second machine learning model inputs the feature map from which the predetermined peripheral region has been removed to a fully connected layer or a pooling layer corresponding to the fully connected layer, and outputs the analysis result.
6. The information processing device according to claim 1 or 2, characterized in that the second machine learning model removes a predetermined peripheral area corresponding to the first image acquired by the image acquisition means from the image showing the analysis result.
7. The information processing device according to claim 6 , wherein the predetermined peripheral region removed by the second machine learning model coincides with the peripheral region to which an image is added in the second image.
8. 3. The information processing device according to claim 1, wherein when the analysis result includes an index indicating an object located in a predetermined peripheral region of the second image, the second machine learning model corrects the index so that it does not overlap with the predetermined peripheral region and outputs the analysis result.
9. The information processing apparatus according to claim 8 , wherein the predetermined peripheral region of the second image coincides with the peripheral region in the second image to which an image has been added.
10. The information processing apparatus according to claim 8 , wherein the indicator is a bounding box surrounding the object, an identifier of the object, or a mask image overlapping the object.
11. The information processing device according to claim 1 or 2, characterized in that the first machine learning model is trained to add images to the surrounding area so that features that may cause the analysis result to be erroneous are not depicted in the surrounding area.
12. An image acquisition means for acquiring an image of an object to be analyzed; an analysis means for inputting the image into a machine learning model to obtain an analysis result of the image; Equipped with the machine learning model generates a feature amount map indicating features related to an analysis target included in the image, removes a predetermined peripheral region of the feature amount map, and outputs the analysis result using the feature amount map from which the predetermined peripheral region has been removed.
23. An information processing apparatus comprising:
13. image acquisition means for acquiring a first image of an object to be analyzed; an image adding means for acquiring a second image in which an image is added to a peripheral area of the first image that is not depicted in the first image; an analysis means for inputting the second image into a machine learning model to obtain an analysis result of the second image; Equipped with the machine learning model generates a feature amount map indicating features related to an analysis target object included in the second image, removes a predetermined peripheral region of the feature amount map, and outputs the analysis result using the feature amount map from which the predetermined peripheral region has been removed.
23. An information processing apparatus comprising:
14. 14. The information processing apparatus according to claim 1, wherein the image added to the peripheral region is an image in which the aspect depicted in the first image is depicted plausibly and continuously.
15. an image acquisition step of acquiring a first image of an object to be analyzed; an image addition step of inputting the first image into a first machine learning model to obtain a second image in which an image is added to a peripheral area of the first image that is not depicted in the first image; an analysis step of inputting the second image into a second machine learning model to obtain an analysis result of the second image; 13. An information processing method comprising:
16. An image acquisition step of acquiring an image to be analyzed; an analysis step of inputting the image into a machine learning model to obtain an analysis result of the image; having the machine learning model generates a feature amount map indicating features related to an analysis target included in the image, removes a predetermined peripheral region of the feature amount map, and outputs the analysis result using the feature amount map from which the predetermined peripheral region has been removed.
23. An information processing method comprising:
17. an image acquisition step of acquiring a first image of an object to be analyzed; an image addition step of acquiring a second image in which an image is added to a peripheral area of the first image that is not depicted in the first image; an analysis step of inputting the second image into a machine learning model to obtain an analysis result of the second image; having the machine learning model generates a feature amount map indicating features related to an analysis target object included in the second image, removes a predetermined peripheral region of the feature amount map, and outputs the analysis result using the feature amount map from which the predetermined peripheral region has been removed. An information processing method comprising:
18. A computer is provided with a program for executing each step of the information processing method according to any one of claims 15 to 17. A program for executing a task.
Citation Information
Cited By
Information processing device, information processing method, and program
WO2025127032A1