Feature extraction device, feature extraction method, program, and information recording medium
Patent Information
- Application Number
- JP2023523407
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2022-05-12
- Filing Date
- 2022-05-12
- Publication Date
- 2025-05-20
Abstract
Description
Feature extraction device, feature extraction method, program, and information recording medium
[0001] The present invention relates to a feature extraction device, a feature extraction method, a program, and an information recording medium for extracting features of an object from a plurality of images of the object.
[0002] Conventionally, techniques have been proposed that use neural networks to process photographs of objects to extract features, classify the objects based on the features, and use the results for various purposes, including medical diagnosis.
[0003] For example, Patent Document 1 discloses a technology in which an image of a target and one or more attribute parameters associated with the target are received, and when classifying the target using a neural network, each element of a given feature map is convolved with the received one or more attribute parameters.
[0004] On the other hand, when taking photographs of the body parts of a subject to be examined using ultrasound or other methods, multiple images may be obtained for one subject. Also, when multiple organs with different functions are photographed in one photograph, this single photograph may be divided into multiple smaller images so that each part can be processed separately.
[0005] In such cases, it is conceivable that there will be many images of patients with lesions or the like in which areas are indistinguishable from those of healthy individuals.
[0006] On the other hand, in prognosis diagnosis such as predicting the recurrence of prostate cancer, the target area resected from the subject is used as a specimen, and doctors, based on their medical knowledge, use pathological photographs of the specimen to narrow down and enclose the area where the cancer is located (the area where the lesion is located) from other areas (normal areas). For example, in the Gleason classification, which is widely used to classify the malignancy of cancer, after narrowing down the area where the cancer is located, the Gleason score, which indicates the malignancy, is measured by further examining the tissue morphology of the cancer.
[0007] Such narrowing down and confinement requires a great deal of time and effort, and the accuracy varies depending on the doctor. In addition, only appearances that can be recognized using existing medical knowledge can be analyzed.
[0008] Furthermore, when diagnosing whether or not a patient has prostate cancer, if useful information useful for diagnosis can be obtained from ultrasound images, etc., it is expected that the burden of subsequent tests such as biopsies can be reduced.
[0009] Patent No. 6345332
[0010] Therefore, in order to estimate whether or not a subject is suffering from a particular disease, it is necessary to appropriately extract features of the subject from a large number of images of the subject.
[0011] Furthermore, in various applications other than disease incidence estimation, if the features of an object can be appropriately extracted from a large number of images relating to the object, the object can be appropriately classified.
[0012] Therefore, there is a need for a technique that can appropriately extract features of an object from a large number of images of the object, which are useful for classifying the object.
[0013] The present invention is devised to solve the above-mentioned problems, and aims to provide a feature extraction device, a feature extraction method, a program, and an information recording medium for extracting features of an object from a plurality of images relating to the object.
[0014] The feature extraction device according to the present invention comprises: an image processing unit that, when an image is input, calculates the likelihood that the input image belongs to a first image class and feature parameters of the input image using an image model; and a feature processing unit that, when an image group is input, inputs images included in the input image group to the image processing unit to calculate the likelihood and feature parameters, selects a predetermined number of representative images from the input image group based on the calculated likelihood, and outputs the calculated feature parameters for the selected predetermined number of representative images as features of the target.
[0015] According to the present invention, it is possible to provide a feature extraction device, a feature extraction method, a program, and an information recording medium for extracting features of an object from a plurality of images relating to the object.
[0016] FIG. 1 is an explanatory diagram showing a schematic configuration of a feature extraction device according to an embodiment of the present invention. FIG. 2 is a flowchart showing the control flow of a learning process for training an image model. FIG. 3 is a flowchart showing the control flow of a learning process for training a classification model. FIG. 4 is a flowchart showing the control flow of image processing for obtaining feature information from a group of images. FIG. 5 is a flowchart showing the control flow of a feature extraction process. FIG. 6 is a flowchart showing the control flow of a classification process. FIG. 7 is a graph showing the experimental results of classification according to a conventional method. FIG. 8 is a graph showing the experimental results of classification according to the present embodiment. FIG. 9 is an explanatory diagram comparing the graph of the experimental results of classification according to the present embodiment with the graph of the experimental results of classification according to the conventional method by superimposing them.
[0017] The following describes embodiments of the present invention. Note that these embodiments are for illustrative purposes only and do not limit the scope of the present invention. Therefore, those skilled in the art can adopt embodiments in which each or all of the elements of the present embodiments are replaced with equivalents. Furthermore, elements described in each example can be omitted as appropriate depending on the application. In this way, all embodiments constructed in accordance with the principles of the present invention are included in the scope of the present invention.
[0018] (Configuration) The feature extraction device according to this embodiment is typically implemented by a computer running a program. The computer is connected to various output devices and input devices, and transmits and receives information to and from these devices.
[0019] A program executed by a computer can be distributed or sold by a server to which the computer is connected for communication, or it can be recorded on a non-transitory information recording medium such as a CD-ROM (Compact Disk Read Only Memory), flash memory, or EEPROM (Electrically Erasable Programmable ROM), and then the information recording medium can be distributed, sold, etc.
[0020] The program is installed on a non-transitory information recording medium such as a hard disk, solid-state drive, flash memory, EEPROM, etc., possessed by the computer. The information processing device of this embodiment is then realized by the computer. Generally, the computer's central processing unit (CPU) reads the program from the information recording medium into random access memory (RAM) under the control of the computer's operating system (OS), and then interprets and executes the code contained in the program. However, in an architecture in which the information recording medium can be mapped within a memory space accessible by the CPU, explicit loading of the program into RAM may not be necessary. Various pieces of information required during program execution can be temporarily stored in RAM.
[0021] Furthermore, as mentioned above, it is desirable for the computer to be equipped with a GPU (Graphics Processing Unit) for performing various image processing calculations at high speed. By using the GPU and libraries such as TensorFlow and PyTorch, it becomes possible to utilize the learning and classification functions in various artificial intelligence processes under the control of the CPU.
[0022] It should be noted that the information processing device of this embodiment may be configured using a dedicated electronic circuit rather than a general-purpose computer. In this embodiment, the program may be used as a resource for generating wiring diagrams, timing charts, and the like for the electronic circuit. In this embodiment, an electronic circuit that meets the specifications defined in the program is configured using an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the electronic circuit functions as a dedicated device that performs the functions defined in the program, thereby realizing the information processing device of this embodiment.
[0023] For ease of understanding, the following description will be given assuming that the feature extraction device 101 is realized by a computer executing a program. Fig. 1 is an explanatory diagram showing the general configuration of a feature extraction device according to an embodiment of the present invention.
[0024] As shown in the figure, the feature extraction device 101 according to this embodiment includes an image processing unit 111 and a feature processing unit 112. The feature extraction device 101 may also include a classification processing unit 113 as an optional element.
[0025] The image processing unit 111 refers to the image model 151. The feature extraction device 101 may include, as an optional element, an image training unit 131 for training the image model 151. For example, when an image model that has already been trained is used as the image model 151, the image training unit 131 can be omitted.
[0026] The feature processing unit 112 refers to the classification model 153. The feature extraction device 101 may include, as an optional element, a classification training unit 133 for training the classification model 153. For example, when a classification model that has already been trained is used as the classification model 153, the classification training unit 133 can be omitted.
[0027] Furthermore, the image training unit 131 and the classification training unit 133 can also be implemented as devices independent of the feature extraction device 101. In this embodiment, the trained parameters constituting the trained image model 151 and the classification model 153 and the inference program using the trained parameters are transferred from the image training unit 131 and the classification training unit 133 to the feature extraction device 101 via an information recording medium, a computer communication network, or the like. In this application, for ease of understanding, learning the parameters of models such as the image model 151 and the classification model 153 may be expressed as training, learning, or updating the model.
[0028] Now, when an image is input, the image processing unit 111 calculates the likelihood that the input image belongs to the first image class and the feature parameters of the input image using the image model 151. Therefore, when a plurality of images are input to the image processing unit 111 sequentially (or in parallel, or all at once), the likelihood and feature parameters for each image are output from the image processing unit 111 sequentially (or in parallel, or all at once).
[0029] Here, various models can be adopted as the image model 151, such as a model related to a deep convolutional neural network.
[0030] The image processing unit 111 calculates likelihood from a vector of pixel values, which can be considered to reduce the dimension of the vector value. If the image model 151 is related to a neural network or the like, information is exchanged across multiple layers to reduce the dimension. Therefore, information output from intermediate layers can be used as feature parameters. In other words, intermediate vectors during the dimension reduction in the image model 151 can be used as feature parameters.
[0031] Also, most simply, the likelihood of the image can be used as a feature parameter as it is. That is, the likelihood that is the final result of the dimension reduction in the image model 151 can be used as a feature parameter.
[0032] On the other hand, when an image group is input, the feature processing unit 112 outputs feature information of the image group.
[0033] First, the feature processing unit 112 inputs the images included in the input image group to the image processing unit 111, and causes the image processing unit 111 to calculate likelihoods and feature parameters.
[0034] Next, the feature processing unit 112 selects a predetermined number of representative images from the input image group based on the calculated likelihood.
[0035] Then, the feature processing unit 112 outputs the feature parameters calculated for the selected predetermined number of representative images as feature information for the image group.
[0036] The number of representative images selected for one image group can be any number equal to or greater than one.
[0037] For example, when one representative image is selected, the feature information is the feature parameters of the representative image, and when the feature parameters are the likelihood itself, the feature information is a scalar value consisting of the likelihood.
[0038] When three representative images are selected, the feature information is a vector, tensor, or array of the feature parameters of the representative images. When the feature parameters are likelihoods themselves, the feature information is a vector value of the three likelihoods.
[0039] Generally, when the feature parameters for one image are an N-dimensional vector and M representative images are selected, the feature information output from feature processing unit 112 for one image group will be an N×M-dimensional vector.
[0040] The simplest method for selecting representative images based on likelihood is to select a predetermined number of representative images in descending order of likelihood.
[0041] In this method, the feature information emphasizes features that correspond to the first image class among the images.
[0042] The next possible method is to select a predetermined number of representative images in descending order of the absolute value of the difference between the likelihood and a predetermined reference value. For example, if the likelihood is assumed to be the probability that an image belongs to the first image class, the likelihood will be a value between 0 and 1, and the predetermined reference value is a boundary value for determining whether or not an image belongs to the first image class, and can be 0.5.
[0043] In this method, compared to the above method, the feature information emphasizes the contrast between whether or not the image group conforms to the first image class.
[0044] As another method, images with the minimum, median, and maximum likelihoods can be selected as a predetermined number of representative images.
[0045] In this method, compared to the above method, the feature information emphasizes the degree to which the image group is dispersed with respect to the first image class.
[0046] Now, when a group of target images relating to an object is input, the classification processing unit 113 inputs the input group of target images to the feature processing unit 112, and estimates whether or not the object belongs to the first object class using the classification model 153 from the feature information output from the feature processing unit 112.
[0047] Here, if the image group input to the feature processing unit 112 is an image group consisting of images of a single common object, the feature information output from the feature processing unit 112 will represent the relationship between the object and the first image class.
[0048] Therefore, if the first image class and the first object class are set so that the fact that an image included in the target image group relating to the target belongs to the first image class correlates with the target belonging to the first object class as the first image class, and if the feature processing unit 112 selects a representative image so as to emphasize the features of the image group, then the target can be appropriately classified by utilizing the feature information output from the feature processing unit 112.
[0049] At this time, the classification processing unit 113 can be configured to receive additional data related to the object in addition to a group of object images related to the object. In this mode, the classification processing unit 113 inputs the input group of object images to the feature processing unit 112, and estimates whether the object belongs to the first object class using the classification model 153, based on the feature information output from the feature processing unit 112 and the input additional data.
[0050] Here, various models can be adopted as the classification model 153, such as linear regression, logistic regression, ridge regression, lasso regression, or a model related to a support vector machine.
[0051] In addition, the image training unit 131 updates the image model 151 and proceeds with learning using training data consisting of a set of an image and a label indicating whether the image belongs to the first image class.
[0052] In addition, the classification training unit 133 updates the classification model 153 and proceeds with learning using training data consisting of a set of feature information of a group of target images related to the target, additional data related to the target if there is any, and a label indicating whether the target belongs to the first target class.
[0053] An example of applying this embodiment to prostate cancer diagnosis will be described below. First, the subject is a test subject or patient who will be diagnosed with prostate cancer. Therefore, the first subject class is a class that indicates that the subject is (highly likely to be) suffering from prostate cancer.
[0054] Furthermore, as the target image group, a plurality of images taken by ultrasound or a large number of images obtained by dividing a taken photograph into pieces of a predetermined size are adopted.
[0055] Additional data that can be used include the subject's age, PSA (Prostate Specific Antigen) value, TPV (Total Prostate Volume) value, and PSAD (PSA Density) value.
[0056] The simplest first image class is a class that indicates that the subject photographed in the image is suffering from prostate cancer.
[0057] In this case, the training data required to advance the learning of image model 151 will be a large number of pairs of a single image of the subject and a label indicating whether the subject has prostate cancer, i.e., whether the subject belongs to the first subject class.
[0058] In this form of training data, images of the same subject will share the same label.
[0059] In addition, if there is information on patients who were discovered in previous prostate cancer tests, or on subjects who were diagnosed as suspected in ultrasound tests but were followed up by a biopsy in which a portion of tissue was removed and observed under a microscope, and who did not develop the disease, the first image class can be a class that indicates that the Gleason score assigned to the specimen area in the biopsy specimen that corresponds to the image area depicted in the image is above a predetermined value.
[0060] In this case, the training data required to advance the learning of image model 151 is image training data that includes a large number of pairs of an image of the target area and a label indicating whether the Gleason score assigned to the area based on the biopsy specimen is equal to or greater than a predetermined value.
[0061] In this type of training data, even if the images relate to the same subject, different labels may be assigned if the parts captured in the images are different.
[0062] After the image model 151 has been trained, the classification model 153 can be trained. In this application example, the data required to train the classification model 153 includes: feature information obtained by the image model 151 from a group of target images relating to the target (photographic images of the subject taken by ultrasound or the like, or images obtained by dividing the photographed images into predetermined sizes), additional data such as the age, PSA value, TPV value, and PSAD value of the target, if available, and a label indicating the final diagnosis result of whether or not the subject is positive for prostate cancer. Classification training data is prepared that includes a large number of sets of the following:
[0063] (Learning Process for Feature Extraction) Fig. 2 is a flowchart showing the control flow of the learning process for training an image model. The following description will be made with reference to this figure.
[0064] When this process starts, the image training unit 131 first accepts input of image training data (step S201).
[0065] Then, until the training of the image model 151 is completed (step S202; No), the image training unit 131 repeats the following process.
[0066] That is, the image training unit 131 repeats the following process for each of the pairs included in the image training data (step S204).
[0067] The image training unit 131 acquires the images and labels contained in the set (step S205), provides the acquired images as inputs to the neural network associated with the image model 151 (step S206), obtains the results output from the neural network (step S207), and calculates the difference between the output results and the labels (step S208).
[0068] After completing the repetition for each pair (step S209), the image training unit 131 calculates the value of the evaluation function based on the difference found for each pair, updates the image model 151 (step S210), and returns control to step S202.
[0069] When the training of the image model 151 is completed (step S202; Yes), the image training unit 131 ends this process.
[0070] (Classification Learning Process) Figure 3 is a flowchart showing the control flow of the learning process for training a classification model. The following description will be made with reference to this figure.
[0071] When this process starts, the classification training unit 133 first accepts input of classification training data (step S301).
[0072] Then, until the training of the classification model 153 is completed (step S302; No), the image training unit 131 repeats the following process.
[0073] That is, the image training unit 131 repeats the following process for each of the pairs included in the image training data (step S304).
[0074] The image training unit 131 acquires the image group included in the set, the additional data (if any) and the label (step S305), and provides the image group as input to the image processing unit 111, which operates based on the trained image model 151 (step S306).
[0075] Then, the image processing unit 111 and the feature extraction unit 112 execute image processing (step S307).
[0076] Here, the image processing unit 111 can be implemented within the feature extraction device 101, or may be implemented in a device independent of the feature extraction device 101 by referring to the same image model 151.
[0077] 4 is a flowchart showing the control flow of image processing for obtaining feature information from a group of images. The following description will be made with reference to this figure.
[0078] When image processing starts, the image processing unit 111 accepts input of a group of images sequentially, in parallel, or all at once (step S401), and repeats the following processing for each of the images included in the input group of images (step S402).
[0079] That is, the image processing unit 111 provides the image to the neural network associated with the image model 151 (step S403), and obtains likelihoods and feature parameters output from the neural network (step S404).
[0080] After the repetition is completed for all images (step S405), the feature extraction unit 112 selects a predetermined number of representative images based on the obtained likelihoods (step S406).
[0081] Then, the feature extraction unit 112 collects the feature parameters obtained for the selected representative image and outputs them as feature information (step S407), and ends this process.
[0082] Returning to FIG. 3, the classification training unit 133 acquires the feature information output from the image processing unit 111 (step S308).
[0083] Then, the classification training unit 133 provides the acquired feature information and, if input, additional data as input to the classifier associated with the classification model 153 (step S309), obtains the results output from the classifier (step S310), and determines the difference between the output results and the labels (step S311).
[0084] After the repetition for each pair is completed (step S312), classification training unit 133 calculates the value of the evaluation function based on the difference found for each pair, updates classification model 153 (step S313), and returns control to step S302.
[0085] When the training of the classification model is completed (step S302; Yes), the classification training unit 133 ends this process.
[0086] In addition, by using a library, it is possible to perform feature extraction and classification learning processes in parallel or in parallel at high speed.
[0087] Furthermore, training of the image model 151 and the classification model 153 may be completed when the number of times the model updates are repeated reaches a predetermined number, or when a predetermined convergence condition is satisfied.
[0088] (Feature Extraction Processing) Fig. 5 is a flowchart showing the control flow of the feature extraction processing, which will be described below with reference to this figure.
[0089] When this process starts, the feature extraction device 101 receives an input of a group of images related to a target (step S501).
[0090] Then, the feature extraction device 101 provides the input image group to the image processing unit 111 (step S502), and causes the image processing unit 111 and the feature extraction unit 112 to execute the above-mentioned image processing (step S503). The image processing unit 111 then calculates the likelihood and feature parameters of each image in the image group, and the feature extraction unit 112 selects a predetermined number of representative images from the image group based on the likelihood, compiles the feature parameters of the representative images, and outputs them as feature information of the image group.
[0091] Then, the feature extraction device 112 acquires the feature information of the image group output from the feature processing unit 111 (step S504).
[0092] Next, the feature processing unit 112 outputs the acquired feature information as feature information relating to the target (step S505), and ends this process.
[0093] (Classification Process) Fig. 6 is a flowchart showing the control flow of the classification process, which will be described below with reference to this figure.
[0094] When this process starts, the classification processing unit 113 of the feature extraction device 101 receives input of a group of images related to the target and additional data (if any) (step S601).
[0095] The input image group is then given as input to the feature processing unit 112 (step S602), which then executes the feature extraction process described above (step S603).
[0096] Then, the classification processing unit 113 acquires the feature information output from the feature processing unit 112 (step S604), and inputs the acquired feature information and the additional data (if input) to a classifier based on the classification model 153 (step S605).
[0097] Then, the classification processing unit 113 causes the classifier to estimate whether or not the object belongs to the first class based on the classification model 153 (step S606), outputs the result (step S607), and ends this processing. The output result may include information on whether or not the object belongs to the first class, as well as the probability thereof.
[0098] (Experimental Results) Hereinafter, experimental results for the embodiment of estimating the presence or absence of prostate cancer using ultrasound images will be described.
[0099] For training and validation, we prepared 2,899 ultrasound images for 772 subjects obtained between November 2017 and June 2020.
[0100] The size of each image is normalized to 256 x 256 pixels.
[0101] Each image is accompanied by information about whether the subject was affected or not.
[0102] The image is associated with a Gleason score assigned by an expert through microscopic observation of a biopsy specimen separately obtained from the subject.
[0103] For the first image class, two methods were tested: one based on whether the subject was diseased (Cancer classification), and the other based on whether the Gleason score assigned to the image was 8 or higher (High-grade cancer classification).
[0104] The likelihood was calculated as the probability that an image belongs to the first image class, and the likelihood was used as the feature parameter. For each image group, three representative images were selected in descending order of the absolute value of the difference from 0.5.
[0105] In addition, clinical data such as age, PSA value, TPV value, and PSAD value were used as additional data.
[0106] The neural network related to the image model 151 was applied to Xception, inceptionV3, and VGG16.
[0107] For the classification model 153, three types were used: Ridge, Lasso, and Support Vector Machine (SVM).
[0108] First, the experimental results for image model 151 showed that Xception performed best, with an accuracy of 0.693 (95% confidence interval: 0.640-0.746) for Cancer classification and 0.723 (95% confidence interval: 0.659-0.788) for High-grade cancer classification.
[0109] Next, the experimental results for classification model 153 showed that SVM was the best, with an accuracy of 0.807 (95% confidence interval: 0.719-0.894) for Cancer classification and 0.835 (95% confidence interval: 0.753-0.916) for High-grade cancer classification.
[0110] In addition, when subjects were classified using only clinical data without using ultrasound images (prior art), the accuracy of the SVM that achieved the best results was 0.722 (95% confidence interval range: 0.620-0.824), which shows that the accuracy is significantly improved by applying the feature extraction device 101 of this embodiment.
[0111] Below, the conventional technology and this embodiment will be compared using ROC curves (Receiver Operating Characteristic Curves). FIG. 7 is a graph showing experimental results of classification using the conventional method. FIG. 8 is a graph showing experimental results of classification using this embodiment. FIG. 9 is an explanatory diagram comparing the graph showing experimental results of classification using this embodiment and the graph showing experimental results of classification using the conventional method, superimposed on each other. These figures show two ROC curves: one using the conventional method based only on clinical data, and one using this embodiment. The ROC curve for this embodiment is shifted to the upper left compared to the conventional method, and the area below is larger for this embodiment than for the conventional method. Therefore, it can be seen that the method according to this embodiment is more effective than the conventional method.
[0112] Additionally, as a supplementary experiment, we attempted to classify subjects using Ridge and Lasso without additional data. When classification was attempted without selecting a representative image according to this embodiment, the accuracies were 0.722 and 0.769, respectively. However, in this embodiment, the accuracies were 0.801 and 0.802, respectively. It was found that the feature extraction device 101 according to this embodiment selects a representative image, thereby improving performance. In this embodiment, even for subjects with cancer, if the cancer site is not captured in the image, that image is not selected as the representative image, which is thought to be effective in reducing noise.
[0113] In the above experiment, this embodiment was used to estimate the presence or absence of prostate cancer using ultrasound images, but as mentioned above, this embodiment can also be applied to diseases other than prostate cancer and to images other than ultrasound images.
[0114] In other words, it can be used to extract features of a subject from a large number of images of the subject to estimate whether the subject is suffering from a particular disease, or more generally, to extract features of an object from a large number of images of the object to help classify the object.
[0115] (Summary) As described above, the feature extraction device of this embodiment includes: an image processing unit that, when an image is input, calculates the likelihood that the input image belongs to a first image class and the feature parameters of the input image using an image model; and a feature processing unit that, when an image group is input, inputs images included in the input image group to the image processing unit to calculate the likelihood and feature parameters, selects a predetermined number of representative images from the input image group based on the calculated likelihood, and outputs the calculated feature parameters for the selected predetermined number of representative images as feature information of the image group.
[0116] Furthermore, the feature extraction device according to this embodiment may further include a classification processing unit that, when a group of target images relating to an object is input, inputs the input group of target images to the feature processing unit and estimates, using a classification model, whether or not the object belongs to a first object class from feature information output from the feature processing unit, and may be configured such that the fact that an image included in the group of target images relating to the object belongs to the first image class correlates with the object belonging to the first object class.
[0117] Furthermore, in the feature extraction device according to this embodiment, the classification processing unit can be configured to further receive additional data relating to the object, and to estimate, using the classification model, whether or not the object belongs to the first object class from the output feature information and the input additional data.
[0118] Furthermore, in the feature extraction device according to this embodiment, the group of target images can be configured to consist of a plurality of images of the subject's prostate taken by ultrasound, the additional data includes the subject's age, PSA value, TPV value, and PSAD value, and the first target class is a class indicating that the subject is suffering from prostate cancer.
[0119] Furthermore, in the feature extraction device according to this embodiment, in the training data for the image model, the first image class can be configured to be a class representing that the Gleason score assigned to a specimen portion corresponding to an image portion depicted in the image in a biopsy specimen is equal to or greater than a predetermined value.
[0120] Furthermore, in the feature extraction device according to this embodiment, in the training data for the image model, the first image class can be configured to be a class representing that the subject related to the image is suffering from prostate cancer.
[0121] In the feature extraction device according to this embodiment, the feature processing unit may be configured to select the predetermined number of representative images in descending order of the absolute value of the difference between the likelihood and a predetermined reference value.
[0122] Furthermore, in the feature extraction device according to this embodiment, the likelihood can be configured to be a value between 0 and 1, and the predetermined reference value can be configured to be 0.5.
[0123] In the feature extraction device according to this embodiment, the feature processing unit may be configured to select the predetermined number of representative images in descending order of likelihood.
[0124] In addition, in the feature extraction device according to this embodiment, the feature processing unit can be configured to select images for which the likelihood is the minimum, median, or maximum as the predetermined number of representative images.
[0125] Furthermore, the feature extraction device according to this embodiment can be configured such that the feature parameter calculated for the image is a likelihood calculated for the image.
[0126] Furthermore, the feature extraction device according to this embodiment can be configured such that the feature parameters calculated for the image are intermediate vectors of the image in the image model.
[0127] In addition, in the feature extraction device according to this embodiment, the image model can be configured to be a model related to a deep convolutional neural network.
[0128] In addition, in the feature extraction device according to this embodiment, the classification model can be configured to be a model related to linear regression, logistic regression, ridge regression, lasso regression, or support vector machine.
[0129] The feature extraction method according to this embodiment comprises the steps of: receiving a group of target images relating to a target into a feature extraction device; calculating, by the feature extraction device, the likelihood that an image included in the input group of target images belongs to a first image class and the feature parameters of the image using an image model; selecting, by the feature extraction device, a predetermined number of representative images from the input group of target images based on the calculated likelihood; and outputting, by the feature extraction device, the calculated feature parameters for the selected predetermined number of representative images as feature information of the group of images.
[0130] The program according to this embodiment causes a computer to function as an image processing unit that, when an image is input, calculates the likelihood that the input image belongs to a first image class and the feature parameters of the input image using an image model; and, when a group of images is input, inputs images included in the input group of images into the image processing unit to calculate the likelihood and feature parameters, selects a predetermined number of representative images from the input group of images based on the calculated likelihood, and outputs the calculated feature parameters for the selected predetermined number of representative images as features of the target.
[0131] The computer-readable non-transitory information recording medium according to this embodiment stores the above program.
[0132] The present invention allows for various embodiments and modifications without departing from the broad spirit and scope of the present invention. Furthermore, the above-described embodiments are intended to illustrate the present invention and do not limit the scope of the present invention. That is, the scope of the present invention is defined by the claims, not the embodiments. Various modifications made within the scope of the claims and their equivalents are deemed to be within the scope of the present invention. This application claims priority based on patent application No. 2021-089721, filed in Japan on Friday, May 28, 2021, and the contents of that basic application are incorporated herein to the extent permitted by the laws and regulations of the designated countries.
[0133] According to the present invention, it is possible to provide a feature extraction device, a feature extraction method, a program, and an information recording medium for extracting features of an object from a plurality of images relating to the object.
[0134] 101 Feature extraction device 111 Image processing unit 112 Feature processing unit 113 Classification processing unit 131 Image training unit 133 Classification training unit 151 Image model 153 Classification model
Claims
1. an image processing unit that, when an image is input, reduces the dimensions of the input image using an image model using a network having a plurality of layers, calculates a likelihood that the input image belongs to a first image class, and sets an intermediate vector output from an intermediate layer in the middle of the plurality of layers as a feature parameter of the input image; When a group of images is input, inputting an image included in the input image group into the image processing unit, and calculating a likelihood and a feature parameter; selecting a predetermined number of representative images from the input image group based on the calculated likelihood; A vector, tensor, or array in which the intermediate vectors that are set as feature parameters of the selected predetermined number of representative images are arranged is output as feature information of the image group. a feature processing unit for When a group of target images related to a target and additional data related to the target are input, Feature information output from the feature processing unit by inputting the input target image group into the feature processing unit; and The input additional data a classification processing unit that estimates whether the object belongs to a first object class using a classification model. A feature extraction device comprising:
2. The fact that an image included in the target image group relating to the target belongs to the first image class correlates with the object belonging to the first object class.
2. The feature extraction device according to claim 1 .
3. The group of target images is made up of a plurality of small images obtained by dividing a single photograph of the target into a predetermined size.
2. The feature extraction device according to claim 1 .
4. the target image group is made up of a plurality of images obtained by photographing the prostate of the target using ultrasound; The additional data includes the subject's age, PSA value, TPV value, and PSAD value; The first subject class is a class representative of the subject suffering from prostate cancer.
2. The feature extraction device according to claim 1 .
5. In the training data for the image model, the first image class is a class representing that the Gleason score given to a specimen portion corresponding to an image portion depicted in the image in a biopsy specimen is equal to or greater than a predetermined value.
5. The feature extraction device according to claim 4.
6. In the training data for the image model, the first image class is a class representing that a subject related to the image has prostate cancer.
5. The feature extraction device according to claim 4.
7. the feature processing unit selects the predetermined number of representative images in descending order of absolute value of the difference between the likelihood and a predetermined reference value; The feature parameter is an N-dimensional (N≧2) vector, The predetermined number is M (M≧2), The feature information is an N×M dimensional vector.
2. The feature extraction device according to claim 1 .
8. The likelihood is a value between 0 and 1, The predetermined reference value is 0.
5.
8. The feature extraction device according to claim 7.
9. the feature processing unit selects the predetermined number of representative images in descending order of the likelihood; The feature parameter is an N-dimensional (N≧2) vector, The predetermined number is M (M≧2), The feature information is an N×M dimensional vector.
2. The feature extraction device according to claim 1 .
10. the feature processing unit selects images for which the likelihood is a minimum value, a median value, and a maximum value as the predetermined number of representative images; The feature parameter is an N-dimensional (N≧2) vector, the predetermined number is 3, The feature information is an N×3 dimensional vector.
2. The feature extraction device according to claim 1 .
11. The image model is a model related to a deep convolutional neural network.
2. The feature extraction device according to claim 1 .
12. The classification model is a model related to linear regression, logistic regression, ridge regression, lasso regression, or support vector machine.
2. The feature extraction device according to claim 1 .
13. A step of inputting a group of target images relating to a target and additional data relating to the target to a feature extraction device; a step in which the feature extraction device reduces the dimensions of an image included in the input target image group by an image model using a network having multiple layers, calculates a likelihood that the image belongs to a first image class, and sets an intermediate vector output from an intermediate layer in the multiple layers as a feature parameter of the image; a step of the feature extraction device selecting a predetermined number of representative images from the input target image group based on the calculated likelihood; a step of outputting, by the feature extraction device, a vector, a tensor, or an array in which the intermediate vectors that are set as feature parameters of the selected predetermined number of representative images are arranged, as feature information of the image group; a step of estimating, based on the output feature information and the input additional data, whether the object belongs to a first object class using a classification model. A feature extraction method comprising:
14. Computer, an image processing unit that, when an image is input, reduces the dimensions of the input image using an image model using a network having a plurality of layers, calculates a likelihood that the input image belongs to a first image class, and sets an intermediate vector output from an intermediate layer in the middle of the plurality of layers as a feature parameter of the input image; When a group of images is input, inputting an image included in the input image group into the image processing unit, and calculating a likelihood and a feature parameter; selecting a predetermined number of representative images from the input image group based on the calculated likelihood; A vector, tensor, or array in which the intermediate vectors that are set as feature parameters of the selected predetermined number of representative images are arranged is output as feature information of the image group. a feature processing unit for When a group of target images related to a target and additional data related to the target are input, Feature information output from the feature processing unit by inputting the input target image group into the feature processing unit; and The input additional data a classification processing unit that estimates whether the object belongs to a first object class using a classification model. A program characterized by causing the program to function as a 15. A computer-readable non-transitory information recording medium having the program according to claim 14 recorded thereon.