Medical information processing apparatus, medical information processing method, and program

Iterative Metamorphic Testing in medical data analysis automatically distinguishes between fundamental and accidental errors in model outputs, enhancing model accuracy and reducing maintenance costs by generating and comparing distributions of model outputs.

JP2025112139APending Publication Date: 2025-07-31CANON MEDICAL SYST CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024006245
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing metamorphic testing methods struggle to automatically determine whether an error output is fundamental or accidental in medical data analysis, complicating the identification of bugs in models due to variations in follow-up outputs.

Method used

A medical information processing apparatus and method utilizing Iterative Metamorphic Testing, which generates follow-up data from source data, calculates distributions of model outputs, and compares them to determine if errors are inherent or accidental, without requiring user labeling.

Benefits of technology

Enables automatic verification of medical models to distinguish between fundamental and accidental errors, reducing maintenance costs and improving model accuracy by identifying the need for re-learning or input data preprocessing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025112139000001_ABST
    Figure 2025112139000001_ABST
Patent Text Reader

Abstract

To automatically determine whether output of an error in a model is theoretical.SOLUTION: A medical information processing apparatus comprises data generation means, first distribution calculation means, second distribution acquisition means, and output means. The data generation means applies a change to a source acquired from first storage means to generate follow-up data. The first distribution calculation means inputs the source and the follow-up data to a sorter, and acquires a first distribution that is the distribution of outputs from the sorter. The second distribution acquisition means acquires to which class in the sorter the source belongs, and acquires, from second storage means, a second distribution that is the distribution of outputs related to the class of the source. The output means compares the first distribution with the second distribution to output a result of inspection of the sorter that is an object to be inspected.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed in this specification and the drawings relate to a medical information processing device, a medical information processing method, and a program.

Background Art

[0002] Metamorphic Testing (MT) is a technique for generating new test cases from existing test cases, and is a test technique for finding processing errors when there is no test oracle. MT determines whether there is an error based on whether a pre-defined metamorphic relation (MR, Metamorphic Properties) is satisfied or not.

[0003] This MT has the merit that generation of correct data is unnecessary. Therefore, it is effective as a technique for testing a system that inputs medical data that requires time-consuming labeling. For example, by repeatedly using MT for a certain source input data, it is possible to analyze bugs potentially included in the model itself (including a learned model).

[0004] However, in MT, variations may occur in the follow-up output, and in this case, analysis becomes difficult. It is desirable to be able to automatically determine whether the model accidentally outputs an error or outputs an error due to a potential bug from multiple follow-up outputs.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] One of the problems to be solved by the embodiments disclosed in this specification and the drawings is to automatically determine whether an error output is fundamental. However, the problems to be solved by the embodiments disclosed in this specification and the drawings are not limited to the above problem. The problems corresponding to the respective effects of the respective configurations shown in the embodiments described later can also be regarded as other problems.

Means for Solving the Problems

[0007] According to one embodiment, a medical information processing apparatus includes at least a data generation means, a first distribution calculation means, a second distribution acquisition means, a determination means, and an iteration means. The medical information processing apparatus further includes, or is accessible to, a first storage means and a second storage means. The first storage means stores a source which is data related to medical information to be verified. The second storage means stores, in association with each other, a class classified by a classifier to be inspected and a distribution of outputs related to the class. The data generation means generates follow-up data by making a change to the source acquired from the first storage means. The first distribution calculation means inputs the source and the follow-up data to the classifier, and acquires a first distribution which is a distribution of outputs of the classifier. The second distribution acquisition means acquires to which class among the classes in the classifier the source belongs, and acquires from the second storage means a second distribution which is a distribution of outputs related to the class of the source. The output means outputs an inspection result of the classifier to be inspected by comparing the first distribution and the second distribution.

[0008] According to one embodiment, a medical information processing method is a first storage means that stores a source which is data related to medical information to be verified, a second storage means that stores, in association with each other, a class classified by a classifier to be inspected and a distribution of outputs related to the class, in a medical information processing apparatus including, or accessible to, these. The data generation means generates follow-up data by making changes to the source obtained from the first storage means, the first distribution calculation means inputs the source and the follow-up data into the classifier, and obtains a first distribution that is the distribution of the output of the classifier, the second distribution acquisition means obtains to which class among the classes in the classifier the source belongs, and obtains from the second storage means a second distribution that is the distribution of the output related to the class of the source, the output means outputs an inspection result of the classifier that is the inspection target by comparing the first distribution and the second distribution.

[0009] According to one embodiment, the program a first storage means for storing a source that is data related to medical information to be verified, a second storage means for associating and storing a class classified in the classifier that is the inspection target and a distribution of the output related to the class, comprises, or is accessible to, a medical information processing apparatus a data generation means for generating follow-up data by making changes to the source obtained from the first storage means, a first distribution calculation means for inputting the source and the follow-up data into the classifier and obtaining a first distribution that is the distribution of the output of the classifier, a second distribution acquisition means for obtaining to which class among the classes in the classifier the source belongs and obtaining from the second storage means a second distribution that is the distribution of the output related to the class of the source, an output means for outputting an inspection result of the classifier that is the inspection target by comparing the first distribution and the second distribution, is caused to function as.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Mode for Carrying Out the Invention

[0011] Hereinafter, embodiments will be described with reference to the drawings. In the present disclosure, regarding the form in which information processing by software is specifically realized using hardware resources, this software can be executed by a program and can be implemented by the program itself or a non-transitory computer-readable medium storing the program. Also, among various data in this specification and the drawings, those mainly used by the information processing apparatus are typically digital data.

[0012] (First Embodiment)

[0013] A metamorphic test is a test to check whether a classifier (which may hereinafter be referred to as a model, including a segmentation model) produces appropriate outputs for verification data when a test oracle cannot be defined or is difficult to define. The present disclosure uses Iterative Metamorphic Testing, which repeatedly performs metamorphic tests using different follow-up data generated thereby, to automatically determine whether the target model lacks the performance to correctly determine in the first place because the model's inference is incorrect (hereinafter, for convenience, it may be described as making a fundamentally and logically incorrect determination), or whether the model's inference is correct but the determination is accidentally incorrect due to noise, deformation, etc. in the input data of a person for whom the model is constructed to perform correct classification (hereinafter, for convenience, it may be described as accidentally making an incorrect determination). Some embodiments of a medical information processing apparatus will be described while giving examples.

[0014] FIG. 1 is a block diagram schematically showing a medical information processing apparatus according to an embodiment. The medical information processing apparatus 1 includes an input unit 10, a storage unit 20, a processing unit 30, and an output unit 40. When an incorrect determination is made in the model, the medical information processing apparatus 1 determines whether the model makes a fundamentally and logically incorrect determination or just accidentally makes an incorrect determination regarding this incorrect determination. In addition to what is shown in the figure, the medical information processing apparatus 1 can be appropriately provided with a power supply unit that supplies appropriate power to each component of the medical information processing apparatus 1, a control unit that controls each component of the medical information processing apparatus 1, etc. as necessary.

[0015] The input unit 10 receives data from the outside. Also, the processing unit 30 can acquire data from the outside via the input unit 10. The input unit 10 may be configured by an appropriate input interface. The input unit 10 can operate as input means in the present disclosure.

[0016] The storage unit 20 stores data necessary for the processing of the medical information processing apparatus 1. The storage unit 20 may be configured to include an arbitrary storage circuit. The processing unit 30 can execute processing, for example, by referring to the data in the storage unit 20. The processing unit 30 can store the processed content in the storage unit 20 temporarily or non-temporarily, for example. The storage unit 20 can operate as the storage means in the present disclosure.

[0017] A part of the storage unit 20 may be provided as an external storage device instead of being provided in the medical information processing apparatus 1. In this case, the processing unit 30 can also access the external storage device via the input unit 10 and receive data. That is, the medical information processing apparatus 1 may be configured to include the storage unit 20, or may be configured to be able to access the storage unit 20.

[0018] The storage unit 20 may store, for example, data used as a source for performing a metamorphic test of a model. Further, the storage unit 20 may store, for example, by associating the class classified by the model to be subjected to the metamorphic test and the distribution of the output classes (distribution including the result of misjudgment) of the results of inputting the teacher data of each class.

[0019] The processing unit 30 makes a determination of the model. The processing unit 30 can be configured by an arbitrary processing circuit. The processing unit 30 includes at least a data acquisition means 300, a data generation means 302, a first distribution calculation means 304, a second distribution acquisition means 306, and a determination means 308 as means for processing. The processing unit 30 can operate as various processing means in the present disclosure.

[0020] Note that the processing unit 30 does not necessarily have all the means in one processing circuit, and the above processing means may be realized by a plurality of processing units 30. Further, it is not necessary for each means of the processing unit 30 to be realized in one medical information processing apparatus 1, and the processing units 30 provided in a plurality of medical information processing apparatuses 1 may cooperate to realize the above means.

[0021] The data acquisition means 300 acquires data serving as a source for executing a metamorphic test on the model from an external storage unit via the input unit 10 or from the storage unit 20.

[0022] The external storage unit or the storage unit 20 can operate as a first storage means for storing a source which is data for verifying the model. The first storage means may include, for example, a database storing medical data.

[0023] The data generation means 302 creates one or more follow-up data for following up the source from the source acquired by the data acquisition means 300. The follow-up data is data obtained by making a change to the source data and is data for executing a metamorphic test. For example, if the source data is medical image data, the data generation means 302 can obtain, as follow-up data, image data obtained by rotating the medical image data by an arbitrary angle, image data obtained by adding noise to the medical image data, image data obtained by changing the distribution by increasing or decreasing the distribution of the luminance values of the medical image data, and the like.

[0024] The first distribution calculation means 304 inputs the source acquired by the data acquisition means 300 and the follow-up data acquired by the data generation means 302 into the model, and acquires a first distribution which is the distribution of the output classes.

[0025] The second distribution acquisition means 306 acquires, from an external storage unit via the input unit 10 or from the storage unit 20, a second distribution that is the distribution of the classes (including the results of misjudgments) output by the model into which a plurality of correct data of the class have been input, for the classes classified by the model of the source acquired by the data acquisition means 300.

[0026] The external storage unit or the storage unit 20 can operate as second storage means for associating and storing a class and a second distribution that is the distribution of the classes obtained by inputting a plurality of data of the class into the model.

[0027] The determination means 308 compares the first distribution and the second distribution to determine whether there is essentially no error in the model.

[0028] The output unit 40 outputs the determination result after the comparison and determination process by the determination means 308 is completed. The output unit 40 may be connected to an external storage means by an appropriate means. Also, the output unit 40 may have an arbitrary user interface and can output the result to means such as a display, a speaker, a printer, etc.

[0029] FIG. 2 is a flowchart showing the processing of the medical information processing apparatus 1 according to an embodiment.

[0030] First, a model to be inspected, the distribution of the classes output for the class of the correct data when the correct data is input into the model, and the source to be classified by the model are prepared (S100).

[0031] The model only needs to be a classifier. As a non-limiting example, it may be a model trained by a machine learning method or a model that performs rule-based processing based on statistical quantities or the like. Further, the model may be a segmentation model that classifies information in a broad sense. As a typical example, a trained model that outputs, as classes, the types of organs (such as the liver, pancreas, large intestine, etc.) captured in a CT image by inputting a CT image of abdominal organs can be used.

[0032] The distribution prepared in this process indicates the distribution of the output classes when data with the correct class assigned as a label is input to this model. This distribution can indicate, for example, how frequently an incorrect class is output when an output result is a class different from the correct class (an incorrect class).

[0033] In addition, the medical information processing device 1 executes a metamorphic test using the data to be classified by the model. For this purpose, it is necessary to prepare the data to be classified by the model. When the model is a trained model, it is desirable that this data is not the data used when learning the machine learning.

[0034] Examples of the data may include medical information actually obtained in a medical field. Examples of the types of data include image data such as X-ray image data, CT image data, MRI image data, ultrasonic image data, and tomographic data, vital data such as electrocardiograms, pulse, body temperature, blood pressure, blood test results, weight, and body fat percentage, voice data, interview data, or data related to various specimens. Any data that can be classified and segmented using the model may be used, and the model only needs to be a model that appropriately classifies these data.

[0035] Instead of these medical-related data and models, for example, it may be data that can be classified by other models that perform general image classification and segmentation. The data may be appropriate data to be input into the model, as shown in the above examples.

[0036] Note that "preparing" may mean that the medical information processing apparatus 1 stores the data serving as the model, distribution, and source in the storage unit 20, or may mean making these information accessible.

[0037] The data acquisition means 300 acquires the source via the input unit 10 or acquires the source stored in the storage unit 20 (S102).

[0038] The data generation means 302 repeatedly generates a plurality of different follow-up data from the source acquired by the data acquisition means 300 (S104). The data generation means 302 generates a plurality of different follow-up data by performing various processes on the source. At this time, the data generation means 302 generates follow-up data by making changes to the source. Although the type and amount of the change are adjusted by the user, it is desirable to generate follow-up data to such an extent that the class expected when the source is classified does not change.

[0039] If the source is data including an image, as a non-limiting example, the data generation means 302 can generate follow-up data by performing processes such as rotation of the image, filtering, change in color tone such as brightness, saturation, or lightness, making pixels missing, adding noise, or changing contrast.

[0040] If the source is data including voice, as non-limiting examples, processing such as noise addition, frequency modulation, and data loss may be performed. If the source is vital data, as non-limiting examples, processing such as noise addition may be performed on the vital data. Even for other types of data, the data generation means 302 generates follow-up data by performing appropriate processing.

[0041] FIG. 3 is a diagram showing non-limiting examples of sources. As shown in this figure, the medical information processing apparatus 1 can use, for example, a CT image or the like that is the subject of inspection for model verification. For example, although a tumor has occurred at the location indicated by the circle in this figure, this model may be a model that extracts this tumor.

[0042] Such CT images or the like are captured in the external storage means or storage unit 20 outside the medical information processing apparatus 1. Such CT images or the like can be used as sources in the present disclosure. The data acquisition means 300 acquires a source for verification from these data.

[0043] FIG. 4 is a diagram showing non-limiting examples of follow-up data. The data generation means 302 obtains different follow-up data by repeatedly performing appropriate processing on the source acquired by the data acquisition means 300. For example, as shown in the figure, the data generation means 302 can obtain, as follow-up data, an image obtained by rotating the source image clockwise or counterclockwise. Without being limited thereto, as described above, the data generation means 302 obtains follow-up data by performing processing to the extent that the class does not change.

[0044] The first distribution calculation means 304 inputs the source and the follow-up data into the model, outputs the class related to each data, and obtains the first distribution that is the distribution of the output classes (S106).

[0045] The second distribution acquisition means 306 acquires the second distribution, which is the output distribution of the model corresponding to the class of the source, from the second storage means (S108). The second distribution acquisition means 306 may acquire the second distribution based on the class of the source determined by the first distribution calculation means 304, may acquire the second distribution based on the class of the source determined by means other than the model, or may acquire the second distribution based on the class to which the source is pre-labeled or the like.

[0046] The determination means 308 makes a determination of the model based on the first distribution acquired by the first distribution calculation means 304 and the second distribution acquired by the second distribution acquisition means 306 (S110). The determination means 308, for example, calculates the distance between the first distribution and the second distribution, calculates the accuracy of the model based on this distance, and determines whether the metamorphic relationship is satisfied. Whether the metamorphic relationship is satisfied can be examined, for example, as follows.

[0047] The determination means 308 may calculate the distance based on a statistic such as, for example, the sum of the absolute values of the differences in the probabilities that can occur for each class of the first distribution and the second distribution, the sum of the squared errors, the sum of the square roots of the squared errors, etc. The determination means 308 may calculate the distance based on, for example, the cross entropy between the first distribution and the second distribution, the KL information amount (Kullback-Leibler divergence), etc. In this way, the determination means 308 calculates the distance between the first distribution and the second distribution by a method that can appropriately obtain the distance between the distributions.

[0048] The determination means 308 compares the calculated distance between the distributions with a predetermined threshold value to determine whether the error in the source and the follow-up data is a fundamental error of the model or an error that occurred by chance. The determination means 308, for example, in the case of the distance according to the above example, determines that there is a possibility that the model includes a fundamental error when the distance is greater than the predetermined threshold value, and determines that it was an error that could occur within the second distribution when the distance is smaller than the predetermined threshold value.

[0049] The determination means 308 can also calculate the difference between the first distribution and the second distribution using an index other than the distance. In this case, the determination means 308 can make an appropriate determination based on whether the index is greater than or less than a predetermined threshold value. For example, if the smaller the value of a certain index, the greater the difference between the first distribution and the second distribution, when the value of this index is smaller than the predetermined threshold value, it can be determined that there is a possibility that the model contains a fundamental error.

[0050] In addition, the determination means 308 can also set a score based on the distance as an index for determination. This score can be defined as the sum of the distance and an arbitrarily set penalty term. The determination means 308 can define the penalty term by various methods.

[0051] For example, when the randomness in the first distribution, that is, the randomness of the class output when source and follow-up data are input to the model, is high, the determination means 308 can set the penalty term so as to impose a large penalty.

[0052] For example, the determination means 308 can define the penalty term based on the probability of outputting a class different from the class of the source in the first distribution. The determination means 308 can use the product of the probability of outputting a class different from the class of the source and a predetermined coefficient as the penalty term. Also, when another class other than the class of the source appears in the first distribution, the determination means 308 can determine a lower limit value of the number of appearances, and it may be assumed that a penalty occurs when the lower limit value is exceeded. Furthermore, the determination means 308 can also define the penalty term based on the variation of the output classes, that is, the variance, standard deviation, etc.

[0053] For example, in the first distribution, when there is a class that does not exist in the second distribution, the determination means 308 can increase the penalty term. In this case, when the class that exists in the first distribution also exists in the second distribution, the determination means 308 can also set the penalty to 0. The determination means 308 may set the penalty term based on the existence distribution of the class that does not exist in the second distribution in the first distribution. Further, the determination means 308 can also define the penalty term in combination with the above definition of the penalty term.

[0054] The setting of the penalty term is given as a non-limiting example and is not limited to these.

[0055] Further, the determination means 308 may calculate the penalty term such that the value for the corresponding follow-up data decreases according to the number of iterations by the data generation means 302. For example, the determination means 308 may reduce the value added to the penalty term for the follow-up data as the number of iterations for obtaining the follow-up data increases, even when the follow-up data is not the class with the maximum value in the second distribution or is a class that does not exist in the second distribution.

[0056] This is because the influence of the change increases as the generation of the follow-up data is repeated, and the influence of the follow-up data on the distribution of the source itself decreases. Further, the data generation means 302 may generate follow-up data with a greater degree of deformation as the number of iterations of generation increases.

[0057] For example, when the data generation means 302 generates follow-up data with the angle of the source changed, it can generate follow-up data such that the angle increases according to the number of iterations. For example, the data generation means 302 can first rotate the source to the right once, rotate the source to the left once in the next iteration, rotate the source to the right twice in the next iteration, and so on, to generate follow-up data. By generating such follow-up data, the effect of setting a penalty term according to the number of iterations can be enhanced.

[0058] In addition, when the class for the source can be clinically determined with a high probability, the determination means 308 can define a penalty term so as to impose a large penalty when a class different from this clinical finding appears in the first distribution. For example, when a class different from the clinical finding appears, the determination means 308 can add a larger penalty compared to the penalty for each of the above-mentioned follow-up data.

[0059] The determination means 308 determines whether the model is fundamentally incorrect based on the above-mentioned distance calculated from the comparison between the first distribution and the second distribution, and the penalty term. This determination result may be output sequentially. The determination means 308 can output the calculation result and / or the determination result at an appropriate timing to the outside via the output unit 40 or to the storage unit 20 (S112).

[0060] As a result of the above processing, the medical information processing apparatus 1 can execute verification of the model while reducing the influence when the model accidentally outputs an incorrect result by comparing the first distribution and the second distribution.

[0061] As described above, according to the present embodiment, it is possible to determine whether a model has a fundamental error in a situation where a test oracle cannot be generated or in a situation where it is difficult to generate a test oracle. By this determination, it becomes possible to appropriately judge the accuracy of the model. When the model has a fundamental error, that is, when the inference is not correctly performed, an approach of re-learning or reconstructing the model is effective in order to improve the inference accuracy of the model. On the other hand, when the model has made a mistake by chance, that is, when the inference itself is correctly performed but the classification is accidentally incorrect due to a change in the input data, an approach such as pre-processing the input data input to the model instead of re-learning or reconstructing the model can be taken to avoid re-learning and reconstruction with high computational load and efficiently obtain a highly accurate classification result.

[0062] For this determination, it is not necessary to label data by a user, for example, a medical professional. Therefore, it can be automatically realized based on data such as medical information for which a test for a certain model can be obtained. By automatically realizing this, it becomes possible to reduce the maintenance cost of the model. The medical information processing apparatus 1 can also realize the verification of the model even in the phase of operating the model.

[0063] (Second Embodiment)

[0064] FIG. 5 is a block diagram schematically showing a medical information processing apparatus according to an embodiment. The medical information processing apparatus 1 can further include data extraction means 310 in addition to the configuration in the above-described first embodiment.

[0065] The data extraction means 310 extracts follow-up data for which the class has not been appropriately determined. The data extraction means 310 extracts, for example, data for which the class is not appropriate from the follow-up data generated by the data generation means 302. The data extraction means 310 may extract, for example, data to which a penalty is added in the determination by the determination means 308 from the follow-up data generated by the data generation means 302.

[0066] The data extraction means 310 transmits the extracted data to the outside via the output unit 40 or stores it in the storage unit 20. The extracted data can be used as data for re-learning the model used for the determination of the data. In the medical information processing apparatus 1 or another information processing apparatus other than the medical information processing apparatus 1, by performing re-learning of the model including the data extracted by the data extraction means 310, it is possible to make the model a more robust model.

[0067] This extracted data can also be used as data for fine-tuning in addition to re-learning.

[0068]

[0069] (Third Embodiment)

[0070] The medical information processing apparatus 1 can further realize information processing for reinforcing the model together with verification of the model. When the medical information processing apparatus 1 executes processing according to the processing of FIG. 2, it is also possible to input data to be verified by the actual model instead of the source.

[0071] FIG. 6 is a block diagram schematically showing a medical information processing apparatus according to an embodiment. The medical information processing apparatus 1 can include a classification means 312 in addition to the configuration shown in FIG. 1. Note that the medical information processing apparatus 1 can further include the data extraction means 310.

[0072] ​The medical information processing apparatus 1, for example, after acquiring data to be inspected via the input unit 10, creates follow-up data of the data by the data generation means 302. The first distribution calculation means 304 generates a first distribution based on this data and the follow-up data.

[0073] The classification means 312 estimates the class of the inspection target data from the first distribution in the inspection target data. The classification means 312 can, for example, estimate the class having the mode value of the first distribution as the class of the inspection target data. In addition, the classification means 312 can also perform estimation based on various statistical quantities.

[0074] Furthermore, the second distribution acquisition means 306 acquires a second distribution based on the class of the inspection target data estimated by the classification means 312, and the determination means 308 can also realize the verification of the model by comparing this second distribution with the first distribution.

[0075] As described above, according to the medical information processing apparatus 1 according to the present embodiment, inspection using a model can be realized, and the model can be verified.

[0076] Also, in the above-described embodiment, for example, in order to determine the class of the source, it is also possible to use the estimation result of the class by the classification means 312.

[0077] (Fourth Embodiment)

[0078] The medical information processing apparatus 1 may further repeatedly perform processing from the acquisition of the source to improve the robustness of the test.

[0079] FIG. 7 is a block diagram schematically showing a medical information processing apparatus according to an embodiment. The medical information processing apparatus 1 can include an iteration means 314 in addition to the configuration shown in FIG. 1. Note that the medical information processing apparatus 1 can also include this iteration means 314 in the configurations shown in FIGS. 5 and 6.

[0080] After the determination for one source acquired by the data acquisition means 300 is completed, for example, the medical information processing apparatus 1 determines whether to continue the test. If the test is to be continued, the processing from the acquisition of the source may be repeatedly executed.

[0081] By repeatedly executing the test for different sources, the medical information processing apparatus 1 can determine the accuracy of the model for more diverse input data.

[0082] The iteration means 314 may, for example, prepare in advance a plurality of sources to be used for the test and repeatedly execute the metamorphic test until the test using these sources is completed. Also, an arbitrary test completion condition may be set, and the iteration means 314 may repeatedly execute the processing while changing the source until this condition is satisfied.

[0083] As another example, the iteration means 314 may repeatedly execute the acquisition of different follow-up data for the same source. The iteration means 314 can also, for example, repeatedly execute the processing from the acquisition of different follow-up data according to the test result for the acquired follow-up data to perform a test on the robustness of the accuracy of the model.

[0084] For example, when the determination means 308 determines that the model is highly likely to include a fundamental error, the iteration means 314 can control to generate follow-up data with a smaller degree of deformation from the source by the data generation means 302 and repeat the test.

[0085] As another example, when the determination means 308 determines that the model is unlikely to include a fundamental error, the iteration means 314 can control to generate follow-up data with a larger degree of deformation from the source by the data generation means 302 and repeat the test.

[0086] In any case, the setting of the penalty term can also be changed according to the number of iterations of the iterative means 314.

[0087] In all of the above-described embodiments, it is desirable to appropriately define the metamorphic relationship. However, it is also possible to realize the verification of the model without clearly defining the metamorphic relationship by comparing the output distribution of the model and the output distribution of the data for model verification. As a result, it is possible to suppress the need for the user or maintainer to put in a great deal of effort such as pre-defining the metamorphic relationship of the verification data.

[0088] In the above embodiment, the input interface can be realized by a trackball, a switch button, a mouse, a keyboard, a touch pad for performing an input operation by touching an operation surface, a touch screen in which a display screen and a touch pad are integrated, a non-contact input circuit using an optical sensor, a voice input circuit, and the like for performing various settings and the like. The input interface is connected to the control circuit, converts the input operation received from the operator into an electrical signal, and outputs it to the control circuit. Note that in this specification, the input interface is not limited to those including physical operation components such as a mouse and a keyboard. For example, an electrical signal processing circuit that receives an electrical signal corresponding to an input operation from an external input device provided separately from the apparatus and outputs this electrical signal to the control circuit is also included in the example of the input interface.

[0089] In the above-described embodiment, each processing function of the information processing function is recorded in the storage circuit in the form of a program executable by a computer. The processing circuit can include a processor. For example, the processing circuit reads and executes the program from the storage circuit to realize the functions corresponding to the respective programs. In other words, the processing circuit in the state of having read each program has each function shown in the processing circuit illustrated in the drawings. Although the drawings have described that each processing function is realized by a single processor, a processing circuit may be configured by combining a plurality of independent processors, and each processor may execute a program to realize the functions. Also, although the drawings have described that a single storage circuit stores the programs corresponding to the respective processing functions, a plurality of storage circuits may be distributed and arranged, and the processing circuit may be configured to read the corresponding programs from the respective storage circuits.

[0090] In the above description, an example in which the "processor" reads and executes a program corresponding to each function from the storage circuit has been described, but the embodiment is not limited to this. The term "processor" can mean, for example, a circuit such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an application specific integrated circuit (ASIC), a programmable logic device (for example, a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)). When the processor is, for example, a CPU, the processor realizes its function by reading and executing a program stored in the storage circuit. On the other hand, when the processor is an ASIC, instead of storing the program in the storage circuit, the function is directly incorporated as a logic circuit in the circuit of the processor. Note that each processor of the present embodiment is not limited to being configured as a single circuit for each processor, and a plurality of independent circuits may be combined to form one processor to realize its function. Further, a plurality of components in the drawings may be integrated into one processor to realize its function.

[0091] According to at least one embodiment described above, it is possible to automatically determine whether or not the output of an error is principled.

[0092] Although some embodiments have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, changes, and combinations of embodiments can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, as well as in the invention described in the claims and the equivalent scope thereof.

Explanation of Reference Numerals

[0093] 1: Medical information processing device, 10: Input unit, 20: Storage unit, 30: Processing unit, 300: Data acquisition means, 302: Data generation means, 304: First distribution calculation means, 306: Second distribution acquisition means, 308: Determination means, 310: Data extraction means, 312: Classification means, 314: Iteration means, 40: Output unit

Claims

1. A first storage means for storing a source which is data related to medical information to be verified; A second storage means for associating and storing a class classified by a classifier to be inspected and a distribution of outputs related to the class; Comprising, or being accessible to the first storage means and the second storage means; A data generation means for generating follow-up data by making a change to the source acquired from the first storage means; A first distribution calculation means for inputting the source and the follow-up data to the classifier and acquiring a first distribution which is a distribution of the output of the classifier; A second distribution acquisition means for acquiring to which class in the classifier the source belongs and acquiring from the second storage means a second distribution which is a distribution of outputs related to the class of the source; An output means for outputting an inspection result of the classifier to be inspected by comparing the first distribution and the second distribution; A medical information processing apparatus comprising the above.

2. The output means: Calculates a distance between the first distribution and the second distribution; Determines whether there is an error in principle in the classifier based on the distance; The medical information processing apparatus according to Claim 1.

3. The output means: Adds a penalty term to the distance to calculate a score related to the classifier; Acquires an inspection result of the classifier based on the score; The medical information processing apparatus according to Claim 2.

4. The output means: Calculates the penalty term based on the randomness in the first distribution; The medical information processing apparatus according to Claim 3.

5. The output means: Increases the penalty term when a class appearing in the first distribution does not appear in the second distribution; The medical information processing apparatus according to Claim 3.

6. The data generation means: Repeatedly generates the follow-up data from the source at different parameters with a degree of deformation increasing according to the number of repetitions; The output means: Calculates the penalty term for a certain follow-up data based on the number of repetitions in the data generation means for generating the follow-up data with respect to the follow-up data; The medical information processing apparatus according to Claim 3.

7. The classifier is a model trained by machine learning; The medical information processing apparatus according to any one of Claims 1 to 6.

8. In the first distribution for a source determined to be fundamentally incorrect in the model by the output means, data extraction means for extracting a source or follow-up data that was not classified into an appropriate class as data for re-learning. The medical information processing apparatus according to claim 7, further comprising the above.

9. The classifier is a classifier used for medical information. The first storage means stores diagnostic data. The medical information processing apparatus according to any one of claims 1 to 6.

10. The medical information includes at least one of image data, vital data, or specimen test data. The medical information processing apparatus according to claim 9.

11. Further comprising classification means, The data generation means generates follow-up data for data to be inspected. The first distribution calculation means obtains the distribution of the class to be inspected from the data to be inspected and the follow-up data for the data to be inspected. The classification means estimates the class of the data to be inspected based on the distribution of the class to be inspected. The medical information processing apparatus according to any one of claims 1 to 6.

12. A first storage means for storing a source which is data related to medical information to be verified; A second storage means for storing in association the classes classified in the classifier to be inspected and the distribution of the outputs related to the classes; In a medical information processing apparatus comprising the above, or accessible to the first storage means and the second storage means, The data generation means generates follow-up data by applying variation candidates to the source obtained from the first storage means. The first distribution calculation means inputs the source and the follow-up data to the classifier and obtains a first distribution which is the distribution of the outputs of the classifier. The second distribution acquisition means obtains to which class among the classes in the classifier the source belongs, and obtains from the second storage means a second distribution which is the distribution of the outputs related to the class of the source. The output means outputs the inspection result of the classifier to be inspected by comparing the first distribution and the second distribution. Medical information processing method.

13. A first storage means for storing a source which is data related to medical information to be verified; A second storage means for storing in association the classes classified in the classifier to be inspected and the distribution of the outputs related to the classes; A medical information processing device that includes or is accessible to the first storage means and the second storage means, Data generation means for generating follow-up data by modifying the source obtained from the first storage means, First distribution calculation means for inputting the source and the follow-up data into the classifier and obtaining a first distribution that is the distribution of the output of the classifier, Second distribution acquisition means for acquiring to which class in the classifier the source belongs and acquiring from the second storage means a second distribution that is the distribution of the output related to the class of the source, Output means for outputting an inspection result of the classifier that is the inspection target by comparing the first distribution and the second distribution, A program that functions as.

Citation Information

Patent Citations

  • Determination method, information processing device and determination program

    JP2023068467A