Verification device, verification method, and verification program
The collation device improves fetal condition prediction by using a grouping and collation process to assign reliable annotations and exclude inconsistent data, enhancing the accuracy of pH value predictions.
Patent Information
- Application Number
- JP2022015123
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-02
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2042-02-02
AI Technical Summary
Existing methods for predicting fetal status from CTG waveforms rely heavily on doctor annotations, which are unreliable due to variability in experience and limited classification capabilities, failing to predict pH values accurately.
A collation device that uses a grouping processing unit to divide samples into safe and risk groups based on explanatory variables, with a collation unit comparing these groups to objective variables to assign reliable annotations, and a learning model that improves prediction accuracy by excluding inconsistent data.
Enhances the reliability of annotations and improves the accuracy of fetal condition prediction by automatically organizing data and generating a learning model that accurately predicts pH values.
Smart Images

Figure 0007778312000001 
Figure 0007778312000002 
Figure 0007778312000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a verification device, a verification method, and a verification program for verifying data. [Background technology]
[0002] CTG (CardioTocoGram) is a waveform that shows the time-dependent changes in fetal heart rate and tocogram (uterine contractions) obtained from a fetal heart rate monitor and an external tocogram, respectively, and is used to evaluate fetal well-being. It is an essential test for assessing the fetus during labor. CTG is useful for early detection of fetal hypoxia and acidosis that can occur during labor and for reducing hypoxic encephalopathy and cerebral palsy in the fetus at birth.
[0003] Doctors evaluate cardiotocography (CTG) by level classification, which evaluates the fetal heart rate baseline and bradycardia pattern, but the waveform of CTG is complex and depends on the experience of the medical professional making the judgment. Also, Non-Patent Document 1 discloses a technology for predicting the fetal condition from the fetal heart rate signal. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Ramanujam, E., et al. "Prediction of Fetal Distress Using Linear and Non-linear Features of CTG Signals." International Conference On Computational Vision and Bio Inspired Computing. Springer, Cham, 2019. Summary of the Invention [Problem to be solved by the invention]
[0005] However, the prediction method presented in the aforementioned Non-Patent Document 1 uses data annotated by doctors to extract features from the waveforms of fetal heart rate and uterine contractions to predict fetal status, but it requires well-organized annotated data. Data annotation depends on the doctor's experience, so reliability varies. Furthermore, doctor annotation alone is limited to classifying data into a few categories (e.g., safe or dangerous) and cannot predict pH values.
[0006] An object of the present invention is to improve the reliability of annotations. [Means for solving the problem]
[0007] A collation device according to one aspect of the invention disclosed in the present application includes a grouping processing unit that divides a group of samples into a first group indicating a first classification and a second group indicating a second classification that is lower in evaluation than the first classification based on explanatory variables of each of the group of samples, and a collation processing unit that compares the classifications of the first and second groups divided by the grouping processing unit with classifications specified by objective variables of each of the group of samples. and if the classification specified by the objective variable of a first sample belonging to the first group among the sample group is the first classification, an annotation indicating the first classification is assigned to the first sample. an output unit that outputs a collation result by the collation unit; the grouping processing unit divides the group of samples to be predicted into the first group and the second group based on the explanatory variables of the group of samples to be predicted; the comparison unit compares the classifications of the first and second groups into which the group of samples to be predicted is divided by the grouping processing unit with a classification specified by the predicted value calculated by the calculation unit; and if the classification specified by the predicted value of a first sample to be predicted that belongs to the first group among the group of samples to be predicted is the second classification, the comparison unit assigns an annotation indicating a contradiction to the first sample to be predicted. It is characterized by: [Effects of the Invention]
[0008] According to the exemplary embodiment of the present invention, it is possible to improve the reliability of annotation. Problems, configurations, and effects other than those described above will become clear from the following description of the embodiment. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is an explanatory diagram showing an example of how a learning DB is prepared. [Figure 2] FIG. 2 is a block diagram illustrating an example of the hardware configuration of the verification device. [Figure 3] FIG. 3 is an explanatory diagram illustrating an example of the learning DB. [Figure 4] FIG. 4 is an explanatory diagram showing an example 1 of measurement factors included in the feature amount. [Figure 5] FIG. 5 is an explanatory diagram showing an example 2 of a measurement factor included in a feature amount. [Figure 6] FIG. 6 is an explanatory diagram showing an example of the basic information DB. [Figure 7] FIG. 7 is a block diagram illustrating an example of the functional configuration of the verification device. [Figure 8] FIG. 8 is a flowchart illustrating an example of a preprocessing procedure performed by the preprocessing unit. [Figure 9] FIG. 9 is an explanatory diagram showing an example of a selection screen displayed during learning. [Figure 10] FIG. 10 is an explanatory diagram showing an example of a display of a selection screen at the time of prediction. [Figure 11] FIG. 11 is a flowchart showing an example of a procedure for a matching process during learning. [Figure 12] FIG. 12 is an explanatory diagram showing an example of the feature amount selection screen. [Figure 13] FIG. 13 is an explanatory diagram showing an example of the matching result display screen. [Figure 14] FIG. 14 is an explanatory diagram showing an example of the matching result display screen. [Figure 15] FIG. 15 is an explanatory diagram showing an example of the matching result display screen. [Figure 16] FIG. 16 is an explanatory diagram showing an example of the matching result display screen. [Figure 17] FIG. 17 is a flowchart illustrating an example of a procedure for a matching process at the time of prediction. DETAILED DESCRIPTION OF THE INVENTION
[0010] <Example of developing a learning database> FIG. 1 is an explanatory diagram showing an example of the preparation of a learning DB. The learning DB 100 is a database that stores, as sample data, a combination of feature amounts serving as explanatory variables and correct answer data serving as target variables for each sample of a pregnant woman. During learning, the correct answer data is, for example, annotations by a doctor for samples after delivery (whether the fetus's condition is safe or dangerous), and during pH prediction, the correct answer data is annotations obtained during learning for samples before delivery (whether the fetus's condition is safe or dangerous).
[0011] Each piece of sample data is clustered into a safe group where the fetus's condition is safe and a risk group where the fetus's condition is dangerous. Clustering is performed by unsupervised learning that does not use correct answer data or by supervised learning that uses correct answer data. Here, unsupervised learning that does not use correct answer data is taken as an example. In unsupervised learning that does not use correct answer data, the sample data group is clustered into two groups, and the group with the larger number of sample data is designated as the safe group, and the group with the smaller number of sample data is designated as the risk group.
[0012] In addition, pH is an indicator of fetal hypoxia or acidosis, and is a value measured by umbilical artery blood gas analysis immediately after delivery. For example, a pH of 7.2 is used as the standard, and if pH > 7.2, the fetus is safe, and if pH ≤ 7.2, the fetus is at risk. For samples from a delivered baby, the pH value is the actual value measured immediately after delivery, and for samples from a pre-delivery baby, the pH value is the predicted value before delivery.
[0013] If the sample data belongs to the safe group and the matching result 101 indicates that the pH value is a safe value (both are safe), the sample data is stored in the learning DB 100 as a learning target and is annotated as "safe." The annotation indicating "safe" is displayed to the user.
[0014] If the matching result 102 (inconsistency) indicates that the sample data belongs to the safe group and the pH value is a dangerous value, the sample data is set as an exception to learning and unavailable for learning in the learning DB 100, and the user is informed that the sample attribute "safe" and the safe pH value are "inconsistent."
[0015] If the matching result 103 (inconsistency) indicates that the sample data belongs to the danger group and the pH value is a safe value, the sample data is set as an exception to the learning target and unavailable for learning in the learning DB 100, and the user is informed that the sample attribute "danger" and the safe pH value are "inconsistent."
[0016] If the matching result 104 indicates that the sample data belongs to the danger group and the pH value is a danger value (both danger), the sample data is set as unusable for learning in the learning DB 100 as an exception to learning, and an annotation indicating "danger" is added. The annotation indicating "danger" is displayed to the user.
[0017] In this way, the matching results 101 to 104 indicating "safe," "inconsistency," and "danger" are automatically added as annotations. This improves the reliability of the annotations. Furthermore, sample data annotated with matching results 102 to 104 indicating "inconsistency" or "danger" are excluded from learning, so only sample data annotated with matching result 101 indicating "safe" remains. This allows automatic organization of the learning DB 100. Furthermore, by generating a learning model using the remaining sample data group, the accuracy of the predicted pH value can be improved.
[0018] <Example of hardware configuration of a verification device> FIG. 2 is a block diagram showing an example of the hardware configuration of a verification device. The verification device 200 includes a processor 201, a storage device 202, an input device 203, an output device 204, and a communication interface (communication IF) 205. The processor 201, the storage device 202, the input device 203, the output device 204, and the communication IF 205 are connected via a bus 206. The processor 201 controls the verification device 200. The storage device 202 serves as a working area for the processor 201. The storage device 202 is a non-transitory or temporary recording medium that stores various programs and data. Examples of the storage device 202 include a read-only memory (ROM), a random access memory (RAM), a hard disk drive (HDD), and a flash memory. The input device 203 inputs data. Examples of the input device 203 include a keyboard, a mouse, a touch panel, a numeric keypad, a scanner, a microphone, and a sensor. The output device 204 outputs data. The output device 204 includes, for example, a display, a printer, and a speaker. The communication IF 205 connects to a network and transmits and receives data.
[0019] <Learning DB100> 3 is an explanatory diagram showing an example of the learning DB 100. The learning DB 100 is stored in the storage device 202 or in a storage device of another computer accessible to the collation device.
[0020] The learning DB 100 has the following fields: sample ID 301, feature amount 302, pH 303, group 304, annotation 305, and non-learning flag 306. A combination of values in the fields 301 to 306 in the same row defines an entry that indicates sample data for one sample. A sample indicates a pregnant woman, and if the heart rate and intensity of labor pains of the same pregnant woman are measured on different dates and times, the resulting samples will be different. In FIG. 3, sample data for m samples (m is an integer greater than or equal to 1) is registered.
[0021] The sample ID 301 is identification information that uniquely identifies a sample. The feature amount 302 is data that indicates the characteristics of the sample identified by the sample ID 301. The feature amount 302 is composed of factors F1 to Fn (n is an integer greater than or equal to 1) such as basic factors such as the age of the pregnant woman, the number of weeks of pregnancy, fetal developmental delay, the number of fetuses, the method of delivery, medication information, and smoking, and measurement factors such as measurement results obtained by a measuring device that measures the heart rate and the intensity of labor pains.
[0022] pH 303 is the umbilical arterial blood gas analysis value immediately after delivery, and a predicted value 331 and an actual measurement value 332 are stored. Before delivery, only the predicted value 331 is calculated and stored by the calculation unit 703 (described later). After delivery, the umbilical arterial blood gas analysis value is measured and stored as the actual measurement value 332.
[0023] The group 304 is the group to which the sample identified by the sample ID 301 belongs. There are two groups: a safe group and a risk group. As explained in FIG. 1, when clustering is performed using unsupervised learning, the group with the larger number of sample data is designated as the safe group, and the group with the smaller number of sample data is designated as the risk group.
[0024] On the other hand, the group 304 to which the sample data belongs may be determined by performing supervised learning using the actual measurement value 332 as a reference, with a pH value of 7.2, and assigning a correct answer label indicating safety if the actual measurement value 332 is pH>7.2, and a correct answer label indicating danger if the actual measurement value 332 is pH≦7.2.
[0025] The annotation 305 is an evaluation index for the sample fetus, and is given by a doctor or as a collation result by a collation unit 704, which will be described later.
[0026] The non-learning target flag 306 is a flag for setting the sample data as non-learning target. "0" indicates a learning target and is the default value, and "1" indicates non-learning target. Note that the matching device may delete the entry of sample data that is not to be learned without using the non-learning target flag 306.
[0027] <Feature 302> Fig. 4 is an explanatory diagram showing an example 1 of measurement factors included in the feature 302. Fig. 4 shows a histogram 400 obtained from the waveform of a CTG. Factors of the feature 302 obtained from the histogram 400 include, for example, the mode, the mean, the median, the variance, the histogram width, the low frequency, and the tendency.
[0028] FIG. 5 is an explanatory diagram showing an example 2 of measurement factors included in the feature 302. FIG. 5 shows a CTG waveform. The waveform includes a heart rate waveform 501 and a labor pain intensity waveform 502. Factors obtained from the heart rate waveform 501 include, for example, the baseline, waveform maximum value, waveform minimum value, waveform mean value, and waveform variance. Factors obtained from the labor pain intensity waveform 502 include, for example, the proportion of labor pains (total labor time / measurement period). Factors obtained as medical knowledge include, for example, the number of times the fetal heart rate drops below 80 bpm, the number of times the fetal heart rate drops by 30 bpm or more from the baseline, the duration that the fetal heart rate drops below 80 bpm, the duration that the fetal heart rate drops by 30 bpm or more from the baseline, and the time from the lowest fetal heart rate to birth.
[0029] <Basic information DB> 6 is an explanatory diagram showing an example of a basic information DB 600. The basic information DB 600 is stored in the storage device 202 or in a storage device of another computer accessible to the collation device.
[0030] The basic information DB 600 stores data used for cleansing sample data. Specifically, the basic information DB 600 includes, for example, conditions 601 related to a basic factor group and data quality conditions 602. The conditions 601 related to the basic factor group are conditions under which a basic factor group such as the age of the sample pregnant woman, the number of weeks of pregnancy, the method of delivery, the number of fetuses, fetal developmental delay, medication information, and smoking are excluded from learning. For example, in the case of age, sample data of a person 45 years of age or older is set as an excluded subject from learning. In addition, in the case of the number of weeks of pregnancy, sample data of less than 27 weeks or 43 weeks or more is set as an excluded subject from learning. In this way, conditions for excluding a sample from learning are set for each factor.
[0031] Unlike the condition 601 related to the basic factor group, the data quality condition 602 is a condition related to the quality of the sample data values, and for example, sample data with 20% or more missing values is set as ineligible for learning. Note that ineligible for learning may mean that the sample data itself is simply left in the learning DB 100 if it is not used for learning, or may be deleted from the learning DB 100. Note that even if it is deleted from the learning DB 100, it may remain in the storage device 202.
[0032] <Example of functional configuration of verification device> 7 is a block diagram showing an example of the functional configuration of the matching device 200. The matching device 200 can access the learning DB 100 and the basic information DB 600. The matching device 200 also includes a preprocessing unit 701, a grouping processing unit 702, a calculation unit 703, a matching unit 704, an output unit 705, and a learning unit 706. Specifically, the preprocessing unit 701, the grouping processing unit 702, the calculation unit 703, the matching unit 704, the output unit 705, and the learning unit 706 are realized, for example, by having the processor 201 execute a program stored in the storage device 202 shown in FIG. 2.
[0033] The preprocessing unit 701 refers to the basic information DB 600, cleanses the sample data group in the learning DB 100, and excludes unnecessary sample data from the learning target.
[0034] The grouping processing unit 702 groups the sample data in the learning DB 100 into a safe group and a risk group. Specifically, for example, as described above, the grouping processing unit 702 performs unsupervised learning to group sample data whose feature quantities 302 (specifically, for example, measurement factor groups) are close to each other, and continues grouping until the data finally converges to two groups. The final two groups are the safe group and the risk group, with the larger number of sample data being the safe group and the smaller number being the risk group.
[0035] The grouping processing unit 702 also uses the combination of the feature 302 (specifically, the group of measurement factors) and the correct label based on the measured value 332 of pH 303 (safe if pH>7.2, dangerous if pH≦7.2) as training data, and classifies the data into a safe group and a dangerous group through supervised learning. Because the measured value 332 of pH 303 is used, the sample data to be grouped is sample data of pregnant women who have already given birth.
[0036] The calculation unit 703 calculates the predicted value 331 of pH303 by inputting a specific group of measurement factors (a group of factors obtained as medical knowledge) of sample data (sample data to be predicted) for which the predicted value 331 of pH303 has not been calculated into the learning model generated by the learning unit 706.
[0037] 1, the collation unit 704 collates the safe value (pH>7.2) or dangerous value (pH≦7.2) obtained from the pH 303 for each piece of sample data with the annotation 305. For the pH 303, the actual measured value 332 is used for sample data with the actual measured value 332, and the predicted value 331 is used for sample data without the actual measured value 332.
[0038] Before the collation by the collation unit 704 , the annotation 305 is an annotation by a doctor, and after the collation by the collation unit 704 , the annotation 305 becomes the collation result by the collation unit 704 .
[0039] The output unit 705 displays the preprocessing results from the preprocessing unit 701 and the matching results from the matching unit 704. Specifically, for example, the output unit 705 displays the preprocessing results and the matching results on a display device, which is an example of the output device 204, or transmits the preprocessing results and the matching results to another computer via the communication IF 206 to display them on the other computer.
[0040] The learning unit 706 generates a multiple regression model as a learning model using a group of factors obtained as medical knowledge for each sample data in the feature amount 302 and the measured value 332 of pH 303.
[0041] <Matching process procedure> Next, an example of a procedure for the matching process performed by the matching device 200 will be explained for each function. The matching device 200 performs learning of a learning model and prediction of pH303 using the learning model, and the processing by each function described below will be explained by indicating whether it is applied to learning or prediction.
[0042] [Example of pre-processing procedure] FIG. 8 is a flowchart showing an example of a preprocessing procedure performed by the preprocessing unit 701. The preprocessing procedure is performed in both learning and prediction. The preprocessing unit 701 determines whether there is unselected sample data in the learning DB 100 (step S801). If there is unselected sample data (step S801: Yes), the preprocessing unit 701 selects one unselected sample data (step S802). The preprocessing unit 701 performs preprocessing (cleansing) on the selected sample data using the basic information DB 600, that is, performs a process of determining whether the selected sample data satisfies the conditions 601 related to the basic factor group and the data quality conditions 602 (step S802).
[0043] Then, the preprocessing unit 701 outputs the selection screen so that it can be displayed (step S804). Specifically, for example, the preprocessing unit 701 displays the selection screen on a display device, which is an example of the output device 204, or on another computer operated by the user.
[0044] 9 is an explanatory diagram showing an example of a selection screen display during learning. A selection screen 900 displays a first radio button 901, a second radio button 902, an execute button 903, a first character string 910, and a second character string 920 as preprocessing results.
[0045] The first radio button 901 is a selection button that the user uses to select the preprocessed selected sample data as a learning target, and the reason for this is displayed as a first character string 910. The "all conditions" in the first character string 910 refers to the conditions 601 related to the basic factor group in the basic information DB 600 and the data quality conditions 602. In other words, this indicates that the selected sample data is a learning target.
[0046] The second radio button 902 is a selection button that the user uses to exclude the preprocessed selected sample data from the learning target, and the reason for this is displayed as a second character string 920. The second character string 920 is, for example, one of the conditions 601 related to the basic factor group and the data quality condition 602 in the basic information DB 600 that the selected sample data meets.
[0047] The execute button 903 is a button for confirming the selection of either the first radio button 901 or the second radio button 902, and the pre-processing unit 701 receives a signal indicating whether the selected sample data has been selected as a learning object or not as a learning object.
[0048] 10 is an explanatory diagram showing an example of a selection screen display during prediction. A selection screen 1000 displays a first radio button 1001, a second radio button 1002, an execute button 1003, a first character string 1010, and a second character string 1020 as preprocessing results.
[0049] A first radio button 1001 is a selection button that allows the user to select preprocessed selected sample data as a learning target, and the reason for this is displayed as a first character string 1010. "All conditions" in the first character string 1010 refers to the conditions 601 related to the basic factor group in the basic information DB 600 and the data quality conditions 602. In other words, this indicates that the selected sample data is a learning target.
[0050] Second radio button 1002 is a selection button for the user to exclude the preprocessed selected sample data from the learning target and to reconfirm the attachment of the measuring device, and the message urging the user to reconfirm is displayed as second character string 1020. Execute button 1003 is a button for confirming the selection of either first radio button 1001 or second radio button 1002, and preprocessing unit 701 receives a signal indicating whether the selected sample data has been selected as a learning target or as an exemption from learning.
[0051] 8, when the execute button 903 or the execute button 1003 is pressed, the preprocessing unit 701 receives a selection signal (step S805). If the selection signal indicates a learning target (step S806: Yes), the process returns to step S801. On the other hand, if the selection signal indicates a non-learning target (step S806: No), the preprocessing unit 701 sets the non-learning target flag 306 of the selected sample data to non-learning target (step S807) and returns to step S801. If there is no unselected sample data in step S801, the preprocessing unit 701 ends preprocessing.
[0052] [Example of matching process procedure during learning] 11 is a flowchart showing an example of a procedure for a matching process during learning. The grouping processing unit 702 selects a factor of a feature amount (step S1101). Specifically, for example, the grouping processing unit 702 outputs a feature amount selection screen so that it can be displayed when matching starts.
[0053] 12 is an explanatory diagram showing an example of a feature selection screen. The feature selection screen 1200 has an automatic selection button 1201, a number selection button 1202, a factor group list 1203, and an execute button 1204. The automatic selection button 1201 is a button for the matching device 200 to automatically select a factor of the feature 302. The number selection button 1202 is a button for selecting any one or more factors from the factor group in the factor group list 1203. The user selects the radio button (shown as a circle) for each factor on the left side of the factor group list 1203. When the execute button 1204 is pressed, the grouping processing unit 702 receives a selection method signal indicating automatic selection or number selection (including the selection number).
[0054] 11, upon receiving the selection method signal, the grouping processing unit 702 selects a factor of the feature amount. Specifically, for example, in the case of automatic selection, the grouping processing unit 702 automatically selects a factor of the feature amount 302 by a T-test or principal component analysis. In addition, in the case of number selection, the grouping processing unit 702 selects a factor selected by number.
[0055] The grouping processing unit 702 performs grouping processing on the sample data group in the learning DB 100 using the factors selected in step S1101 as the features 302 (step S1102). The sample data that is the target of the grouping processing (step S1102) is sample data that has the measured value 332 of pH 303 and whose non-learning target flag 306 is "0" (learning target). This prevents bad data from being applied to the grouping processing (step S1102), improving the accuracy of the grouping processing (step S1102) during learning.
[0056] Furthermore, during learning, in the grouping process (step S1102), the grouping processing unit 702 may perform supervised learning using the selection factors of the sample data and the correct answer data (the correct answer label obtained from the measured value 332 of pH 303) as a learning data set, or may perform unsupervised learning using only the selection factors of the sample data. In either case, the grouping processing unit 702 groups the sample data group into a safe group and a dangerous group, and sets the belonging group 304.
[0057] Next, as shown in FIG. 1, the matching unit 704 matches the classification (safe or dangerous) of the measured value 332 of pH 303 with the classification (safe or dangerous) of the group 304 to which the sample data belongs (step S1103). The matching unit 704 assigns the matching results 101 to 104 as annotations 305 (step S1104), and updates the non-learning target flags 306 based on the matching results 101 to 104 (step S1105). Specifically, for example, the matching unit 704 updates the non-learning target flags 306 to "1" for the sample data of the matching results 101 to 104. The output unit 705 then outputs the matching results 101 to 104 in a displayable manner (step S1106).
[0058] 13 to 16 are explanatory diagrams showing examples of matching result display screens. Matching result display screen 1300 displays matching result 101, matching result display screen 1400 displays matching result 102, matching result display screen 1500 displays matching result 103, and matching result display screen 1600 displays matching result 104.
[0059] 11, the learning unit 706 performs learning using a learning data set that is a combination of each selection factor of the sample data group of the matching result 101 and the measured value 332 of pH 303, and generates a learning model (for example, a multiple regression model) (step S1107). This completes the matching process during learning by the grouping processing unit 702, the matching unit 704, the output unit 705, and the learning unit 706.
[0060] [Example of matching process procedure during prediction] 17 is a flowchart showing an example of a matching process procedure during prediction. As in step S1101, the grouping processing unit 702 selects a factor of the feature (step S1701). The grouping processing unit 702 performs grouping processing on a group of sample data in the learning DB 100 using the factor selected in step S1701 as the feature 302 (step S1702).
[0061] The sample data to be subjected to the grouping process (step S1702) is sample data that does not have the measured value 332 of pH 303 and whose non-learning target flag 306 is "0" (learning target). This prevents bad data from being applied to the grouping process (step S1102), improving the accuracy of the grouping process (step S1702) during prediction.
[0062] Also, during learning, in the grouping process (step S1702), the grouping processing unit 702 may perform unsupervised learning using only the selection factors of the sample data, or may perform supervised learning using the selection factors of the sample data and the correct answer data as the learning dataset.
[0063] In the case of supervised learning, the correct answer data is, for example, the classification of the annotation 305 assigned in step S1104 (the correct answer label for "both safe" is "safe", and the correct answer label for "conflict" and "both dangerous" is "dangerous"). In either case, the grouping processing unit 702 groups the sample data group into a safe group and a dangerous group, and sets the group to which the data belongs 304. By performing supervised learning using the correct answer data, the matching result 101 assigned as the annotation 305 in step S1104 can be reflected in the learning model.
[0064] Next, the calculation unit 703 inputs the values of the selection factors of the sample data into the learning model to calculate the predicted value 331 of the pH 303 of the sample data, and stores it in the learning DB (step S1703).
[0065] 1, the matching unit 704 matches the classification (safe or dangerous) of the predicted value 331 of pH 303 with the classification (safe or dangerous) of the group 304 to which the sample data belongs (step S1704). The matching unit 704 assigns the matching results 101 to 104 as annotations 305 (step S1705), and updates the non-learning target flag 306 based on the matching results 101 to 104 (step S1706).
[0066] Specifically, for example, the matching unit 704 updates the non-learning target flag 306 to "1" for the sample data of the matching results 101 to 104. Then, the matching unit 704 outputs the matching results 101 to 104 in a displayable manner (step S1707), as shown in FIGS. 14 to 17. This completes the matching process during prediction by the grouping processing unit 702, the calculation unit 703, the matching unit 704, and the output unit 705.
[0067] Note that even for a sample that does not have a measured value 332 of pH 303 before delivery, the measured value 332 of pH 303 can be obtained after delivery. In this case, the collation device can update the annotation 305 by executing the collation process during learning shown in FIG. 11 for the measured value 332 of pH 303. The collation device also retrains the learning model by backpropagating errors based on the difference between the predicted value 331 and the measured value 332 of pH 303. This can improve the accuracy of the learning model.
[0068] As described above, the matching results 101 to 104 indicating "safe," "inconsistent," and "dangerous" are automatically assigned as annotation 305, thereby improving the reliability of annotation 305. Furthermore, sample data to which the matching results 102 to 104 indicating "inconsistent" or "dangerous" are assigned as annotation 305 are excluded from learning, leaving only sample data to which the matching result 101 indicating "safe" is assigned as annotation 305. Therefore, automatic organization of the learning DB 100 can be achieved. Furthermore, by generating a learning model using the remaining sample data group, the predicted value 331 of pH 303 can be made more accurate. Therefore, fetal condition can be predicted without incurring data processing costs.
[0069] In the above-described embodiment, the samples are divided into two groups, safe and dangerous, but the classification of the groups is not limited to safe and dangerous, and may be any classification of evaluations that express the degree of accuracy, reliability, etc. In the above-described embodiment, the samples are pregnant women, but the samples are not limited to pregnant women and may be various measurement targets such as people and devices.
[0070] The present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the spirit and scope of the appended claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to configurations including all of the described configurations. Furthermore, part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Furthermore, the configuration of another embodiment may be added to the configuration of one embodiment. Furthermore, part of the configuration of each embodiment may be added to, deleted from, or replaced with other configurations.
[0071] Furthermore, the aforementioned configurations, functions, processing units, processing means, etc. may be realized in part or in whole in hardware, for example by designing them as integrated circuits, or may be realized in software by having a processor interpret and execute a program that realizes each function.
[0072] Information such as programs, tables, files, etc. that realize each function can be stored in storage devices such as memory, hard disks, SSDs (Solid State Drives), or recording media such as IC (Integrated Circuit) cards, SD cards, and DVDs (Digital Versatile Discs).
[0073] In addition, the control lines and information lines shown are those that are considered necessary for explanation, and do not necessarily represent all the control lines and information lines that are necessary for implementation. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]
[0074] 101~104 Matching results 200 Collation Device 201 processor 202 Storage Devices 302 Features 304 Group Affiliation 305 Annotations 306 Not applicable to learning 331 predicted value 332 actual measurements 701 Pre-processing section 702 Grouping processing unit 703 Calculation Unit 704 Matching Unit 705 Output Section 706 Learning Department
Claims
1. a grouping processing unit that groups the sample group into a first group indicating a first classification and a second group indicating a second classification that is lower in evaluation than the first classification, based on each explanatory variable of the sample group; a matching unit that matches the classification of the first group and the second group grouped by the grouping processing unit with a classification specified by each objective variable of the sample group, and if the classification specified by the objective variable of a first sample belonging to the first group of the sample group is the first classification, assigns an annotation indicating the first classification to the first sample; an output unit that outputs a collation result by the collation unit; a learning unit that performs learning based on explanatory variables of a first sample to which an annotation indicating the first classification is added and an annotation indicating the first classification that serves as a target variable, and generates a learning model; a calculation unit that calculates a predicted value of each dependent variable of the sample group to be predicted by inputting each explanatory variable of the sample group to be predicted into the learning model generated by the learning unit, the grouping processing unit groups the group of samples to be predicted into the first group and the second group based on each explanatory variable of the group of samples to be predicted; the comparison unit compares the classifications of the first group and the second group into which the group of samples to be predicted is grouped by the grouping processing unit with a classification specified by the predicted value calculated by the calculation unit, and if the classification specified by the predicted value of a first sample to be predicted that belongs to the first group among the group of samples to be predicted is the second classification, assigns an annotation indicating a contradiction to the first sample to be predicted. A collation device characterized by:
2. A grouping processing unit that groups the sample group into a first group indicating a first classification and a second group indicating a second classification that is lower in evaluation than the first classification based on each explanatory variable of the sample group; a matching unit that matches the classification of the first group and the second group grouped by the grouping processing unit with a classification specified by each objective variable of the sample group, and if the classification specified by the objective variable of a first sample belonging to the first group of the sample group is the first classification, assigns an annotation indicating the first classification to the first sample; an output unit that outputs a collation result by the collation unit; a learning unit that performs learning based on explanatory variables of a first sample to which an annotation indicating the first classification is added and an annotation indicating the first classification that serves as a target variable, and generates a learning model; a calculation unit that calculates a predicted value of each dependent variable of the sample group to be predicted by inputting each explanatory variable of the sample group to be predicted into the learning model generated by the learning unit, the grouping processing unit groups the group of samples to be predicted into the first group and the second group based on each explanatory variable of the group of samples to be predicted; the comparison unit compares the classifications of the first group and the second group into which the group of samples to be predicted is grouped by the grouping processing unit with a classification specified by the predicted value calculated by the calculation unit, and if the classification specified by the predicted value of a second sample to be predicted that belongs to the second group of the group of samples to be predicted is the first classification, assigns an annotation indicating a contradiction to the second sample to be predicted. A collation device characterized by:
3. 3. The verification device according to claim 1, the collation unit assigns an annotation indicating a contradiction to a first sample when a classification specified by a dependent variable of a first sample belonging to the first group among the sample group is the second classification; A collation device characterized by:
4. 3. The verification device according to claim 1, When a classification specified by a dependent variable of a second sample belonging to the second group among the sample group is the second classification, the matching unit assigns an annotation indicating the second classification to the second sample. A collation device characterized by:
5. 3. The verification device according to claim 1, the collation unit assigns an annotation indicating a contradiction to a second sample when a classification specified by a dependent variable of a second sample belonging to the second group among the sample group is the first classification; A collation device characterized by:
6. 3. The verification device according to claim 1, the grouping processing unit performs clustering by unsupervised learning using the explanatory variables to group the sample group into the first group and the second group; A collation device characterized by:
7. 3. The verification device according to claim 1, the grouping processing unit groups the sample group into the first group and the second group based on supervised learning using the explanatory variables and a classification identified by the actual measured value of the objective variable; A collation device characterized by:
8. A collation method using a collation device having a processor that executes a program and a storage device that stores the program, comprising: the processor: a first grouping process for grouping the sample group into a first group indicating a first classification and a second group indicating a second classification having a lower evaluation than the first classification, based on each explanatory variable of the sample group; a first matching process of matching the classifications of the first group and the second group obtained by the first grouping process with classifications specified by the objective variables of each of the sample groups, and, if the classification specified by the objective variables of a first sample belonging to the first group among the sample groups is the first classification, adding an annotation indicating the first classification to the first sample; an output process for outputting a collation result obtained by the first collation process; a learning process for generating a learning model by learning based on explanatory variables of a first sample to which annotations indicating the first classification are added and the annotations indicating the first classification that serve as a target variable; a calculation process of calculating a predicted value of each objective variable of a group of samples to be predicted by inputting each explanatory variable of the group of samples to be predicted into a learning model generated by the learning process; a second grouping process of grouping the group of samples to be predicted into the first group and the second group based on each explanatory variable of the group of samples to be predicted; a second comparison process for comparing the classifications of the first group and the second group into which the group of samples to be predicted is divided by the second grouping process with a classification specified by the predicted value calculated by the calculation process, and for adding an annotation indicating a contradiction to the first predicted sample if the classification specified by the predicted value of a first predicted sample belonging to the first group among the group of samples to be predicted is the second classification; A matching method characterized by performing the following steps.
9. A matching method using a matching device having a processor that executes a program and a storage device that stores the program, comprising: the processor: a first grouping process for grouping the sample group into a first group indicating a first classification and a second group indicating a second classification having a lower evaluation than the first classification, based on each explanatory variable of the sample group; a first matching process of matching the classifications of the first group and the second group obtained by the first grouping process with classifications specified by the objective variables of each of the sample groups, and, if the classification specified by the objective variables of a first sample belonging to the first group among the sample groups is the first classification, adding an annotation indicating the first classification to the first sample; an output process for outputting a collation result obtained by the first collation process; a learning process for generating a learning model by learning based on explanatory variables of a first sample to which annotations indicating the first classification are added and the annotations indicating the first classification that serve as a target variable; a calculation process of calculating a predicted value of each objective variable of a group of samples to be predicted by inputting each explanatory variable of the group of samples to be predicted into a learning model generated by the learning process; a second grouping process of grouping the group of samples to be predicted into the first group and the second group based on each explanatory variable of the group of samples to be predicted; a second matching process for matching the classifications of the first group and the second group into which the group of samples to be predicted is grouped by the second grouping process with a classification specified by the predicted value calculated by the calculation process, and for adding an annotation indicating a contradiction to the second predicted sample if the classification specified by the predicted value of a second predicted sample belonging to the second group of the group of samples to be predicted is the first classification; A matching method characterized by performing the following steps.
10. A processor, a first grouping process for grouping the sample group into a first group indicating a first classification and a second group indicating a second classification having a lower evaluation than the first classification, based on each explanatory variable of the sample group; a first matching process of matching the classifications of the first group and the second group obtained by the first grouping process with classifications specified by the objective variables of each of the sample groups, and, if the classification specified by the objective variables of a first sample belonging to the first group among the sample groups is the first classification, adding an annotation indicating the first classification to the first sample; an output process for outputting a collation result obtained by the first collation process; a learning process for generating a learning model by learning based on explanatory variables of a first sample to which annotations indicating the first classification are added and the annotations indicating the first classification that serve as a target variable; a calculation process of calculating a predicted value of each objective variable of a group of samples to be predicted by inputting each explanatory variable of the group of samples to be predicted into a learning model generated by the learning process; a second grouping process of grouping the group of samples to be predicted into the first group and the second group based on each explanatory variable of the group of samples to be predicted; a second comparison process for comparing the classifications of the first group and the second group into which the group of samples to be predicted is divided by the second grouping process with a classification specified by the predicted value calculated by the calculation process, and for adding an annotation indicating a contradiction to the first predicted sample if the classification specified by the predicted value of a first predicted sample belonging to the first group among the group of samples to be predicted is the second classification; A matching program characterized by executing the above.
11. A processor, a first grouping process for grouping the sample group into a first group indicating a first classification and a second group indicating a second classification having a lower evaluation than the first classification, based on each explanatory variable of the sample group; a first matching process of matching the classifications of the first group and the second group obtained by the first grouping process with classifications specified by the objective variables of each of the sample groups, and, if the classification specified by the objective variables of a first sample belonging to the first group among the sample groups is the first classification, adding an annotation indicating the first classification to the first sample; an output process for outputting a collation result obtained by the first collation process; a learning process for generating a learning model by learning based on explanatory variables of a first sample to which annotations indicating the first classification are added and the annotations indicating the first classification that serve as a target variable; a calculation process of calculating a predicted value of each objective variable of a group of samples to be predicted by inputting each explanatory variable of the group of samples to be predicted into a learning model generated by the learning process; a second grouping process of grouping the group of samples to be predicted into the first group and the second group based on each explanatory variable of the group of samples to be predicted; a second matching process for matching the classifications of the first group and the second group into which the group of samples to be predicted is grouped by the second grouping process with a classification specified by the predicted value calculated by the calculation process, and for adding an annotation indicating a contradiction to the second predicted sample if the classification specified by the predicted value of a second predicted sample belonging to the second group of the group of samples to be predicted is the first classification; A matching program characterized by executing the above.
Citation Information
Patent Citations
Data screening method and device, storage medium and electronic equipment
CN111046969A
Image recognition method and device, computer equipment and storage medium
CN113392867A
Learning dataset generation system, learning server, and learning dataset generation program
JP2020204800A
Method and system for detecting anomalies in data labels
US20190205794A1
Data processing device, data processing method, program, and integrated circuit
WO2010125781A1