Machine learning data processing method and electronic device using same
By balancing the number of sources, regularizing processing and screening of machine learning data, the problem of degradation in prediction accuracy caused by different source data sets is solved, and the prediction accuracy and data adaptability of the model are improved.
Patent Information
- Application Number
- CN202410236681.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-01
- Publication Date
- 2025-09-02
AI Technical Summary
In machine learning, the use of multiple inappropriate detection items will reduce prediction accuracy, especially when data sets from different sources vary greatly, such as test data from different hospitals, resulting in poor prediction results in model prediction.
By balancing the source quantity, regularizing the subject and regularizing the source of the original measurement data, generating balanced allocation charts and regularizing measurement data, and through segmentation and sampling analysis of the prediction capabilities of the detection items, appropriate detection items are selected for machine learning model training and prediction.
It improves the prediction accuracy of machine learning models and improves the amplification of data, allowing the model to better adapt to data from different sources, and improves the accuracy of training and prediction.
Smart Images

Figure CN120579653A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a processing method and an electronic device using the same, and in particular to a data processing method for machine learning and an electronic device using the same. Background Art
[0002] In machine learning, using more test items for modeling generally improves prediction accuracy. However, adding inappropriate test items can actually reduce prediction accuracy. Therefore, careful selection and decision-making of test items are necessary for machine learning models to generate modeling and predictive inferences.
[0003] Furthermore, in real-world applications, data comes from a variety of sources. For example, in the medical field, common source differences can arise from differences in testing machine types or hospitals. Such source variations can significantly reduce the effectiveness of machine learning in modeling and inference. For example, for a blood test, different hospitals may use different testing machines or different numerical units. Blood values on different testing machines may have different ranges: 20 to 2000 on machine A, 100 to 10,000 on machine B. Obviously, because the values on different machines represent different meanings, simply training a model using data from both ranges will result in poor predictions. Summary of the Invention
[0004] This invention relates to a data processing method for machine learning and an electronic device employing the method. During the screening process of raw measurement data for test items, the method appropriately processes raw measurement data from various sources, thereby improving the prediction accuracy of the machine learning model. Furthermore, the raw measurement data, after quantitative balance and numerical normalization, also exhibits good scalability, facilitating the training or correction of the machine learning model.
[0005] According to one aspect of the present invention, a data processing method for machine learning is provided. The data processing method for machine learning includes the following steps. A source balancing procedure is performed on raw measurement data for multiple sources to obtain a balanced distribution diagram. In the balanced distribution diagram, the raw measurement data includes multiple test values corresponding to multiple test items for multiple subjects, corresponding to the same number of different sources. A subject normalization procedure is performed on these test values for each subject to obtain subject normalized measurement data. In the subject normalized measurement data, these test values for each subject are scaled to the same numerical range. A source scaling procedure is performed on these test values for each source to obtain source normalized measurement data. In the source normalized measurement data, these test values for each source are scaled to the same numerical range. The balanced distribution diagram, the subject normalized measurement data, and the source normalized measurement data are merged to obtain balanced subject normalized data and balanced source normalized data. The balanced subject-normalized data and the balanced source-normalized data are divided into several splits. Each split corresponds to all sources. Each split is sampled and analyzed to generate a predictive power table. The predictive power table includes a predictive power for each test item. Based on the predictive power table, a subset of test items are output. The output test items are used by a machine learning model for modeling, training, or predictive inference.
[0006] According to another aspect of the present invention, an electronic device is provided. The electronic device includes a source quantity balancing unit, a subject normalization unit, a source normalization unit, a merging unit, and an extraction unit. The source quantity balancing unit is used to perform a source quantity balancing procedure on raw measurement data for multiple sources to obtain a balanced distribution diagram. In the balanced distribution diagram, the quantities corresponding to different sources are the same. The raw measurement data includes multiple test values corresponding to multiple test items for multiple subjects. The subject normalization unit is used to perform a subject normalization procedure on these test values for each subject to obtain subject normalized measurement data. In the subject normalized measurement data, these test values for each subject are scaled to the same numerical range. The source normalization unit is used to perform a source normalization procedure on these test values for each source to obtain source normalized measurement data. In the source normalized measurement data, these test values for each source are scaled to the same numerical range. The merging unit is used to merge the balanced distribution map, the subject normalized measurement data and the source normalized measurement data to obtain a balanced subject normalized data and a balanced source normalized data. The extraction unit includes a splitter, a calculator and a selector. The splitter is used to cut the balanced subject normalized data and the balanced source normalized data into several splits. Each split corresponds to all sources. The calculator is used to sample each split and analyze a prediction capability table. The prediction capability table includes a prediction capability for each test item. The selector is used to output part of the test items based on the prediction capability table. The output test items are used for modeling, training or predictive inference by a machine learning model.
[0007] In order to better understand the above and other aspects of the present invention, embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 This is a data processing process of machine learning in one embodiment.
[0009] Figure 2 FIG. 1 is a schematic diagram of an electronic device according to an embodiment.
[0010] Figure 3 The present invention is a flowchart of a data processing method for machine learning according to an embodiment.
[0011] Figure 4 yes Figure 3 An example of how machine learning data processing methods work.
[0012] Figure 5 This is an example of the operation of the extraction unit.
[0013] Figure 6 This is step S151 of an embodiment.
[0014] Reference numerals:
[0015] 100: Electronic devices
[0016] 110: Source Quantity Balance Unit
[0017] 120: Subject Normalization Unit
[0018] 130: Source Normalization Unit
[0019] 140:Merge unit
[0020] 150: Extraction unit
[0021] 151: Splitter
[0022] 152:Calculator
[0023] 153:Selector
[0024] 160: Training Unit
[0025] DM: Agency
[0026] DT0: original measurement data
[0027] DT1: Normalized measurement data of subjects
[0028] DT1': Balanced subject normalization data
[0029] DT2: Source normalized measurement data
[0030] DT2': Balanced Source Normalized Data
[0031] EP: Instrument
[0032] IT,IT*:Test items
[0033] MD: Machine Learning Model
[0034] MP: Balanced Distribution Graph
[0035] P1: Source Quantity Balancing Procedure
[0036] P2: Subject regularization procedure
[0037] P3: Source Regularization Procedure
[0038] S110, S120, S130, S140, S151, S152, S153: Steps
[0039] SP: Segmentation
[0040] SR: Source
[0041] TB: Predictive Power Table
[0042] US: Subject
[0043] VL: Detection value DETAILED DESCRIPTION
[0044] The technical terms used in this specification are based on customary terms in the technical field. If some terms are explained or defined in this specification, the interpretation of these terms shall be based on the explanations or definitions in this specification. Each embodiment of the present invention has one or more technical features. Under the premise of possible implementation, those with ordinary knowledge in this technical field may selectively implement some or all of the technical features in any embodiment, or selectively combine some or all of the technical features in these embodiments.
[0045] Please refer to Figure 1 , which illustrates the data processing process for machine learning according to one embodiment of the present invention. In machine learning, using more test items IT to build a machine learning model MD generally improves the prediction accuracy of the machine learning model MD. However, adding inappropriate test items IT can actually reduce prediction accuracy. Therefore, it is necessary to screen the test items IT for the machine learning model MD to perform modeling and predictive inference.
[0046] When building a machine learning model (MD), a large amount of data is required for training. To increase the amount of raw measurement data (DT0), it is often necessary to obtain raw measurement data (DT0) from multiple sources (SR). These sources (SR) can be, for example, different institutions (DM) (e.g., different hospitals, different research institutes) or different instruments (EP).
[0047] However, the original measurement data DT0 from different sources SR may have different value ranges. If the original measurement data DT0 is directly used to train the machine learning model MD, the prediction effect of the machine learning model MD will be poor.
[0048] Please refer to Figure 2, which illustrates a schematic diagram of an electronic device 100 according to an embodiment of the present invention. The electronic device 100 includes a source quantity balancing unit 110, a subject normalization unit 120, a source normalization unit 130, a merging unit 140, an extraction unit 150, and a training unit 160. The extraction unit 150 includes a splitter 151, a calculator 152, and a selector 153. The source quantity balancing unit 110, the subject normalization unit 120, the source normalization unit 130, the merging unit 140, the extraction unit 150, and / or the training unit 160 are configured to execute various control, processing, and analysis procedures and may be, for example, a circuit, a circuit board, a storage device storing program code, or a chip. The chip may be, for example, a central processing unit (CPU), or other programmable general-purpose or special-purpose microcontrol unit (MCU), microprocessor, digital signal processor (DSP), programmable controller, application specific integrated circuit (ASIC), graphics processing unit (GPU), image signal processor (ISP), image processing unit (IPU), arithmetic logic unit (ALU), complex programmable logic device (CPLD), field programmable gate array (FPGA), or other similar components or combinations of the above components.
[0049] In this embodiment, during the screening of the detection items IT in the raw measurement data DT0, the raw measurement data DT0, which includes data from different sources SR, undergoes appropriate quantitative balancing and numerical normalization, thereby improving the prediction accuracy of the machine learning model MD. This quantitative balancing and numerical normalization of the raw measurement data DT0 enhances scalability, facilitating the further training or correction of the machine learning model MD.
[0050] Please refer to Figures 2 to 3 , Figure 3 is a flowchart of a data processing method for machine learning according to an embodiment. Figure 4 yes Figure 3The machine learning data processing method operates. In step S110, the source balancing unit 110 performs a source balancing procedure P1 on the raw measurement data DT0 for all source SRs to obtain a balanced distribution map MP. Table 1 below illustrates various source SRs for the raw measurement data DT0.
[0051]
[0052]
[0053] Table 1 (Various sources SR of raw measurement data DT0)
[0054] The original measurement data DT0 includes the subject US and the source SR. In the original measurement data DT0 in Table 1, there are three source SRs from "1", three source SRs from "2", and four source SRs from "3". In order to make the numbers corresponding to all source SRs the same, an upsampling method can be used to increase the data volume of the original measurement data DT0. For example, as shown in Table 2, the subject US of "C" (source SR from "1") is repeatedly sampled, and the subject US of "E" (source SR from "2") is repeatedly sampled, so that the source SRs from "1", "2", and "3" are all four, to obtain a balanced distribution map MP. In the balanced distribution map MP, the numbers corresponding to these source SRs are the same.
[0055] In addition, weight adjustment can also be used to reduce the problem of the same subject US being repeated too many times. For example, a subject US that is repeatedly sampled a times can be given a weight of 1 / a.
[0056] Subject US Source SR First 1 Second 1 C 1 Man 2 Wu 2 Self 2 Geng 3 pungent 3 the ninth of the ten Heavenly Stems 3 Gui 3 C 1 Wu 2
[0057] Table 2 (Balanced Distribution Map MP)
[0058] Table 3 below illustrates an example of raw measurement data DT0. The raw measurement data DT0 includes a number of test values VL (e.g., "325," "270," ...) corresponding to a number of test items IT (e.g., "Protein Value 1," "Protein Value 2," ...) for a number of subjects US (e.g., "A," "B," ...).
[0059]
[0060]
[0061] Table 3 (Original measurement data DT0)
[0062] Next, in step S120, the subject normalization unit 120 performs a personalization scaling procedure P2 on the detection values VL for each subject US to obtain subject normalized measurement data DT1. Please refer to Table 4 for an example of the subject normalized measurement data DT1.
[0063]
[0064] Table 4 (Normalized measurement data DT1 of the subjects)
[0065] As shown in Table 4, in the subject-normalized measurement data DT1, the measured values VL for each subject US are scaled to the same numerical range. For example, among the measured values VL of "325," "270," and "200" for subject "A," "325" is normalized to "1," "200" is normalized to "0," and "270" is normalized proportionally to "0.56." Among the measured values VL of "155," "810," and "310" for subject "B," "810" is normalized to "1," "155" is normalized to "0," and "310" is normalized proportionally to "0.24." Similarly, the maximum and minimum measured values VL for the same subject US are normalized to 1 and 0, respectively, and the remaining measured values VL are normalized proportionally. In this way, the differences between different test items IT of each subject US can be preserved, and the test values VL of each subject US are maintained at the same order of magnitude.
[0066] Then, in step S130, the source normalization unit 130 performs a source normalization procedure P3 on the detection values VL for each source SR to obtain source normalized measurement data DT2. Please refer to Table 5 for an example of the source normalized measurement data DT2.
[0067]
[0068] Table 5 (Source: Normalized Measurement Data DT2)
[0069] In the source-normalized measurement data DT2, the detection values VL of each source SR are scaled to the same numerical range. The source normalization unit 130 performs a z-score transformation on the detection values VL of the same source SR. For example, the detection values VL of "325," "155," and "160" from the source SR "1" have a mean of "213.3" and a standard deviation of "78.99." Subtracting the mean from "325" and dividing by the standard deviation normalizes the value to "1.41." Subtracting the mean from "155" and dividing by the standard deviation normalizes the value to "-0.74." Subtracting the mean from "160" and dividing by the standard deviation normalizes the value to "-0.67." Similarly, the detection values VL of "270," "30," and "265" from the source SR "2" also undergo a z-score transformation. In this way, the differences between each source SR and different subjects US can be preserved, and the detection values VL of each source SR are maintained at the same order of magnitude.
[0070] After the above steps, the balanced distribution diagram MP in Table 2, the subject normalized measurement data DT1 in Table 4, and the source normalized measurement data DT2 in Table 5 are obtained.
[0071] Then, in step S140, the merging unit 140 merges the balanced distribution map MP, the subject-normalized measurement data DT1, and the source-normalized measurement data DT2 to obtain balanced subject-normalized data DT1' and balanced source-normalized data DT2'. See Table 6 for an example of balanced subject-normalized data DT1' and balanced source-normalized data DT2'.
[0072]
[0073] Table 6 (Balanced Subject Normalized Data DT1' and Balanced Source Normalized Data DT2')
[0074] Then, proceed to steps S151 to S153. Figure 5 , which illustrates the operation of the extraction unit 150. The divider 151, calculator 152 and selector 153 of the extraction unit 150 execute steps S151 to S153 in sequence.
[0075] In step S151 , the splitter 151 of the extraction unit 150 splits the balanced subject normalized data DT1 ′ and the balanced source normalized data DT2 ′ into a plurality of splits SP, each of which corresponds to the entire source SR.
[0076] Please refer to Figure 6 , which illustrates step S151. Figure 6This example illustrates splitting into two separate SPs. These separate SPs have the same amount of data (e.g., 6 data sets each), and the union of these separate SPs covers all subjects US (e.g., subjects "A" through "Duke"). Furthermore, the subjects US in different separate SPs are not exactly the same.
[0077] Then, in step S152, the calculator 152 of the extraction unit 150 samples each segmented SP and analyzes it to generate a prediction capability table TB. Please refer to Table 7 for an example of the prediction capability table TB.
[0078]
[0079] Table 7 (Prediction Ability Table TB)
[0080] The prediction capability table TB includes a prediction capability for each test item IT. The calculator 152 uses random sampling to sample each segmented SP. For example, after calculating the prediction capability 10 times for each segmented SP, the average value is obtained to obtain the values in Table 7 above.
[0081] Next, in step S153, the selector 153 outputs some of the test items IT* according to the prediction capability table TB. Please refer to the following Table 8, which illustrates the predicted total scores of the test items IT.
[0082]
[0083] Table 8
[0084] Selector 153, for example, sums the predictive abilities of all segmented SPs to obtain a total predictive ability score and selects and outputs the top two test items IT* with the highest scores. Taking Table 8 above as an example, "Protein 1" from the balanced source normalized data DT2' and "Protein 3" from the balanced subject normalized data DT1' are output.
[0085] Please refer to Table 9 below, which illustrates the final output of “Protein 1” of the balanced source normalized data DT2 ′ and “Protein 3” of the balanced subject normalized data DT1 ′.
[0086]
[0087]
[0088] Table 9
[0089] These output detection items IT* are used for modeling, training or prediction inference by machine learning models MD. Figure 2As shown, the test item IT* can be provided to the training unit 160 to train a highly accurate machine learning model MD. In addition, when performing prediction inference, the test value of the test item IT* is input into the machine learning model MD to obtain a highly accurate prediction result.
[0090] Please refer to Table 10 below. The technology of the present invention has excellent scalability. Whenever new source data is available, existing models can be used for prediction or added directly to the training data without modifying the data format or adding new source fields. In a prediction experiment for physical cognitive decline syndrome (PCDS), the technology of the present invention can be used to combine multiple source data to train the model. Compared with models trained with single source data, the technology of the present invention has higher accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and area under the receiver operating characteristic (AUC) curve.
[0091]
[0092]
[0093] Table 10
[0094] According to the above embodiment, during the screening of test items IT for the raw measurement data DT0, the raw measurement data DT0, which includes measurements from different sources SR, undergoes source balancing procedures P1, subject normalization procedures P2, and source normalization procedures P3. This improves the prediction accuracy of the machine learning model MD. Furthermore, the raw measurement data DT0, after balancing and normalization, also exhibits excellent scalability, facilitating the training or correction of the machine learning model MD.
[0095] The above disclosure provides different features for implementing some embodiments or examples of the present invention. The specific examples of the components and configurations described above (e.g., the numerical values or names mentioned) are to simplify / illustrate some embodiments of the present invention. Of course, these components and configurations are merely examples and are not intended to be limiting. In addition, some embodiments of the present invention may repeat reference symbols and / or letters in various examples. This repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or configurations discussed.
[0096] In summary, although the present invention has been disclosed above with reference to the embodiments, these are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations can be made without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A data processing method for machine learning, comprising: A source balancing procedure is performed on raw measurement data for a plurality of sources to obtain a balanced distribution diagram, wherein the quantities corresponding to different sources are the same, and the raw measurement data includes a plurality of test values corresponding to a plurality of test items for a plurality of subjects; For each of the subjects, performing a personalization scaling procedure on the test values to obtain a personalization measurement data, in which the test values of each of the subjects are scaled to the same value range; For each of the sources, performing a source scaling procedure on the detection values to obtain source normalized measurement data, in which the detection values of each of the sources are scaled to the same value range; Combining the balanced distribution map, the subject-normalized measurement data, and the source-normalized measurement data to obtain balanced subject-normalized data and balanced source-normalized data; dividing the balanced subject-normalized data and the balanced source-normalized data into a plurality of splits, each of the splits corresponding to all of the sources; Sampling each of the segments and analyzing a prediction capability table, the prediction capability table including a prediction capability of each of the test items; as well as According to the prediction capability table, part of the detection items are output, and the output detection items are used for modeling, training or prediction inference by a machine learning model.
2. The data processing method for machine learning according to claim 1, wherein: In the source quantity balancing process, an upsampling method is used to increase the data volume of the original measurement data.
3. The data processing method for machine learning according to claim 1, wherein: In the subject normalization process, the maximum value and the minimum value of the detection values of the same subject are normalized to 1 and 0 respectively.
4. The data processing method for machine learning according to claim 1, wherein: In the source normalization process, the detection values of the same source are subjected to a z-score transformation.
5. An electronic device comprising: a source quantity balancing unit for performing a source quantity balancing procedure on raw measurement data for a plurality of sources to obtain a balanced distribution diagram, wherein the quantities corresponding to different sources are the same in the balanced distribution diagram, and the raw measurement data includes a plurality of test values corresponding to a plurality of test items for a plurality of test subjects; a subject normalization unit for performing a personalization scaling procedure on the test values for each subject to obtain subject normalized measurement data, wherein the test values of each subject are scaled to the same value range; a source normalization unit for performing a source normalization procedure on the detection values for each source to obtain source normalized measurement data, in which the detection values of each source are scaled to the same value range; a merging unit for merging the balanced distribution map, the subject-normalized measurement data, and the source-normalized measurement data to obtain balanced subject-normalized data and balanced source-normalized data; and An extraction unit comprising: a splitter for splitting the balanced subject-normalized data and the balanced source-normalized data into a plurality of splits, each of which corresponds to all of the sources; a calculator for sampling each of the segments and analyzing a prediction capability table, the prediction capability table including a prediction capability of each of the test items; and A selector is used to output part of the detection items according to the prediction capability table, and the output detection items are used for modeling, training or prediction inference by a machine learning model.
6. The electronic device according to claim 5, wherein: The source quantity balancing unit uses an upsampling method to increase the data volume of the original measurement data.
7. The electronic device according to claim 5, wherein: The subject normalization unit normalizes the maximum value and the minimum value of the detection values of the same subject to 1 and 0 respectively. 8 . The electronic device as claimed in claim 5 , wherein the source normalization unit performs a z-score transform on the detection values of the same source.
9. The electronic device according to claim 5, wherein: The amount of data obtained by the segmentation device is the same.
10. The electronic device according to claim 5, wherein: The union of the segmentations obtained by the segmenter covers all the subjects.