Artificial intelligence-oriented quantitative data standardization method and device

By using a data standardization network model with a fully connected linear projection layer and a layer normalization layer, medical test reports are sorted and transformed to a normal distribution. This solves the problem of quantitative data standardization, improves the degree of data standardization, and enhances the diagnostic accuracy of AI models.

CN122073154APending Publication Date: 2026-05-22SHANGHAI FUZHI MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI FUZHI MEDICAL TECHNOLOGY CO LTD
Filing Date
2024-11-22
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively extract and standardize quantitative data from medical test reports, resulting in insufficient accuracy and consistency for AI models when processing this data.

Method used

A data normalization network model with a fully connected linear projection layer and a normalization layer is used to train the sample dataset. The data is transformed into a standard normal distribution through data sorting and normal distribution transformation, and necessary preprocessing is performed.

Benefits of technology

It improves the standardization of data, enhances the comparability between data texts, reduces the adverse effects of measurement causes or numerical constraints on AI training, optimizes data analysis results, and improves the diagnostic accuracy and flexibility of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122073154A_ABST
    Figure CN122073154A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence-oriented quantitative data standardization method, which comprises the following steps: (a) obtaining a plurality of sample data sets to be standardized, each sample data set comprising a plurality of indexes and a numerical value of the index corresponding to each index; (b) performing normalization preprocessing on the plurality of sample data sets to be standardized to obtain a plurality of preprocessed sample data sets in normal distribution; (c) respectively inputting each preprocessed sample data set in the plurality of preprocessed sample data sets in normal distribution into the data standardization network model for training, the data standardization network model outputs first output data of the sample data set, the data standardization network model comprises one or more combined data processing layers, and the combined data processing layers sequentially comprise a full-connection linear projection layer and a layer normalization layer; and (d) inputting the first output data of the sample data set into a subsequent deep learning network for training. According to the method, the deep learning network is combined to sort and convert the quantitative data, so that the standardization of the data is improved, and the comparability between data texts is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a quantitative data standardization method and apparatus for artificial intelligence. Background Technology

[0002] In the field of medical laboratory testing, the application of artificial intelligence (AI) and deep learning technologies is becoming increasingly widespread, demonstrating enormous potential in improving diagnostic accuracy, optimizing the allocation of medical resources, and developing personalized treatment plans. Currently, the application of AI technology in clinical testing and intelligent diagnosis mainly focuses on image recognition, pattern recognition, and data analysis. For example, deep learning algorithms can assist doctors in identifying lesions and abnormalities by analyzing medical images, such as X-rays, CT scans, and MRI images. Furthermore, AI systems can process and analyze large amounts of clinical data, including laboratory test reports, to identify disease patterns and risk factors.

[0003] However, despite the promising prospects of AI technology in the field of medical testing, several challenges remain in practical applications. One key issue is how to effectively extract quantitative data from medical test reports and standardize it for use by AI models. Medical test reports typically contain a large number of medical terms and numerical indicators, the format and units of which may vary depending on the laboratory and region. Furthermore, some test indicators may have upper and lower limits, requiring appropriate conversion or normalization so that AI models can accurately understand and analyze them.

[0004] To overcome these challenges, developing a standardized method for quantitative data in medical laboratory reports for artificial intelligence is crucial. This method needs to automatically identify and parse numerical data in laboratory reports, converting it into a standardized format and units, and performing necessary data preprocessing, such as normalization and noise reduction, to improve data quality and usability. This approach provides AI models with accurate, consistent, and high-quality input data, thereby improving model performance and diagnostic accuracy. Furthermore, this method should possess flexibility and scalability to adapt to evolving medical laboratory standards and the development of AI technologies. Summary of the Invention

[0005] The purpose of this application is to provide a quantitative data standardization method for artificial intelligence. By combining deep learning networks to sort and transform quantitative data, the standardization of data is improved and the comparability between data texts is enhanced.

[0006] The first aspect of this application provides a quantitative data standardization method for artificial intelligence, comprising the following steps:

[0007] (a) Obtain multiple sample datasets to be standardized, each sample dataset including multiple indicators and the numerical values ​​of the indicators corresponding to each indicator;

[0008] (b) Normalize the multiple sample datasets to be standardized, sort the data and transform them into a normal distribution for the values ​​of the same indicator corresponding to the same indicator in the multiple sample datasets to be standardized, and obtain multiple preprocessed sample datasets with normal distribution. The values ​​of the same indicator in the multiple preprocessed sample datasets with normal distribution are in a standard normal distribution in different preprocessed sample datasets.

[0009] (c) Input each of the preprocessed sample datasets in the multiple preprocessed sample datasets of the normal distribution into the data normalization network model for training, and output the first output data of the sample dataset. The data normalization network model includes one or more combined data processing layers, and the combined data processing layer includes a fully connected linear projection layer and a layer normalization layer in sequence.

[0010] (d) Input the first output data of the sample dataset into the subsequent deep learning network for training.

[0011] In another preferred embodiment, in step (b), the values ​​of the indicators corresponding to each indicator in the multiple sample datasets to be standardized are preprocessed for normalization.

[0012] In another preferred example, the metrics in each sample dataset are either correlated or unrelated.

[0013] In another preferred embodiment, the multiple sample datasets to be standardized are blood routine test results reports, urine routine test results reports, blood biochemistry test results reports, or other quantitative clinical test results reports of multiple patients in a hospital or laboratory database.

[0014] In another preferred embodiment, the blood routine test results report includes, but is not limited to, multiple indicators such as white blood cells, red blood cells, hemoglobin, lymphocyte percentage, neutrophil count, lymphocyte count, and mean corpuscular hemoglobin concentration.

[0015] In another preferred embodiment, the urinalysis results report includes, but is not limited to, pH, specific gravity, urobilinogen, occult blood, white blood cells, protein, glucose, bilirubin, ketones, and red blood cells.

[0016] In another preferred embodiment, the blood biochemistry test results report includes, but is not limited to, multiple indicators such as α-L-fucosidase (AFU), alanine aminotransferase (ALT), aspartate aminotransferase (AST), total protein (TP), albumin (ALB), albumin / globulin ratio (A / G), alkaline phosphatase (ALP), L-γ-glutamyl transferase (L-γ-GGT), total bilirubin (TBIL), direct bilirubin (DBIL), indirect bilirubin (IBIL), cholinesterase (CHE), total bile acids (TBA), adenosine deaminase (ADA), prealbumin (PALB), homocysteine ​​(HCY), glucose (GLU), triglycerides (TG), total cholesterol (TCHO), high-density lipoprotein cholesterol (HDL-C), low-density lipoprotein cholesterol (LDL-C), apolipoprotein A1 (ApoA1), apolipoprotein B (ApoB), and lipoprotein a (Lp( a) Creatinine (CREA), uric acid (UA), blood urea nitrogen (BUN or UREA), cystatin C (Cys-C), β2-microglobulin (β2-MG), potassium (K), sodium (NA), chloride (Cl), total carbon dioxide (TCO2), anion gap (AG), phosphorus (P), magnesium (Mg), calcium (Ca), serum iron (Fe), total iron-binding capacity (TIBC), unsaturated iron-binding capacity (UIBC), transferrin (TRF), transferrin saturation (TS), lactate dehydrogenase (LDH), hydroxybutyrate dehydrogenase (HBDH), creatine kinase (CK), creatine kinase isoenzyme (CKMB), myoglobin (Mb or MYO), anti-streptolysin O (ASO), rheumatoid factor (RF), high-sensitivity C-reactive protein (HCRP), anti-cyclic citrullinated peptide antibody (CCP), amylase (AMY), lipase (LIPA or LPS), glycated hemoglobin (HbA1c).

[0017] In another preferred embodiment, in step (c), multiple indicators in each preprocessed sample dataset of the normally distributed multiple preprocessed sample datasets and the value of the indicator corresponding to each indicator are respectively input into a fully connected linear projection layer for linear transformation to extract the primary feature vector of each sample dataset. The primary feature vector of each sample dataset is the value of the first new indicator corresponding to each indicator.

[0018] In another preferred embodiment, in step (c), the primary feature vector of each sample dataset is input into the layer normalization layer, and the value of each corresponding new indicator in the primary feature vector of each sample dataset is internally normalized to obtain the standardized primary feature vector of each sample dataset. The standardized primary feature vector of each sample dataset includes multiple indicators and the value of a second new indicator corresponding to each indicator.

[0019] In another preferred embodiment, in step (b), the values ​​of the same indicator in the multiple sample datasets to be standardized are sorted from smallest to largest, and then transformed into a normal distribution using methods including but not limited to Z-score or Normal score.

[0020] In another preferred embodiment, the number of layers in the combined data processing layer is 1-10, preferably 2-6. More preferably, the number of layers in the combined data processing layer is 3.

[0021] A second aspect of this application provides an apparatus for quantitative data standardization for artificial intelligence, comprising:

[0022] Memory, used to store computer-executable instructions; and,

[0023] A processor, coupled to the memory, is configured to implement the steps of the method described above when executing the computer-executable instructions.

[0024] A third aspect of this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method described above.

[0025] A fourth aspect of this application provides a computer program product including computer-executable instructions, characterized in that the computer-executable instructions, when executed by a processor, implement the steps in the above-described method.

[0026] The fifth aspect of this application provides a quantitative data standardization device for artificial intelligence, comprising:

[0027] The data acquisition module is configured to acquire multiple sample datasets to be standardized, each sample dataset including multiple indicators and the numerical values ​​of the indicators corresponding to each indicator;

[0028] The data preprocessing module is configured to perform normalization preprocessing on the multiple sample datasets to be standardized. The module sorts and transforms the values ​​of the same indicator in each of the sample datasets into normal distributions, thereby obtaining multiple preprocessed sample datasets with normal distributions. The values ​​of the same indicator in the multiple preprocessed sample datasets with normal distributions in different preprocessed sample datasets follow a standard normal distribution.

[0029] The data standardization training module includes a data standardization network model, which is configured to train each of the multiple preprocessed sample datasets in the normally distributed dataset separately and output the first output data of the sample dataset. The data standardization network model includes one or more combined data processing layers, which sequentially include a fully connected linear projection layer and a layer normalization layer.

[0030] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. It should be understood that the accompanying drawings described below are merely some implementation examples of the present invention, and those skilled in the art can obtain other implementation examples based on these drawings without creative effort.

[0032] Figure 1 This is a schematic diagram of a quantitative data standardization method for artificial intelligence based on this application;

[0033] Figure 2 This is a schematic diagram of the data standardization network model structure according to the embodiments of this application;

[0034] Figure 3 This is an example diagram illustrating the application of the method of this application for determining gender from routine blood data;

[0035] Figure 4 This is an example diagram illustrating the application of the method described in this application for determining whether the eyes are open on EEG data. Detailed Implementation

[0036] Through extensive and in-depth research, the inventors have developed for the first time a quantitative data standardization method for artificial intelligence. This method improves the standardization of data in sample datasets and enhances the comparability between data texts by using a data standardization network model with a fully connected linear projection layer and a layer normalization layer to train and process the sample dataset.

[0037] In the following description, many technical details are presented to help the reader better understand this application. However, those skilled in the art will understand that the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.

[0038] This application possesses at least one of the following advantages.

[0039] (a) The quantitative data standardization method for artificial intelligence in this application improves the standardization of data in the sample dataset and enhances the comparability between data texts by using a data standardization network model with a fully connected linear projection layer and a layer normalization layer to train and process the sample dataset.

[0040] (b) The quantitative data standardization method for artificial intelligence in this application reduces the adverse effects of numerical values ​​in the dataset on artificial intelligence training due to measurement reasons or the constraints of the numerical values ​​themselves, optimizes the comparability of data, and improves the effect of data analysis or subsequent model training.

[0041] (c) The quantitative data standardization method for artificial intelligence in this application can efficiently process data: through sorting and normal distribution transformation, it can efficiently process and standardize data such as blood routine data;

[0042] (d) The quantitative data standardization method for artificial intelligence proposed in this application has the function of deep feature extraction. By combining multi-layer fully connected linear projection layers and LayerNorm operators, it can extract deep features of the data to be standardized.

[0043] (e) The quantitative data standardization method for artificial intelligence in this application is highly adaptable and can be applied to various routine blood data analysis tasks, including classification and regression.

[0044] (f) Compared with traditional data standardization methods, the method of this application has a significant advantage in reducing the loss function, thus achieving high accuracy;

[0045] (g) The method of this application can automatically identify and parse numerical data in test reports, convert them into a uniform format and units, and perform necessary data preprocessing, such as normalization and noise reduction, to improve the quality and usability of the data. Through this method, accurate, consistent and high-quality input data can be provided to AI models, thereby improving the performance of the models and the accuracy of diagnosis. In addition, the method also has a certain degree of flexibility and scalability to adapt to the ever-changing medical testing standards and the development of AI technology.

[0046] A Quantitative Data Standardization Method for Artificial Intelligence

[0047] This embodiment provides a quantitative data standardization method for artificial intelligence, which includes the following steps:

[0048] (a) Obtain multiple sample datasets to be standardized, each sample dataset including multiple indicators and the numerical values ​​of the indicators corresponding to each indicator;

[0049] (b) Normalize the values ​​of the indicators corresponding to each indicator in the multiple sample datasets to be standardized. Sort and transform the values ​​of the same indicator corresponding to each indicator in the sample datasets to be standardized into a normal distribution, obtaining multiple preprocessed sample datasets with a normal distribution. The values ​​of the same indicator in the multiple preprocessed sample datasets with a normal distribution in different preprocessed sample datasets exhibit a standard normal distribution. For example...

[0050] (c) Input each of the preprocessed sample datasets in the multiple preprocessed sample datasets of the normal distribution into the data normalization network model for training, and output the first output data of the sample dataset. The data normalization network model includes one or more combined data processing layers, and the combined data processing layer includes a fully connected linear projection layer and a layer normalization layer in sequence.

[0051] (d) Input the first output data of the sample dataset into the subsequent deep learning network for training.

[0052] In one embodiment, the data sorting and normal distribution transformation in step (b) are performed in the following manner.

[0053] For example, the normal distribution transformation can be performed using the following formula.

[0054] P = (Rank + 0.5) / N

[0055] Z = Norm.Inv(P, 0, 1)

[0056] Where P represents the cumulative distribution function (CDF) value of the data point in the population, which can also be regarded as a probability value; Rank refers to the ranking of the data point among all data, usually in ascending order; N represents the total number of samples, i.e., the number of all data points. The function Z = Norm.Inv(P,0,1) finds the quantile of the standard normal distribution (mean 0, standard deviation 1) corresponding to a given probability P.

[0057] A sample set of data is shown in Table 1 below:

[0058]

[0059]

[0060] In one embodiment, in step (c), multiple indicators and the numerical values ​​of each indicator in each of the multiple preprocessed sample datasets of the normal distribution are respectively input into a fully connected linear projection layer for linear transformation to extract the primary feature vector of each sample dataset. The primary feature vector of each sample dataset includes the numerical value of the first new indicator corresponding to each indicator.

[0061] In one embodiment, the primary feature vector of each sample dataset is input into the layer normalization layer, and the value of each corresponding new indicator in the primary feature vector of each sample dataset is internally normalized to obtain the standardized primary feature vector of each sample dataset. The standardized primary feature vector of each sample dataset includes multiple indicators and the value of a second new indicator corresponding to each indicator.

[0062] A quantitative data standardization device for artificial intelligence

[0063] The device comprises the following parts:

[0064] The data acquisition module is configured to acquire multiple sample datasets to be standardized, each sample dataset including multiple indicators and the numerical values ​​of the indicators corresponding to each indicator;

[0065] The data preprocessing module is configured to perform normalization preprocessing on the multiple sample datasets to be standardized. The module sorts and transforms the values ​​of the same indicator in each of the sample datasets into normal distributions, thereby obtaining multiple preprocessed sample datasets with normal distributions. The values ​​of the same indicator in the multiple preprocessed sample datasets with normal distributions in different preprocessed sample datasets follow a standard normal distribution.

[0066] The data standardization training module includes a data standardization network model, which is configured to train each of the multiple preprocessed sample datasets in the normally distributed dataset separately and output the first output data of the sample dataset. The data standardization network model includes one or more combined data processing layers, which sequentially include a fully connected linear projection layer and a layer normalization layer.

[0067] To make the objectives, technical solutions, and advantages of the present invention clearer, embodiments of the present invention will be described in further detail below with reference to the accompanying drawings. It should be understood that these are merely examples provided to the reader of possible implementations of the present invention and are not intended to limit the scope of the invention.

[0068] Example 1

[0069] This embodiment provides a quantitative data standardization method for artificial intelligence. Taking medical blood routine data as an example, the method includes the following steps:

[0070] 1. Data preprocessing steps: The data preprocessing steps are as follows:

[0071] 1) Raw data acquisition: Obtain the patient's blood routine test results report from the hospital or laboratory database. Use all the data in one report as a sample dataset, and n reports constitute a complete dataset consisting of n sample datasets.

[0072] 2) Data sorting: Sort the values ​​of the same blood routine indicator in n reports from smallest to largest. Sort each indicator in the blood routine report using this sorting method.

[0073] 3) Normal Distribution Transformation: Based on the formula, the data is converted to a standard normal distribution, completing the data normalization process. This data normalization involves normalizing the data (values ​​of the same indicator) of a specific indicator across n reports in a blood routine test report; that is, normalizing the data between different report samples. Data normalization methods include, but are not limited to, z-score / normal score methods. Therefore, through preprocessing, multiple preprocessed sample datasets with a normal distribution are obtained, such as z0, z1, ... z n .

[0074] 2. Network Structure Setup

[0075] See Figure 2 A data standardization network model for quantitative data standardization methods oriented towards artificial intelligence is presented.

[0076] 1) Fully connected layer 1 (FC1): The normal distribution data of the input layer is passed to the first fully connected linear projection layer, where a linear transformation is performed to extract the primary feature vector.

[0077] Multiple preprocessed sample datasets, such as z0, z1, ... z, are obtained by preprocessing to obtain a normally distributed dataset. n The inputs are fed into a fully connected layer 1 (FC1), and after linear transformation, the primary feature vectors z′0, Z′1, ..., z′ of each sample dataset are obtained. n That is, the primary feature vector of each sample dataset includes multiple indicators and the value of the first new indicator corresponding to each indicator.

[0078] For example, using the formula O = WI, a weighted average is calculated on the 20 variables that have undergone preprocessing (e.g., data from a preprocessed blood routine report; however, since hospital blood routine data is confidential, the random example data in Table 2 is used to illustrate the steps of this method), generating 20 new variables. Here, O can represent the output or a result, W may represent weights or coefficients, and I represents the input or a variable. The output variable will be used as the input for the next layer.

[0079] Table 2 Random Example Data

[0080] variable enter Output V1 1.169001785 0.895418033 V2 1.236723943 0.733070305 V3 -0.135000664 1.076078213 V4 -0.216861024 0.840640227 V5 -0.336491018 1.425795322 V6 -0.057646649 1.187781164 V7 0.91694467 1.102982261 V8 1.187023841 -0.939909389 V9 1.174667865 -0.969408485 V10 -0.561113931 -1.462324584 V11 -0.70078791 -1.234823758 V12 0.901568563 -0.483218584 V13 0.316773707 0.496786349 V14 -0.67652941 -0.831682071 V15 0.430924837 -1.425101744 V16 1.200439785 -0.934464066 V17 -0.150056071 0.457840829 V18 0.292300848 -0.178742275 V19 1.239462201 -1.353332574 V20 -0.910215625 -1.479761071

[0081] 2) Layer Normalization 1 (LayerNorm Operator 1): Standardizes the primary feature vectors output by FC1 to obtain the standardized primary feature vectors z″0, z″1, ..., z″ for each sample dataset. n The primary feature vector of each standardized sample dataset includes multiple indices and the value of a second new indice corresponding to each indice, thereby eliminating scale differences between different features.

[0082] For example, through layer normalization, 20 new variables are all transformed into a normal distribution, meaning that the mean of the 20 new variables is zero and the variance is one. This completes the internal normalization of, for example, a blood routine report, generating 20 latent variables that have already undergone one standardization process. The main significance is to eliminate scale differences between different samples, reduce the repetitive weighting of correlated indicators in a single blood routine report during data training, and prevent statistical or training result bias.

[0083] The weight parameters of a fully connected layer are not fixed or random; they are trained together with the subsequent data. The weight parameters will change when connected to different training algorithms or training objectives.

[0084] Through formula Standardize the data.

[0085] Where O: represents the standardized value or output (z-score), which is the value of the second new indicator corresponding to each indicator; I: represents the original data value, which is the value of the first new indicator corresponding to each indicator; Average(I): represents the mean of dataset I; SD(I): represents the standard deviation of dataset I.

[0086] For example, in the following case, the input data comes from the output of the previous layer:

[0087] variable enter average -0.1538187949 Output V1 0.895418033 SD 1.05671118303306 0.992926783456573 V2 0.733070305 0.839291872543641 V3 1.076078213 1.16389136055133 V4 0.840640227 0.941088772241437 V5 1.425795322 1.49483998267152 V6 1.187781164 1.26959947939606 V7 1.102982261 1.18935153217552 V8 -0.939909389 -0.743902973563694 V9 -0.969408485 -0.771818923123346 V10 -1.462324584 -1.23828138545861 V11 -1.234823758 -1.02298998861733 V12 -0.483218584 -0.311721679711953 V13 0.496786349 0.615688715434997 V14 -0.831682071 -0.641483958748312 V15 -1.425101744 -1.20305620690036 V16 -0.934464066 -0.738749888119855 V17 0.457840829 0.578833309022055 V18 -0.178742275 -0.0235858912744729 V19 -1.353332574 -1.1351387178472 V20 -1.479761071 -1.25478209760211

[0088] 3) Fully connected layer 2 (FC2): The primary feature vector after normalization 1 is passed to the second fully connected linear projection layer to further extract intermediate feature vectors.

[0089] 4) Layer Normalization 2 (LayerNorm operator 2): Standardizes the intermediate feature vectors output by FC2.

[0090] 5) Fully connected layer 3 (FC3): The normalized intermediate feature vector is passed to the third fully connected linear projection layer to extract the final high-level feature vector.

[0091] 6) Layer Normalization 3 (LayerNorm operator 3): Standardizes the high-level feature vectors output by FC3.

[0092] Subsequent analysis network: The standardized high-level feature vectors are input into the subsequent analysis network for further prediction and classification.

[0093] As can be seen, in this embodiment, the combined data processing layer of the fully connected layer and the layer normalization layer consists of three layers.

[0094] 3. Model Training

[0095] It can connect to standard neural network training modules, such as forward propagation and backpropagation. The data standardization processing apparatus protected in this application (which executes the above-mentioned quantitative data standardization method steps for artificial intelligence) is interconnected with the subsequently connected data training module; data standardization and data training are linked. That is, the weights of the FC layer are adaptively adjusted based on the training results, continuously iterating during the data training process. The weights of the FC layer are only finally fixed when the training results meet the requirements.

[0096] 4. Comparison of data standardization results

[0097] See Figure 3 This document demonstrates the results of gender prediction using a dataset based on nine blood routine indicators (this dataset is confidential and not shown here). The results show data standardization using the commonly used normal score standard, followed by regression using a classifier. It also demonstrates data standardization using the quantitative data standardization method for artificial intelligence proposed in this application, followed by regression using a classifier. Training and testing loss graphs for both conventional methods and this application are presented. The results show that the cross-entropy loss for gender prediction decreases after using the method proposed in this application, indicating a significant improvement in prediction accuracy, regardless of whether it is on the training or testing set.

[0098] Figure 4The paper presents training and testing loss graphs based on the EEG prediction dataset for eye opening (https: / / archive.ics.uci.edu / dataset / 264 / eeg+eye+state), using both conventional data standardization methods and the quantitative data standardization method for artificial intelligence proposed in this application. The results show that the cross-entropy loss for gender determination decreases on both the training and testing sets after applying the method described in this application, indicating a significant improvement in prediction accuracy.

[0099] It should be noted that in this patent application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. In this patent application, if it refers to performing an action according to an element, it means performing the action at least according to that element, including two cases: performing the action only according to that element, and performing the action according to that element and other elements. Expressions such as "multiple," "repeatedly," and "various" include two, two times, two kinds, and more than two, more than two times, and more than two kinds.

[0100] In this invention, all directional indicators (such as up, down, left, right, front, back, etc.) are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.

[0101] The specification of this application contains numerous technical features distributed across various technical solutions. Listing all possible combinations of these technical features (i.e., technical solutions) would make the specification excessively lengthy. To avoid this problem, the various technical features disclosed in the above-described invention, the various technical features disclosed in the following embodiments and examples, and the various technical features disclosed in the accompanying drawings can be freely combined to form various new technical solutions (all of which are considered to have been described in this specification), unless such a combination of technical features is technically infeasible. For example, one example discloses feature A+B+C, and another example discloses feature A+B+D+E. Features C and D are equivalent technical means that serve the same function, and technically only one needs to be used; they cannot be used simultaneously. Feature E can technically be combined with feature C. Therefore, the solution A+B+C+D should not be considered as described because it is technically infeasible, while the solution A+B+C+E should be considered as described.

[0102] All documents mentioned in this application are considered to be incorporated in their entirety into the disclosure of this application so that they can serve as a basis for modifications if necessary. Furthermore, it should be understood that after reading the foregoing disclosure of this application, those skilled in the art can make various alterations or modifications to this application, and these equivalent forms also fall within the scope of protection claimed in this application.

Claims

1. A quantitative data standardization method for artificial intelligence, characterized in that, Includes the following steps: (a) Obtain multiple sample datasets to be standardized, each sample dataset including multiple indicators and the numerical values ​​of the indicators corresponding to each indicator; (b) Normalize the multiple sample datasets to be standardized, sort the data and transform them into a normal distribution for the values ​​of the same indicator corresponding to the same indicator in the multiple sample datasets to be standardized, and obtain multiple preprocessed sample datasets with normal distribution. The values ​​of the same indicator in the multiple preprocessed sample datasets with normal distribution are in a standard normal distribution in different preprocessed sample datasets. (c) Input each of the preprocessed sample datasets in the multiple preprocessed sample datasets of the normal distribution into the data normalization network model for training, and output the first output data of the sample dataset. The data normalization network model includes one or more combined data processing layers, and the combined data processing layer includes a fully connected linear projection layer and a layer normalization layer in sequence. (d) Input the first output data of the sample dataset into the subsequent deep learning network for training.

2. The method as described in claim 1, characterized in that, The metrics in each sample dataset may be correlated or unrelated.

3. The method as described in claim 1, characterized in that, The multiple sample datasets to be standardized are blood routine test results reports, urine routine test results reports, blood biochemistry test results reports, or other quantitative clinical test results reports from multiple patients in a hospital or laboratory database.

4. The method as described in claim 1, characterized in that, In step (c), multiple indicators and the values ​​of each indicator in each of the multiple preprocessed sample datasets of the normal distribution are input into a fully connected linear projection layer for linear transformation to extract the primary feature vector of each sample dataset. The primary feature vector of each sample dataset is the value of the first new indicator corresponding to each indicator.

5. The method as described in claim 4, characterized in that, In step (c), the primary feature vector of each sample dataset is input into the layer normalization layer, and the value of each corresponding new indicator in the primary feature vector of each sample dataset is internally normalized to obtain the standardized primary feature vector of each sample dataset. The standardized primary feature vector of each sample dataset includes multiple indicators and the value of a second new indicator corresponding to each indicator.

6. The method as described in claim 1, characterized in that, In step (b), the values ​​of the same indicator in the multiple sample datasets to be standardized are sorted from smallest to largest, and then transformed into a normal distribution using methods including but not limited to Z-score or Normal score.

7. The method as described in claim 1, characterized in that, The number of layers in the combined data processing layer is 1-10, preferably 2-6.

8. A device for quantitative data standardization for artificial intelligence, characterized in that, include: Memory is used to store executable instructions for a computer; as well as, A processor, coupled to the memory, is configured to implement the steps of the method as described in any one of claims 1 to 7 when executing the computer-executable instructions.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 7.

10. A computer program product comprising computer-executable instructions, characterized in that, When executed by a processor, the computer-executable instructions implement the steps of the method according to any one of claims 1 to 7.