Method for improving accuracy of antenatal Down's syndrome data and related equipment

By obtaining biomarker sample data and combining detection equipment, storage environment and duration information, and using the data influence factor model to correct it, the data accuracy problems caused by detection equipment and environmental factors are solved, and high-precision and personalized auxiliary judgments for prenatal Down syndrome detection are achieved.

CN120299592APending Publication Date: 2025-07-11XIAMEN CHANG GUNG MEMORIAL HOSPITAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510275184.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the existing prenatal Down syndrome screening methods, due to the differences in detection equipment, storage environment and storage duration, the accuracy of biomarker sample data is reduced, affecting the accuracy and reliability of screening results.

Method used

By obtaining biomarker sample data, combining detection equipment information, storage environment information and storage duration information, the data impact factor model is used to correct it, and machine learning training is used to determine the impact factor, quantify the degree of impact of different factors on the sample data, and achieve multi-dimensional data correction.

Benefits of technology

It improves the accuracy of prenatal Down syndrome testing data, provides more reliable auxiliary judgment data, and improves the personalization and accuracy of the testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005303996260000011
    Figure HDA0005303996260000011
  • Figure HDA0005303996260000021
    Figure HDA0005303996260000021
  • Figure HDA0005303996260000031
    Figure HDA0005303996260000031
Patent Text Reader

Abstract

The invention provides a method for improving accuracy of antenatal Down's syndrome data and related equipment, and relates to the technical field of medical care informatics. The method comprises the steps that sample data of a plurality of target biomarkers is obtained to serve as a basic data source, three key influence factors including detection equipment information, storage environment information and storage duration information are combined, and a data influence factor model is obtained through training in a machine learning mode; the influence degree of different factors on sample data can be accurately quantified. Original data is corrected through the synergistic effect of the first, second and third influence factors, and finally more accurate biomarker sample data is obtained. According to the multi-dimensional data correction method, the accuracy of auxiliary judgment data of antenatal Down's syndrome detection is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of healthcare informatics, and particularly to a method for improving the accuracy of prenatal Down syndrome data and related devices. Background Art

[0002] Prenatal Down syndrome is a genetic disease that seriously affects the health of the fetus, and its early diagnosis is of great significance for improving the quality of prenatal and postnatal care. At present, screening for prenatal Down syndrome by detecting biomarkers in maternal serum has become an important means in clinical practice.

[0003] The existing methods for prenatal Down syndrome screening mainly involve collecting maternal serum samples, quantitatively detecting biomarkers such as pregnancy-associated plasma protein A (PAPP-A), alpha-fetoprotein (AFP), unconjugated estriol (uE3), and beta-human chorionic gonadotropin (β-HCG) in the samples using detection equipment, and then inputting the detection data into an evaluation model for risk assessment.

[0004] However, in the actual detection process, due to differences in the detection equipment used by different medical institutions, and the inconsistent storage environment and storage duration of the samples after collection, the detection results of the same patient may fluctuate greatly, thus affecting the accuracy and reliability of the screening results. Summary of the Invention

[0005] This application provides a method for improving the accuracy of prenatal Down syndrome data and related devices, which is used to solve the problem of reduced accuracy of auxiliary data for prenatal Down syndrome caused by multiple factors such as detection equipment, storage environment, and storage duration in the existing prenatal screening process.

[0006] First aspect, this application provides a method for improving the accuracy of prenatal Down syndrome data, which is applied to a data correction device. The method includes: obtaining multiple sets of target biomarker sample data, where the target biomarker sample data at least includes serum pregnancy-associated plasma protein A and β-human chorionic gonadotropin in the first trimester, or alpha-fetoprotein, free estriol, and β-human chorionic gonadotropin in the second trimester; obtaining detection device information, storage environment information, and storage duration information; combining the detection device information, storage environment information, and storage duration information, and determining a first influence factor corresponding to the detection device information, a second influence factor corresponding to the storage environment information, and a third influence factor corresponding to the storage duration information through a data influence factor model. The data influence factor model is obtained through machine learning training in advance according to a detection device data set, a storage environment data set, and a storage duration data set with corresponding influence factor annotation information. The data influence factor model is used to quantify the influence degree of different detection devices, different storage environment conditions, and different storage durations on the accuracy of sample data; correcting the target biomarker sample data according to the first influence factor, the second influence factor, and the third influence factor to determine the final target biomarker sample data; determining the final target biomarker sample data as prenatal Down syndrome auxiliary judgment data.

[0007] By adopting the above technical solution, by obtaining multiple sets of target biomarker sample data as the basic data source, and combining the three key influencing factors of detection device information, storage environment information, and storage duration information, and the data influence factor model is obtained through machine learning training, which can accurately quantify the influence degree of different factors on the sample data. Through the synergistic effect of the first, second, and third influence factors, the original data is corrected, and finally more accurate biomarker sample data is obtained. This multi-dimensional data correction method significantly improves the accuracy of prenatal Down syndrome detection data and provides a more reliable data basis for subsequent auxiliary judgment.

[0008] Second aspect, the present application provides a data correction device, which includes: a sample data acquisition module, configured to acquire multiple target biomarker sample data, where the target biomarker sample data includes at least serum pregnancy-associated plasma protein A and β-human chorionic gonadotropin in the first trimester, or alpha-fetoprotein, free estriol, and β-human chorionic gonadotropin in the second trimester; multiple information acquisition modules, configured to acquire detection device information, storage environment information, and storage duration information; an influencing factor determination module, configured to combine the detection device information, storage environment information, and storage duration information, and determine a first influencing factor corresponding to the detection device information, a second influencing factor corresponding to the storage environment information, and a third influencing factor corresponding to the storage duration information through a data influencing factor model, where the data influencing factor model is obtained by machine learning training in advance according to a detection device data set, a storage environment data set, and a storage duration data set with corresponding influencing factor annotation information, and the data influencing factor model is used to quantify the influence degree of different detection devices, different storage environment conditions, and different storage durations on the accuracy of sample data; a final sample data determination module, configured to correct the target biomarker sample data according to the first influencing factor, the second influencing factor, and the third influencing factor to determine the final target biomarker sample data; an auxiliary judgment data determination module, configured to determine the final target biomarker sample data as prenatal Down syndrome auxiliary judgment data.

[0009] Third aspect, the present application provides a data correction device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, where the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the data correction device to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0010] Fourth aspect, the present application provides a computer-readable storage medium, including instructions, which when running on a data correction device, enable the data correction device to execute the method described in the first aspect and any possible implementation manner in the first aspect. Description of the Drawings

[0011] Figure 1 is a flowchart of a method for improving the accuracy of prenatal Down syndrome data in an embodiment of the present application; Figure 2 is another flowchart of a method for improving the accuracy of prenatal Down syndrome data in an embodiment of the present application; Figure 3 is a schematic module structure diagram of a data correction device in an embodiment of the present application; Figure 4It is a schematic diagram of the hardware structure of the data correction device in an embodiment of the present application. Detailed implementation manners

[0012] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more of the listed items.

[0013] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0014] For ease of understanding, the method provided in this embodiment will be described in terms of its process below. Please refer to Figure 1 , which is a schematic diagram of a process for the method of improving the accuracy of prenatal Down syndrome data in an embodiment of the present application.

[0015] S101. Obtain data of multiple target biomarker samples, where the data of the target biomarker samples includes at least serum pregnancy-associated plasma protein A and β-human chorionic gonadotropin in the first trimester, or alpha-fetoprotein, free estriol, and β-human chorionic gonadotropin in the second trimester; The data correction device is connected to the hospital's laboratory information system (LIS) through a built-in data interface module to automatically obtain detection data of various biomarkers such as serum pregnancy-associated plasma protein A (PAPP-A), alpha-fetoprotein (AFP), free estriol (uE3), and β-human chorionic gonadotropin (β-hCG). Among them, β-human chorionic gonadotropin (β-HCG) can include data of pregnant women in the first trimester and the second trimester. For historical detection data, the device uses the read_excel function to read the stored Excel format data file. During the reading process, the device will automatically detect the file encoding format and perform conversion as needed to ensure that special characters can be correctly displayed. When importing data, the device will record metadata such as the source and collection time of the data for subsequent traceability and quality control.

[0016] The data calibration device performs multi-level format standardization processing: First, it identifies and removes invisible characters and extra spaces in column names through regular expressions. Then, the device uniformly converts all column names into a predefined standard format. For example, different expressions such as "PAPP-A", "serum pregnancy-associated plasma protein A", and "PAPP-A value" are unified into the standard name "PAPP-A_VALUE". The device maintains a standard name mapping table to ensure that data from different sources can be processed uniformly. The device checks data integrity through preset verification rules: First, it verifies whether required fields exist, including sample ID, sampling time, detection values of three biomarkers, etc. Then, it checks whether the data type of each field meets the expectations, such as ensuring that the detection value is numeric and the timestamp format is correct. If a missing or incorrect format is found, the device will generate a detailed error report and decide whether to terminate the processing flow according to the configuration. For repairable problems, the device will call the corresponding repair function for processing.

[0017] S102. Obtain detection device information, save environment information, and save duration information; When the data calibration device obtains detection device information, it mainly uses two methods: directly connecting to the detection device and obtaining information from the device management system. If the detection device has a digital interface (such as USB, Ethernet interface) and supports a specific communication protocol (such as Modbus, TCP / IP), the device will use the corresponding communication module to connect to it. After connection, instructions are sent according to the device's communication protocol, such as querying information such as device model, manufacturer, calibration time, etc. Taking a common biochemical analyzer as an example, the device sends query instructions according to the specified instruction format of the analyzer. After the device receives the instructions, it returns its own information in a specific data frame format, and the device then parses the data frame to obtain the required information.

[0018] For detection devices that cannot be directly connected, the data calibration device will obtain information from the device management system of the hospital or testing institution. By docking with the device management system and based on a standardized data interface (such as HL7 protocol), after obtaining authorization, it reads the detailed information of the device from the system database, including device number, using department, maintenance records, etc. To ensure the accuracy and timeliness of the information, the device will synchronize data from the device management system regularly and perform data verification when obtaining information, verifying the integrity and accuracy of the data through a preset verification algorithm (such as CRC verification).

[0019] When obtaining storage environment information, the data correction device uses a variety of sensors and Internet of Things technologies. Temperature and humidity sensors, light sensors, gas sensors, etc. are deployed in sample storage places (such as refrigerators, sample libraries). These sensors will monitor the temperature, humidity, light intensity, harmful gas concentration and other parameters of the storage environment in real time, and transmit the data to the data correction device through wireless communication modules (such as Wi-Fi, Bluetooth, ZigBee). To ensure the reliability of sensor data, the device will calibrate the sensor regularly. For example, use a high-precision temperature and humidity calibrator to calibrate the temperature and humidity sensor, and adjust the sensor's measurement parameters by comparing the standard value and the sensor measurement value.

[0020] The acquisition of storage duration information depends on the sample collection time record and the device's time management system. When the sample is collected, the data correction device will record the sample collection timestamp, which is accurate to the second and stored in the device's database. The time management system inside the device will regularly synchronize time with the network time server (such as the NTP server) to ensure the accuracy of its own time. When it is necessary to obtain the storage duration information, the device calculates the sample's storage duration based on the current time and the sample collection timestamp. At the same time, in order to prevent errors in the calculation of the storage duration due to time system failures, the device will regularly back up the time data and perform multiple checks when calculating the storage duration.

[0021] After obtaining the detection equipment information, storage environment information and storage time information, the data correction device will integrate and pre-process this information. For the detection equipment information, the information from different sources will be merged and duplicate information will be removed; for the storage environment information, the abnormal data collected by the sensor will be cleaned, such as removing the temperature and humidity data that exceeds the reasonable range; for the storage time information, ensure that its correspondence with the sample is correct. Through these operations, an accurate and complete data basis is provided for the subsequent determination of the influencing factors.

[0022] S103, combining the detection device information, the storage environment information and the storage time information, determining a first influencing factor corresponding to the detection device information, a second influencing factor corresponding to the storage environment information and a third influencing factor corresponding to the storage time information through a data influencing factor model, wherein the data influencing factor model is obtained in advance through machine learning training based on a detection device data set, a storage environment data set and a storage time data set with corresponding influencing factor annotation information, and the data influencing factor model is used to quantify the degree of influence of different detection devices, different storage environment conditions and different storage time on the accuracy of sample data; A large amount of relevant data covering detection equipment, storage environment, and storage duration can be obtained in advance. The detection equipment data includes the brand, model, service life, calibration status, performance indicators, etc. of the equipment. There are differences in detection accuracy, stability, etc. among equipment of different brands and models, and these differences will affect the accuracy of biomarker sample data. The storage environment data involves the temperature, humidity, light, gas composition, etc. of the sample storage site. For example, some biomarkers may degrade in high-temperature and high-humidity environments, thus affecting the test results. The storage duration data records the time elapsed from sample collection to testing. As the storage duration increases, the quality of the sample may gradually decline.

[0023] After collecting the above data, it is necessary to label the corresponding influencing factors for each data sample. This is usually done through experimental research and actual case analysis. Researchers will, under the condition of controlling other variables, change factors such as detection equipment, storage environment, and storage duration respectively, observe the degree of influence of these factors on the accuracy of biomarker sample data, and quantify them as influencing factors respectively, then the first influencing factor, the second influencing factor, and the third influencing factor can be obtained. For example, through multiple experiments, it is found that due to the limitations of the detection method of a certain model of detection equipment, the test results of serum pregnancy-associated plasma protein A (PAPP-A) are on average 5% higher, then the corresponding influencing factor of this equipment can be labeled as 1.05.

[0024] The data influencing factor model in this step can adopt machine learning algorithms including decision trees, random forests, neural networks, etc. based on the obtained influencing factors. Among them, the decision tree algorithm constructs a decision tree model by hierarchically partitioning the data, intuitively showing the relationship between different factors and influencing factors. The random forest is an ensemble learning model composed of multiple decision trees, which can improve the accuracy and stability of the model. The data with labeled influencing factors is divided into a training set and a test set. The training set is used for model training. The model adjusts its own parameters by continuously learning the data features and corresponding influencing factors in the training set to establish a mapping relationship between the input (detection equipment information, storage environment information, storage duration information) and the output (influencing factor). The test set is used to evaluate the performance of the model and test the prediction accuracy of the model on unseen data.

[0025] Before the training process, since there may be problems such as class imbalance in the actually collected training data, for example, the amount of data samples of certain specific detection devices, storage environments, or storage durations is too small, which will cause the model to learn insufficiently about these minority-class data during training, thus affecting the generalization ability of the model to different situations. At this time, the ROSE (Random Over-Sampling Examples) method is used to enhance the training data. The ROSE method replicates the minority-class samples through random oversampling to increase their proportion in the training data, making the number of samples in each category more balanced. Taking the detection device information as an example, if the amount of data samples of a certain new detection device is small, the ROSE method will randomly select these samples and replicate them, and add the replicated samples to the training data set, so that the model can better learn the influence characteristics of this new detection device on the accuracy of the sample data during the training process. Similarly, for the minority-class samples in the storage environment information and storage duration information, the same method is used for processing, so as to enrich the diversity of the training data and improve the robustness and accuracy of the model. After enhancing the training data with the ROSE method, it is also necessary to perform standardization processing on the training data and validation data of the detection device information, storage environment information, and storage duration information to eliminate the dimensional and scale differences between different feature data, making the model training more stable and efficient. The preProcess method is used to achieve this goal. The preProcess method can perform various standardization operations on the data, such as normalization, centering, etc. For categorical data such as device models and manufacturers in the detection device information, the preProcess method will encode them and convert them into numerical data for convenient subsequent model calculation. For continuous data such as temperature and humidity in the storage environment information and numerical data such as storage duration information, the preProcess method will normalize them to a specific interval, such as the [0, 1] interval, or perform standardization processing to make them have a distribution characteristic with a mean of 0 and a standard deviation of 1. In this way, when training the data impact factor model, data with different features can be compared and learned on the same scale, avoiding the model training being biased towards certain features due to data scale differences, thereby improving the performance and stability of the model.

[0026] During the model training process, it is necessary to continuously adjust the hyperparameters of the model, such as the depth of the decision tree, the number of trees in the random forest, the learning rate of the neural network, etc., to optimize the performance of the model. An independent validation dataset is used to validate the trained model to ensure that the model can accurately predict the impact factors in different situations. By calculating the errors between the predicted results and the actual labeled impact factors, such as the mean squared error, mean absolute error, etc., the accuracy and reliability of the model are evaluated. If the error is within the acceptable range, the model is considered to pass the validation; otherwise, the model needs to be further adjusted and optimized. With the continuous development of detection technologies, the emergence of new detection devices, and the in-depth study of the impact of storage environments and storage durations, it is necessary to continuously collect new data to update and optimize the model. For example, when a new type of detection device appears, relevant data of the device need to be collected and incorporated into the training dataset to retrain the model to ensure that the model can accurately reflect the latest situation.

[0027] S104. Correct and determine the final target biomarker sample data based on the first impact factor, the second impact factor, and the third impact factor; After obtaining the first impact factor, the second impact factor, and the third impact factor, the data correction device enters the correction process of the target biomarker sample data. This process uses a set of complex and delicate technical means to ensure that the finally obtained biomarker sample data is more accurate and reliable.

[0028] When determining the first interaction term F of the first impact factor and the second impact factor 12 = F1×F2, the second interaction term F of the second impact factor and the third impact factor 23 = F2×F3, the third interaction term F of the first impact factor and the third impact factor 13 = F1×F3, and the total interaction term F of the first impact factor, the second impact factor, and the third impact factor 123 = F1×F2×F3, the data correction device relies on its powerful computing ability to perform multiplication operations based on the values of each impact factor. For example, if the first impact factor F1 represents the degree of influence of a specific detection device on the detection result of serum pregnancy-associated plasma protein A (PAPP-A) as 1.03 (i.e., the device will make the detection result of serum pregnancy-associated plasma protein A (PAPP-A) 3% higher), and the second impact factor F2 represents the degree of influence of a specific storage environment (such as a storage environment with a higher temperature) on the detection result of serum pregnancy-associated plasma protein A (PAPP-A) as 0.98 (i.e., this environment will make the detection result of serum pregnancy-associated plasma protein A (PAPP-A) 2% lower), then their first interaction term F 12= 1.03 × 0.98 = 1.0094, which means that the comprehensive influence degree of the detection device and the storage environment on the detection result of serum pregnancy-associated plasma protein A (PAPP-A) is 0.94% on the high side. By calculating these interaction terms, the device can consider the synergistic effect between various factors more comprehensively, rather than just the influence of a single factor.

[0029] Next, according to the comprehensive correction formula: to correct the target biomarker sample data. Among them, D is the target biomarker sample data, is the final target biomarker sample data, a1, a2, a3, a 13 、a 23 、a 12 、a 123 are coefficients obtained by fitting a large amount of experimental data. The determination of these coefficients is crucial. They are obtained through extensive experimental research and data analysis on biomarker sample data under different detection devices, storage environments, and storage durations. For example, researchers detected biomarker samples of a large number of pregnant women under various different conditions, collected rich data, and used professional mathematical statistical methods, such as the least squares method, to perform fitting analysis on these data, so as to determine the coefficients that can accurately reflect the action degree of each influencing factor.

[0030] In the actual correction process, the data correction device uses its internal calculation module to substitute the original data D of each sample and the corresponding influencing factors and coefficients into the comprehensive correction formula for calculation. Suppose the original sample data D of serum pregnancy-associated plasma protein A (PAPP-A) is 50 (unit: μg / L), a1 = 0.2, a2 = 0.1, a3 = 0.15, a 12 = 0.05, a 13 = 0.08, a 23 = 0.06, a 123 = 0.03, F1 = 1.03, F2 = 0.98, F3 = 0.95, F 12 = 1.0094, F 13 = 0.9785, F 23 = 0.931, F 123 = 0.9588, then substituting into the formula gives: D' = 50 × (1 + 0.2 × 1.03 + 0.1 × 0.98 + 0.15 × 0.95 + 0.05 × 1.0094 + 0.08 × 0.9785 + 0.06 × 0.931 + 0.03) = 50 × (1 + 0.206 + 0.098 + 0.1425 + 0.05047 + 0.07828 + 0.05586 + 0.028764) = 50 × 1.659874 = 82.9937 (μg / L) During the calculation process, the device will store and process the calculation results of each step with high precision to avoid deviations in the final result due to precision loss. At the same time, to ensure the accuracy of the calculation, the device will also perform multiple calculation verifications. For example, different calculation methods are adopted or backup data is used for calculation comparison. If the result difference is within the allowable error range, the calculation result is considered reliable; if the difference is too large, the data and calculation process will be rechecked to find and solve possible problems.

[0031] In addition, during the correction process, the device will also consider the boundary conditions of the data and the processing of outliers. For some raw data or influencing factors that exceed the reasonable range, the device will activate the outlier processing mechanism. For example, if the preservation duration information of a certain sample is abnormal (such as a negative preservation duration), the device will process it according to the preset rules, which may be to mark that the sample data needs to be further verified manually, or to reasonably correct the outlier based on historical data and experience. For boundary conditions, such as when some key data of the detection device information is missing, the device will refer to the relevant information of similar devices and combine expert experience for estimation and correction to ensure the accuracy of data correction to the greatest extent.

[0032] In some embodiments, the data correction device can interact with the hospital's information management system (such as the electronic medical record system), use a standardized data interface (such as the HL7 interface), and extract the individual information of pregnant women, including age, gestational week, weight, height, race, family medical history, etc. from the system database under the condition of obtaining authorization, and then adjust the final target biomarker sample data. Specifically, it may include but is not limited to the data correction device adjusting the final target biomarker sample data according to the age-related correction coefficients, which are obtained through statistical analysis of a large amount of clinical research data, and different age groups correspond to different correction coefficients. If a pregnant woman has a family history of Down syndrome, the risk of her fetus having Down syndrome will increase, and the biomarker level may also be affected. The data correction device will appropriately adjust the biomarker sample data according to the family medical history information.

[0033] S105. Determine the final target biomarker sample data as the prenatal Down syndrome auxiliary judgment data.

[0034] After the data correction device completes the correction of the target biomarker sample data, it uses the final target biomarker sample data for the auxiliary judgment of prenatal Down syndrome. This process involves multiple key steps, aiming to provide accurate and valuable diagnostic references for doctors. The specific implementation process will be described in detail later and will not be elaborated here.

[0035] In the embodiments of the present application, the method and related devices for improving the accuracy of prenatal Down syndrome data, by obtaining multiple target biomarker sample data, combining detection device information, preservation environment information, and preservation duration information, using a data impact factor model to determine impact factors and correct the sample data, realize high-precision correction of prenatal screening biomarker sample data in the early or middle trimester of pregnancy. It not only effectively solves the problem of data distortion caused by the difficulty of a single data correction method in dealing with the comprehensive influence of multiple factors in the prior art, improves the accuracy of the auxiliary judgment data for prenatal Down syndrome detection, but also realizes the precise quantification and comprehensive correction of the complex correlation between sample data influencing factors by considering the interaction between impact factors, enhancing the personalization, precision, and reliability of prenatal Down syndrome risk assessment.

[0036] In some embodiments, after the determination of the data impact factor model, the model has the ability to quantify the impact of various factors on the accuracy of sample data. Time series data can be input into the model, and this time series data can cover relevant information at different time nodes from sample collection to the detection process. The model will comprehensively analyze the impact of time on the quality of sample data based on the patterns and rules learned during previous training. For example, the model will consider the possible changes in biomarkers such as serum pregnancy-associated plasma protein A (PAPP-A), alpha-fetoprotein (AFP), unconjugated estriol (uE3), and beta-human chorionic gonadotropin (β-HCG) in the sample over time. Some biomarkers may degrade over time, resulting in changes in concentration. Through the learning of a large amount of historical data and experimental data, the model can determine at what storage time the sample data of the target biomarker is least affected by time factors, and thus output the optimal time data for storing the sample data of the target biomarker. Based on the above-obtained optimal time data, the data correction device will further determine the optimal time threshold for the sample data of the target biomarker. This threshold is an important criterion for judging whether the sample data may affect the detection accuracy due to excessive storage time. Usually, the device will set this threshold in combination with clinical practice experience, industry standards, and research on the stability of sample data. For example, if the optimal time data shows that within 48 hours after sample collection, the sample data of the biomarker is less affected by time and is relatively stable, then after comprehensively considering factors such as the actual detection process and sample processing time, the optimal time threshold may be set to 72 hours. In this way, while ensuring a certain time buffer, the reliability of the sample data can be maximally ensured. During the sample storage and subsequent detection process, the data correction device will obtain the storage time information of the sample in real time. This information can be obtained through the time stamp recorded at the time of sample collection and the internal time management system of the device. The device will continuously compare the storage time of the sample with the optimal time threshold. If it is found that the storage time of the sample exceeds the optimal time threshold, the device will immediately activate the reminder mechanism and send a reminder message to the user side. The reminder method can be diversified. For example, a pop-up reminder can be sent to the staff responsible for sample detection through the hospital information system, informing them that the storage time of some samples is too long and may affect the accuracy of the detection results. Through this timely reminder mechanism, the staff can take corresponding measures in a timely manner, such as re-collecting the sample or conducting additional quality inspections on the sample, to avoid deviations in the detection results caused by using sample data affected by time, thereby ensuring the accuracy of prenatal Down syndrome detection.

[0037] After combining the above content, the following is a further and more specific process description of the method provided in this embodiment. Please refer to Figure 2 , which is another process schematic diagram of the method for improving the accuracy of prenatal Down syndrome data in the embodiments of the present application.

[0038] S201. After obtaining individual information of the pregnant woman, input the individual information into a personalized reference interval model to obtain a personalized reference interval, where the personalized reference interval model is obtained by analyzing biomarker data of a plurality of pregnant women with different individual characteristics in advance and then training using a machine learning algorithm; When the data correction device obtains individual information of pregnant women, it mainly interacts with the hospital's information management system or uses standardized data interfaces to extract various information of pregnant women from the system database with authorization, including age, gestational age, weight, height, race, family medical history, etc. During the extraction process, the device will strictly follow data security and privacy protection regulations, encrypt sensitive information, and ensure the security of data transmission and storage.

[0039] After obtaining individual information, the data correction device cleans and preprocesses it. For missing data, the device will fill it in according to the distribution characteristics and statistical methods of the existing data. For example, if the weight data of some pregnant women is missing, the device will refer to the weight distribution of pregnant women of the same gestational age and race, and use the mean, median or predicted value based on the regression model to fill it in. For abnormal values, such as negative gestational age or values ​​far beyond the normal range, the device will mark them and correct or delete them according to the actual situation. At the same time, the device will convert date data in different formats into a standard format to facilitate subsequent processing.

[0040] The data correction device inputs the cleaned and preprocessed individual information into the personalized reference interval model. The researchers first collected a large amount of biomarker data from pregnant women with different individual characteristics. These data came from multiple medical institutions and covered a wide range of age, gestational age, weight range, and different ethnic characteristics. These biomarker data must be confirmed by experts in advance as not abnormal data, so that it is convenient to train an accurate personalized reference interval model in the future. The collected data is divided into training set, validation set, and test set. The training set is used for model training, the validation set is used to adjust the model hyperparameters, and the test set is used to evaluate the generalization ability of the model.

[0041] Model training uses machine learning algorithms, such as neural network algorithms in deep learning. Taking the multi-layer perceptron (MLP) as an example, the input layer receives the encoded individual information, the hidden layer extracts and combines the data through a series of nonlinear transformations, and the output layer outputs the personalized reference interval. During the training process, the model continuously adjusts the weights and biases of the network through the back-propagation algorithm to minimize the difference between the predicted reference interval and the actual biomarker data (such as the mean square error).

[0042] After training is completed, after evaluation by the validation set and the test set, and ensuring that the model performance is good, it is deployed to the data correction device. When the individual information of a pregnant woman is input, the model outputs a personalized reference range for this pregnant woman according to the pattern learned during training. For example, for a pregnant woman who is 30 years old, 16 weeks pregnant, and weighs 60 kg, the model may output a personalized reference range for serum pregnancy-associated plasma protein A (PAPP-A) of [30, 50] μg / L, a personalized reference range for alpha-fetoprotein (AFP) of [30, 50] μg / L, a reference range for unconjugated estriol (uE3) of [1.5, 3.0] nmol / L, and a reference range for beta-human chorionic gonadotropin (β-HCG) of [10,000, 30,000] mIU / mL. These ranges are obtained based on the model's analysis and learning of biomarker data of a large number of pregnant women with similar individual characteristics.

[0043] S202. Compare the final target biomarker sample data of each pregnant woman with the corresponding personalized reference range; when the data correction device compares the final target biomarker sample data with the personalized reference range, it uses an efficient data matching algorithm. For biomarker data such as serum pregnancy-associated plasma protein A (PAPP-A), alpha-fetoprotein (AFP), unconjugated estriol (uE3), and beta-human chorionic gonadotropin (β-HCG) of each pregnant woman, the device will compare them with the corresponding personalized reference ranges respectively. Taking serum pregnancy-associated plasma protein A (PAPP-A) as an example, assume that the final detected value of serum pregnancy-associated plasma protein A (PAPP-A) of a certain pregnant woman is 45 μg / L, and its personalized reference range is [30, 50] μg / L. The device first determines whether the detected value is within the reference range. Through a simple numerical comparison operation, it is confirmed that 45 is between 30 and 50, indicating that the detected value of serum pregnancy-associated plasma protein A (PAPP-A) of this pregnant woman is within the normal range. For alpha-fetoprotein (AFP), unconjugated estriol (uE3), and beta-human chorionic gonadotropin (β-HCG), the device performs similar comparison operations.

[0044] During the comparison process, the device will also consider the data accuracy and error range. Due to certain errors in the detection process, the device will determine a reasonable error range based on the accuracy parameters of the detection equipment and historical data statistical analysis. For example, if the accuracy of the detection equipment for serum pregnancy-associated plasma protein A (PAPP-A) detection is ±5 μg / L, then during the comparison, the device will expand the reference range to [25, 55] μg / L for comparison. If the detected value is within the expanded range and the deviation from the original reference range is within a reasonable range, it is still considered that the detected value is basically normal.

[0045] To improve the comparison efficiency, the device adopts parallel computing technology. For a batch of pregnant women's data, the biomarker data of different pregnant women are distributed to multiple computing cores or threads for simultaneous comparison operations. This can greatly shorten the time required for comparison and improve the data processing efficiency. At the same time, the device will record and store the comparison results in real time for subsequent analysis and query. If the detected value exceeds the reference interval, the device will record the direction of the excess (higher or lower) and the specific value of the excess, providing detailed information for subsequent risk assessment.

[0046] S203. Determine the risk value level of prenatal Down syndrome according to the matching degree between the final target biomarker sample data and the personalized reference interval. The risk value levels include low risk, medium risk, and high risk. The data correction device determines the risk value level based on a pre-set risk assessment rule base, combining biomarker data, pregnant women's individual characteristics, and the interrelationships of multiple risk factors. This rule base can be established based on a large amount of clinical research data and medical expert experience, comprehensively considering the deviation degree of biomarker data and reference intervals in the first and second trimesters, pregnant women's individual characteristics, and the interrelationships among multiple risk factors. Since the biomarkers detected in the first and second trimesters are different, risk assessments need to be carried out separately for these two trimesters.

[0047] In early pregnancy, the main tests are for serum pregnancy-associated plasma protein A (PAPP-A) and β-human chorionic gonadotropin (β-HCG) in early pregnancy. When the test values of both of these biomarkers are within the personalized reference range and the deviation from the central value of the range is small, it can be preliminarily determined that the risk of the pregnant woman having Down syndrome in early pregnancy is low. However, when the test value of a certain biomarker exceeds the reference range but the deviation degree is less than the preset deviation threshold, further comprehensive judgment is required. For example, if the test value of PAPP-A is slightly higher than the upper limit of the reference range while the test value of β-HCG is normal, the device will analyze in combination with the individual characteristics of the pregnant woman. If the pregnant woman is older (usually 35 years old and above) or has a family history of Down syndrome, even if only the biomarker PAPP-A is mildly abnormal, the device may determine the risk value level as medium risk. This is because for older pregnant women, the quality of their eggs relatively declines, increasing the probability of fetal chromosomal abnormalities; while for pregnant women with a family history of Down syndrome, the likelihood of carrying the relevant pathogenic genes is relatively high, increasing the risk of the fetus having Down syndrome. On the contrary, if there are no obvious risk factors in the individual characteristics of the pregnant woman, it may still be determined as low risk, but it will be marked as a low-risk state that requires close attention for subsequent enhanced monitoring. If the test values of both biomarkers, PAPP-A and β-HCG, exceed the reference range and the deviation degree is greater than the preset deviation threshold, it will be determined as high risk. Moreover, the device will subdivide different levels of risk according to the different deviation degrees, such as highly high risk, moderately high risk, etc. Highly high risk may mean that the test value of the biomarker deviates greatly from the normal range, and the possibility of the fetus having Down syndrome is very high; moderately high risk means that the deviation degree is relatively small, but still requires great attention. By subdividing the risk levels, doctors can more accurately assess the severity of the risk, thereby formulating more targeted diagnostic and intervention measures.

[0048] The risk assessment in the second trimester is mainly based on the data of three biomarkers: alpha-fetoprotein (AFP), unconjugated estriol (uE3), and beta-human chorionic gonadotropin (β-HCG) in the second trimester. If the test values of these three biomarkers are all within the personalized reference range and have a small deviation from the central value of the range, after comprehensive judgment, the device will consider that the pregnant woman has a low risk of having a fetus with Down syndrome in the second trimester. When the test value of a certain biomarker exceeds the reference range but the deviation is not large, it is also necessary to make a judgment by comprehensively considering other biomarkers and the individual characteristics of the pregnant woman. For example, if the AFP test value is slightly higher than the upper limit of the reference range, while the uE3 and β-HCG test values are normal, the device will further analyze the individual characteristics of the pregnant woman, such as age and family medical history. If the pregnant woman is older or has a family history of Down syndrome, even if only the AFP is slightly abnormal, the risk value level may be determined as medium risk; if there are no obvious risk factors in the individual characteristics, it may be determined as low risk that requires close attention. If the test values of multiple biomarkers exceed the reference range and the deviation degree is large, such as the AFP is significantly higher than the upper limit of the reference range, the uE3 is lower than the lower limit of the reference range, and the β-HCG also exceeds the normal range, the device will directly determine it as high risk and subdivide the risk level according to the deviation degree. Such subdivision helps doctors intuitively understand the severity of the risk. For highly high-risk cases, it may be recommended that the pregnant woman undergo more invasive examinations as soon as possible, such as amniocentesis or chorionic villus sampling, to determine whether the fetus has Down syndrome.

[0049] By conducting detailed risk assessments in the first trimester and the second trimester respectively, it is possible to provide doctors with a more accurate basis for judging the risk of a pregnant woman having a fetus with Down syndrome. During the risk assessment process, the device will also use machine learning algorithms for auxiliary judgment. By training a risk assessment model and inputting features such as biomarker data, individual information of the pregnant woman, and risk assessment results, the model can learn the relationship between different feature combinations and risk levels. During the actual assessment, the doctor inputs the relevant data of the pregnant woman into the risk assessment model, and the model outputs the predicted risk level, which is compared and verified with the judgment result of the rule base. If the two results are consistent, the risk level is confirmed; if there are differences, the device will initiate an artificial review process, and medical experts will make a final judgment based on the specific situation.

[0050] S204. Record the risk value level in the electronic health record of the target pregnant woman and send the risk value level to the management terminal for display.

[0051] The data calibration device will interact with the hospital's electronic health record system to accurately record the risk value level in the electronic health record of the target pregnant woman. This interaction process strictly follows relevant medical data transmission standards and security protocols, such as using the encrypted HTTPS protocol for data transmission to ensure the security and integrity of the data. After receiving the data, the electronic health record system will accurately store the risk value level data in the corresponding pregnant woman's file record according to its internal storage architecture and indexing mechanism. To facilitate subsequent query and statistical analysis, the system will establish a clear index for each risk value level record, for example, using the unique identifier of the pregnant woman (such as ID number or medical record number) and the evaluation time as a combined index. In this way, when doctors or other authorized personnel need to view the prenatal Down syndrome risk assessment of a pregnant woman, they can quickly and accurately retrieve the relevant information from the electronic health record system.

[0052] The data calibration device will also send the risk value level to the management end for display. The management end can be the doctor workstation in the hospital, the remote medical management platform or other authorized medical information management terminals. To ensure the security of data transmission, the device will encrypt the data. After receiving the risk value level data, the management end will display it to doctors or managers through an intuitive and easy-to-understand interface. On the doctor workstation, the risk value level will be displayed with prominent colors and icons. For example, low risk is represented by a green icon, medium risk by a yellow icon, and high risk by a red icon. At the same time, the relevant information of the pregnant woman and the detailed content of the risk assessment will also be displayed to facilitate doctors to quickly understand the situation. For pregnant women with high risk, the system will also provide a warning prompt function to remind doctors to intervene and conduct further diagnosis in a timely manner.

[0053] In the embodiment of this application, by obtaining the individual information of the pregnant woman and inputting it into the personalized reference interval model, comparing the biomarker sample data with the personalized reference interval, determining the risk value level according to the matching degree, and recording and displaying the risk value level, the personalized, precise and efficient prenatal Down syndrome risk assessment is realized. It not only effectively solves the problem of risk assessment deviation caused by the inability of the unified reference interval in the prior art to adapt to individual differences, but also improves the accuracy and reliability of the risk assessment results, provides more targeted diagnostic references for doctors, helps to detect potential risks in a timely manner and take corresponding measures, which is of great significance for ensuring the health of pregnant women and fetuses.

[0054] The following introduces the data calibration device in the embodiment of this application from the perspective of modules. Please refer to Figure 3 , which is a schematic diagram of a module structure of the data calibration device in the embodiment of this application.

[0055] A sample data acquisition module 301 is configured to acquire multiple target biomarker sample data, where the target biomarker sample data at least includes serum pregnancy-associated plasma protein A and β-human chorionic gonadotropin in the early pregnancy, or alpha-fetoprotein, free estriol, and β-human chorionic gonadotropin in the second trimester; Multiple information acquisition modules 302 are configured to acquire detection device information, storage environment information, and storage duration information; An influence factor determination module 303 is configured to combine the detection device information, storage environment information, and storage duration information, and determine a first influence factor corresponding to the detection device information, a second influence factor corresponding to the storage environment information, and a third influence factor corresponding to the storage duration information through a data influence factor model. The data influence factor model is obtained through machine learning training in advance according to a detection device data set, a storage environment data set, and a storage duration data set with corresponding influence factor annotation information. The data influence factor model is used to quantify the influence degree of different detection devices, different storage environment conditions, and different storage durations on the accuracy of sample data; A final sample data determination module 304 is configured to correct the target biomarker sample data according to the first influence factor, the second influence factor, and the third influence factor to determine the final target biomarker sample data; An auxiliary judgment data determination module 305 is configured to determine the final target biomarker sample data as prenatal Down syndrome auxiliary judgment data.

[0056] The above describes the data correction device in the embodiments of the present application from the perspective of modular functional entities. The following describes the data correction device in the embodiments of the present invention from the perspective of hardware processing. Please refer to Figure 4 which is a schematic hardware structure diagram of the data correction device in the embodiments of the present invention.

[0057] It should be noted that Figure 4 the structure of the data correction device shown is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present invention.

[0058] Such as Figure 4As shown, the data correction device includes a Central Processing Unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 402 or the program loaded from the storage section 408 into the Random Access Memory (RAM) 403, such as executing the method described in the above embodiments. In the RAM 403, various programs and data required for system operation are also stored. The CPU 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An Input / Output (I / O) interface 405 is also connected to the bus 404.

[0059] The following components are connected to the I / O interface 405: an input section 406 including an audio input device, a button switch, etc.; an output section 407 including a Liquid Crystal Display (LCD), an audio output device, an indicator light, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read from it can be installed into the storage section 408 as needed.

[0060] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by the Central Processing Unit (CPU) 401, various functions defined in the present invention are executed.

[0061] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0062] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings.

[0063] Specifically, the data correction device in this embodiment includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, it implements the method for improving the accuracy of prenatal Down syndrome data provided in the above-mentioned embodiment.

[0064] On the other hand, the present invention also provides a computer-readable storage medium. This storage medium may be included in the data correction device described in the above-mentioned embodiment; or it may exist separately and not be assembled into the data correction device. The above-mentioned storage medium carries one or more computer programs. When the one or more computer programs are executed by a processor of the data correction device, the data correction device is enabled to implement the method for improving the accuracy of prenatal Down syndrome data provided in the above-mentioned embodiment.

[0065] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present application.

[0066] As used in the foregoing embodiments, depending on the context, the term "when" may be interpreted to mean "if" or "after" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "upon determining" or "if (the stated condition or event) is detected" may be interpreted to mean "if determined" or "in response to determining" or "when (the stated condition or event) is detected" or "in response to detecting (the stated condition or event)".

[0067] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the foregoing embodiments can be implemented by a computer program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the foregoing method embodiments. The foregoing storage medium includes: various media that can store program codes, such as ROM or random access memory RAM, magnetic disks, or optical discs.

Claims

1. A method for improving the accuracy of prenatal Down syndrome data, applied to a data correction device, characterized in that The method includes: Obtaining multiple target biomarker sample data, where the target biomarker sample data at least includes serum pregnancy-associated plasma protein A and β-human chorionic gonadotropin in early pregnancy, or alpha-fetoprotein, free estriol, and β-human chorionic gonadotropin in mid-pregnancy; Obtaining detection device information, storage environment information, and storage duration information; Combining the detection device information, storage environment information, and storage duration information, and determining a first influence factor corresponding to the detection device information, a second influence factor corresponding to the storage environment information, and a third influence factor corresponding to the storage duration information through a data influence factor model. The data influence factor model is obtained in advance through machine learning training based on a detection device data set, a storage environment data set, and a storage duration data set with corresponding influence factor annotation information. The data influence factor model is used to quantify the influence degree of different detection devices, different storage environment conditions, and different storage durations on the accuracy of sample data; Correcting the target biomarker sample data according to the first influence factor, the second influence factor, and the third influence factor to determine the final target biomarker sample data; Determining the final target biomarker sample data as prenatal Down syndrome auxiliary judgment data.

2. The method according to claim 1, characterized in that After the step of combining the detection device information, storage environment information, and storage duration information and determining a first influence factor corresponding to the detection device information, a second influence factor corresponding to the storage environment information, and a third influence factor corresponding to the storage duration information through a data influence factor model, the following steps are further included: Inputting a time series and determining the optimal time data for storing the target biomarker sample data through the data influence factor model; Determining an optimal time threshold for the target biomarker sample data according to the optimal time data; If the storage time exceeds the optimal time threshold, sending a reminder to the user terminal.

3. The method according to claim 1, wherein In the step of correcting the target biomarker sample data according to the first influence factor, the second influence factor, and the third influence factor to determine the final target biomarker sample data, specifically includes: Determine the first interaction term F of the first influencing factor and the second influencing factor 12 = F1 × F2; Determine the second interaction term F of the second influencing factor and the third influencing factor 23 = F2 × F3; Determine the third interaction term F of the first influencing factor and the third influencing factor 13 = F1 × F3; Determine the total interaction term of the first influence factor, the second influence factor, and the third influence factor, where the total interaction term is F 123 = F1 × F2 × F3.

4. The method according to claim 3, characterized in that, After the step of determining the total interaction term of the first influence factor, the second influence factor, and the third influence factor, the following steps are further included: Correct the target biomarker sample data according to the comprehensive correction formula, and the comprehensive correction formula is as follows: where D is the target biomarker sample data, is the final target biomarker sample data, and a1, a2, a3, a 13 , a 23 , a 12 , a 123 are coefficients obtained by fitting experimental data.

5. The method according to claim 1, wherein In the step of determining the final target biomarker sample data as prenatal Down syndrome auxiliary judgment data, specifically includes: After obtaining the individual information of the pregnant woman, inputting the individual information into a personalized reference interval model to obtain a personalized reference interval. The personalized reference interval model is obtained in advance by analyzing the biomarker data of multiple pregnant women with different individual characteristics and then training using machine learning algorithms; Comparing the final target biomarker sample data of each pregnant woman with the corresponding personalized reference interval; Determining the risk value level of prenatal Down syndrome according to the matching degree between the final target biomarker sample data and the personalized reference interval. The risk value level includes low risk, medium risk, and high risk; Record the risk value level in the electronic health record of the target pregnant woman, and send the risk value level to the management terminal for display.

6. The method according to claim 1, wherein After the steps of obtaining the detection device information, saving the environment information, and saving the duration information, the following steps are further included: Use the ROSE method to enhance the training data of the detection device information, the saved environment information, and the saved duration information; Use the preProcess method to standardize the training data and validation data of the detection device information, the saved environment information, and the saved duration information.

7. The method according to claim 1, characterized in that Before the step of correcting the target biomarker sample data according to the first influencing factor, the second influencing factor, and the third influencing factor to determine the final target biomarker sample data, the following steps are further included: After obtaining the individual information of the pregnant woman, adjust the final target biomarker sample data according to the individual information.

8. A data correction device, characterized in that, The data correction device includes: A sample data acquisition module, configured to acquire a plurality of target biomarker sample data, where the target biomarker sample data at least includes serum pregnancy-associated plasma protein A and β-human chorionic gonadotropin in the first trimester, or alpha-fetoprotein, free estriol, and β-human chorionic gonadotropin in the second trimester; A plurality of information acquisition modules, configured to acquire detection device information, saved environment information, and saved duration information; An influencing factor determination module, configured to combine the detection device information, the saved environment information, and the saved duration information, and determine a first influencing factor corresponding to the detection device information, a second influencing factor corresponding to the saved environment information, and a third influencing factor corresponding to the saved duration information through a data influencing factor model. The data influencing factor model is obtained through machine learning training in advance according to a detection device data set, a saved environment data set, and a saved duration data set with corresponding influencing factor annotation information. The data influencing factor model is used to quantify the influence degree of different detection devices, different saved environment conditions, and different saved durations on the accuracy of sample data; A final sample data determination module, configured to correct the target biomarker sample data according to the first influencing factor, the second influencing factor, and the third influencing factor to determine the final target biomarker sample data; An auxiliary judgment data determination module, configured to determine the final target biomarker sample data as prenatal Down syndrome auxiliary judgment data.

9. A data correction device, characterized in that, The data correction device includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code. The computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the data correction device to execute the method according to any one of claims 1-7.

10. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the data correction device, it enables the data correction device to execute the method according to any one of claims 1-7.