Blood collection tube sample data integration method and system
Through the methods of generating barcodes, fuzzy matching algorithms, data cleaning and blood database verification, the problems of high cost and insufficient standardization in the blood collection management system are solved, the standardization and safety of blood management are achieved, and work efficiency and resource sharing are improved.
Patent Information
- Application Number
- CN202510157270.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The existing blood collection and management system has problems such as high procurement costs, lack of standardized management and lax information management, which makes it difficult to ensure blood safety.
By generating barcodes, collecting information and entering, using the fuzzy matching algorithm to correlate blood vessel sample data from different sources, performing assays and data cleaning, using first-order linear interpolation method to fill in missing values, and establishing a blood database for verification, eliminating outliers, realizing unique identification and standardized management of data.
It improves the standardization and security of blood management, reduces work burden, realizes resource sharing and quick access to information, and improves work efficiency.
Smart Images

Figure CN119597747B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and specifically to a method and system for integrating blood collection tube sample data. Background Art
[0002] Blood is an essential resource in medical treatment. However, during the blood collection and transmission process, unsafe factors are extremely likely to occur. Therefore, in order to ensure the quality and safety of blood, it is necessary to use the implementation of information transfer control and comparison to control blood safety. With the development of computer technology, using information technology to manage blood is an important means to improve blood quality. Although some blood stations have established blood collection and management systems, there are still many problems in actual work, such as high procurement costs for blood collection institutions, lax external audit procedures for blood information management systems, and lack of standardization in management. Summary of the Invention
[0003] To solve the above technical problems, a method and system for integrating blood collection tube sample data are provided, and this technical solution solves the problems raised in the above background art.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0005] A method for integrating blood collection tube sample data, including:
[0006] During the blood collection process, generate a barcode based on the collection information of the sample, and enter the collection information into the blood collection tube sample data. The collection information includes patient information, collection date, and collection site;
[0007] Number each sample data based on the barcode of the blood collection tube sample data, and use the number as the unique identifier of the sample data;
[0008] Use the fuzzy matching algorithm to associate the blood collection tube sample data from different sources;
[0009] Perform a test on the blood collection tube sample and merge the test results with the corresponding blood collection tube sample data;
[0010] Perform preliminary data cleaning on the collected blood collection tube sample data to remove duplicate and redundant blood collection tube sample data;
[0011] Use the first-order linear interpolation method to fill in the missing values in the blood collection tube sample data;
[0012] Establish a blood database for storing blood collection tube sample data, and enter the associated blood collection tube sample data into the blood database;
[0013] Use the interquartile range method to verify the blood collection tube sample data in the blood database, and detect and remove outliers in the blood collection tube sample data;
[0014] The user obtains the number of the corresponding blood collection tube sample data by scanning the bar code, and obtains the test result based on the said number.
[0015] Preferably, the specific steps of using the fuzzy matching algorithm to associate the blood collection tube sample data from different sources include:
[0016] Create a summary record of basic standardized matching, and the summary record includes the matching data types that have been repaired;
[0017] Select and map attributes based on the fuzzy matching that will occur;
[0018] Select a fuzzy matching technique for each attribute. The text information is matched according to the keyboard distance and name variants, and the numerical information is matched according to the numerical similarity index;
[0019] Preset weights for each attribute. The attributes with high weights have a greater impact on the overall matching confidence than those with low weights;
[0020] Run the fuzzy matching algorithm and calculate the fuzzy matching confidence of the blood collection tube sample data from different sources;
[0021] Preset a matching confidence threshold;
[0022] Judge whether the fuzzy matching confidence of the blood collection tube sample data from different sources is higher than the preset matching confidence threshold. If so, output the blood collection tube sample data from different sources of this group as associated data. If not, do not output;
[0023] Associate the blood collection tube sample data of the same patient at different times that are matched, and merge and output them as a dataset of the same patient.
[0024] Preferably, the specific steps of performing preliminary data cleaning on the collected blood collection tube sample data to remove duplicate and redundant blood collection tube sample data include:
[0025] Obtain a dataset of the same patient;
[0026] Judge whether the attribute values between the records in the dataset of the same patient are exactly equal. If so, output the equal records as duplicate data and merge the equal records into one record. If not, do not output;
[0027] Perform multiple regression on each response variable in the blood collection tube sample data with the explanatory variables in the blood collection tube sample data respectively;
[0028] Obtain the fitted value of each response variable through the regression model, that is, the value corresponding on the regression line;
[0029] Output the difference between the observed value and the fitted value of the response variable as the residual.
[0030] Output the fitted value matrix and the residual matrix that contain the fitted values and residuals of all response variables;
[0031] Perform principal component analysis on the fitted value matrix to obtain the canonical eigenvector matrix;
[0032] Use the original data matrix to obtain the quadrat ordination coordinates in the original variable space, and set the obtained coordinates as the quadrat scores;
[0033] Use the fitted value matrix to obtain the quadrat ordination coordinates in the explanatory variable space, and set the obtained coordinates as the quadrat constraints;
[0034] Add the quadrat scores and the quadrat constraints to obtain the redundancy scores of each data;
[0035] Judge whether the redundancy scores of each data are lower than the preset redundancy threshold. If so, output that the data is redundant data and delete the data record. If not, do not output.
[0036] Preferably, the method for filling the missing values in the blood collection tube sample data by using the first-order linear interpolation method specifically includes:
[0037] Read the blood collection tube sample data in the same patient dataset, and mark the known data points and the missing data points;
[0038] Select two known data points adjacent to the missing data point as the basis for interpolation;
[0039] Calculate the interpolation point value of the missing data point by using the first-order linear interpolation formula;
[0040] Replace the missing values in the same patient dataset with the calculated interpolation point values;
[0041] The first-order linear interpolation formula is: ,
[0042] In the formula, are respectively the collection date value and the interpolation point value of the missing data point, are respectively the collection date value and the data value of the adjacent known data point before the missing data point, are respectively the collection date value and the data value of the adjacent known data point after the missing data point.
[0043] Preferably, the method for establishing a blood database for storing the blood collection tube sample data and inputting the associated blood collection tube sample data into the blood database specifically includes:
[0044] Construct an entity-relationship model to describe the relationships between blood collection tube sample data. In the blood database, the entities include the patient entity, the blood collection tube sample entity, and the test result entity, and determine the association relationships between the entities;
[0045] Convert the entity-relationship model into the logical structure of the blood database, and the logical structure includes table structure, field type, and index;
[0046] Based on the logical structure design of the blood database, determine the storage structure, storage path, and storage device of the database;
[0047] Enter the associated blood collection tube sample data into the blood database.
[0048] Preferably, using the interquartile range method to verify the blood collection tube sample data in the blood database, detecting and removing outliers in the blood collection tube sample data specifically includes: sorting the blood collection tube sample data;
[0049] Calculate the data point at the 25% position in the dataset and output it as the first quartile;
[0050] Calculate the data point at the 75% position in the dataset and output it as the third quartile;
[0051] Use the interquartile range method formula to calculate the normal value range of the sample data;
[0052] Judge whether the blood collection tube sample data exceeds the normal value range of the sample data. If so, output the data as an outlier, remove the item from the database, and use the first-order linear interpolation method to fill in the missing values. If not, do not output;
[0053] The interquartile range method formula is: ,
[0054] In the formula, is the normal value of the sample data, is the first quartile, is the third quartile.
[0055] Furthermore, a blood collection tube sample data integration system is proposed to implement the blood collection tube sample data integration method as described above, including:
[0056] A data acquisition module, which is used to generate a barcode based on the collection information of the sample during the blood collection process and enter the collection information into the blood collection tube sample data;
[0057] A data processing module, the data processing module is used to number each sample data based on the barcode of the blood collection tube sample data, and use the number as the unique identifier of the sample data, use a fuzzy matching algorithm to associate blood collection tube sample data from different sources, test the blood collection tube samples, and merge the test results with the corresponding blood collection tube sample data, perform preliminary data cleaning processing on the collected blood collection tube sample data, remove duplicate and redundant blood collection tube sample data, and use a first-order linear interpolation method to fill in missing values in the blood collection tube sample data;
[0058] A data integration module is used to establish a blood database for storing blood collection tube sample data, enter the associated blood collection tube sample data into the blood database, verify the blood collection tube sample data in the blood database using the interquartile range method, detect and remove abnormal values in the blood collection tube sample data, and the user obtains the number of the corresponding blood collection tube sample data by scanning the barcode, and obtains the test result based on the number.
[0059] Optionally, the data processing module specifically includes:
[0060] A data numbering unit, which is used to number each sample data based on the barcode of the blood collection tube sample data, and use the number as a unique identifier of the sample data;
[0061] A data association unit, the data association unit is used to associate blood collection tube sample data from different sources using a fuzzy matching algorithm;
[0062] A sample testing unit, which is used to test the blood collection tube sample and merge the test result with the corresponding blood collection tube sample data;
[0063] A data cleaning unit, the data cleaning unit is used to perform preliminary data cleaning processing on the collected blood collection tube sample data to remove duplicate and redundant blood collection tube sample data;
[0064] A data interpolation unit is used to fill in the missing values in the blood collection tube sample data using a first-order linear interpolation method.
[0065] Optionally, the data integration module specifically includes:
[0066] A database establishment unit, the database establishment unit is used to establish a blood database for storing blood collection tube sample data, and enter the associated blood collection tube sample data into the blood database;
[0067] A data verification unit, the data verification unit is used to verify the blood collection tube sample data in the blood database using the interquartile range method, and detect and remove abnormal values in the blood collection tube sample data;
[0068] A user query unit, which is used for the user to obtain the number of the corresponding blood collection tube sample data by scanning the bar code, and obtain the test result based on the number.
[0069] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0070] By setting up a data acquisition module, a data processing module and a data integration module, taking advantage of the advantages of computers in operation, storage and transmission, eliminating unnecessary working links, improving work efficiency, reducing work burden, increasing operation standardization, all information of this system can be accessed by all legal users on the network, and they can conveniently and quickly obtain the information within their permissions, realizing resource sharing. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 is a flowchart of the method for integrating blood collection tube sample data of the present invention;
[0072] Figure 2 is a flowchart of the method for associating blood collection tube sample data from different sources by using a fuzzy matching algorithm of the present invention;
[0073] Figure 3 is a flowchart of the method for preliminarily cleaning and processing the collected blood collection tube sample data of the present invention;
[0074] Figure 4 is a flowchart of the method for filling in missing values in the blood collection tube sample data of the present invention;
[0075] Figure 5 is a flowchart of the method for establishing a blood database for storing blood collection tube sample data of the present invention;
[0076] Figure 6 is a flowchart of the method for verifying the blood collection tube sample data in the blood database by using the interquartile range method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations.
[0078] Referring to Figure 1 as shown, the method for integrating blood collection tube sample data includes:
[0079] During the blood collection process, a bar code is generated based on the collection information of the sample, and the collection information is entered into the blood collection tube sample data, and the collection information includes patient information, collection date and collection site;
[0080] Barcodes are used to number each sample data based on the blood collection tube sample data, and the number is used as the unique identifier of the sample data;
[0081] The fuzzy matching algorithm is used to associate the blood collection tube sample data from different sources;
[0082] The blood collection tube samples are tested, and the test results are merged with the corresponding blood collection tube sample data;
[0083] The collected blood collection tube sample data is initially processed for data cleaning to remove duplicate and redundant blood collection tube sample data;
[0084] The first-order linear interpolation method is used to fill in the missing values in the blood collection tube sample data;
[0085] A blood database is established to store the blood collection tube sample data, and the associated blood collection tube sample data is entered into the blood database;
[0086] The quartile range method is used to verify the blood collection tube sample data in the blood database, and the outliers in the blood collection tube sample data are detected and removed;
[0087] The user obtains the number of the corresponding blood collection tube sample data by scanning the barcode and obtains the test result based on the number.
[0088] Refer to Figure 2 As shown, using the fuzzy matching algorithm to associate the blood collection tube sample data from different sources specifically includes:
[0089] Create a summary record for basic standardized matching, and the summary record includes the repaired matching data types;
[0090] Select and map attributes based on the fuzzy matching to occur;
[0091] Select a fuzzy matching technique for each attribute. The text information is matched according to the keyboard distance and name variants, and the numerical information is matched according to the numerical similarity index;
[0092] Preset weights for each attribute. Attributes with high weights have a greater impact on the overall matching confidence than those with low weights;
[0093] Run the fuzzy matching algorithm and calculate the fuzzy matching confidence of the blood collection tube sample data from different sources;
[0094] Preset the matching confidence threshold;
[0095] Judge whether the fuzzy matching confidence of the blood collection tube sample data from different sources is higher than the preset matching confidence threshold. If so, output the blood collection tube sample data from this group of different sources as associated data. If not, do not output;
[0096] Associate the data of blood collection tube samples at different times that match the same patient, and merge and output them as a dataset for the same patient.
[0097] Fuzzy matching is a data matching technique that compares two or more records and calculates the likelihood that they belong to the same entity. Fuzzy matching does not result in a match or non-match, but rather a percentage that describes the likelihood that the record belongs to the same customer, product, employee, etc. An effective fuzzy matching algorithm can handle a range of data ambiguities, such as name / surname reversals, acronyms, abbreviations, phonetic and intentional spelling mistakes, contractions, added / removed punctuation, etc.
[0098] Refer to Figure 3 As shown, perform preliminary data cleaning on the collected blood collection tube sample data to remove duplicate and redundant blood collection tube sample data, specifically including:
[0099] Obtain a dataset for the same patient;
[0100] Judge whether the attribute values between the records in the same patient dataset are exactly equal. If so, output the equal records as duplicate data and merge the equal records into one record. If not, do not output.
[0101] Perform multiple regression on each response variable in the blood collection tube sample data with the explanatory variables in the blood collection tube sample data;
[0102] Obtain the fitted value of each response variable through the regression model, that is, the value corresponding on the regression line;
[0103] Output the difference between the observed value and the fitted value of the response variable as the residual;
[0104] Output the fitted value matrix and the residual matrix containing all the fitted values and residuals of the response variables;
[0105] Perform principal component analysis on the fitted value matrix to obtain the canonical eigenvector matrix;
[0106] Use the original data matrix to obtain the quadrat sorting coordinates in the original variable space, and set the obtained coordinates as the quadrat scores;
[0107] Use the fitted value matrix to obtain the quadrat sorting coordinates in the explanatory variable space, and set the obtained coordinates as the quadrat constraints;
[0108] Add the quadrat scores and the quadrat constraints to obtain the redundancy scores of each data;
[0109] Judge whether the redundancy scores of each data are lower than the preset redundancy threshold. If so, output the data as redundant data and delete the data record. If not, do not output.
[0110] In statistics, redundancy analysis is to analyze the causes of the variation of the original variables through the correlation between the original variables and the canonical variables. Specifically, taking the original variables as the dependent variables and the canonical variables as the independent variables, a linear regression model is established, and then the corresponding coefficient of determination is equal to the square of the correlation coefficient between the dependent variable and the canonical variable. It describes the proportion of the variation of the dependent variable caused by the linear relationship between the dependent variable and the canonical variable in the total variation of the dependent variable.
[0111] Refer to Figure 4 As shown, using the first-order linear interpolation method to fill in the missing values in the blood collection tube sample data specifically includes:
[0112] Read the blood collection tube sample data in the same patient dataset and mark the known data points and missing data points;
[0113] Select two adjacent known data points to the missing data point as the basis for interpolation;
[0114] Use the first-order linear interpolation formula to calculate the interpolation point value of the missing data point;
[0115] Replace the missing value in the same patient dataset with the calculated interpolation point value;
[0116] The first-order linear interpolation formula is: ,
[0117] In the formula, are respectively the collection date value and the interpolation point value of the missing data point, are respectively the collection date value and the data value of the adjacent known data point before the missing data point, are respectively the collection date value and the data value of the adjacent known data point after the missing data point.
[0118] Interpolation usually refers to interpolation, which is a term in discrete mathematics. It refers to interpolating a continuous function on the basis of discrete data so that the continuous curve passes through all the given discrete data points. As an important method for approximating discrete functions, interpolation can be used to estimate the approximate values of the function at other points according to the values of the function at a finite number of points.
[0119] Refer to Figure 5 As shown, establishing a blood database for storing blood collection tube sample data and inputting the associated blood collection tube sample data into the blood database specifically includes:
[0120] Construct an entity-relationship model to describe the relationships between the blood collection tube sample data. In the blood database, the entities include the patient entity, the blood collection tube sample entity, and the test result entity, and determine the association relationships between the entities;
[0121] Convert the entity-relationship model into the logical structure of the blood database, where the logical structure includes table structure, field type, and index;
[0122] Based on the design of the logical structure of the blood database, determine the storage structure, storage path, and storage device of the database;
[0123] Enter the associated blood collection tube sample data into the blood database.
[0124] For the blood database, common database systems include relational databases such as MySQL and NoSQL databases. Relational databases are suitable for storing structured data, while NoSQL databases are more suitable for processing unstructured or semi-structured data.
[0125] Refer to Figure 6 As shown, use the interquartile range method to verify the blood collection tube sample data in the blood database, and detect and remove outliers in the blood collection tube sample data, which specifically includes:
[0126] Sort the blood collection tube sample data;
[0127] Calculate the data point at the 25% position in the dataset and output it as the first quartile;
[0128] Calculate the data point at the 75% position in the dataset and output it as the third quartile;
[0129] Use the interquartile range method formula to calculate the normal value range of the sample data;
[0130] Judge whether the blood collection tube sample data exceeds the normal value range of the sample data. If so, output the data as an outlier, remove the data item from the database, and use the first-order linear interpolation method to fill in the missing values. If not, do not output;
[0131] The interquartile range method formula is: ,
[0132] In the formula, is the normal value of the sample data, is the first quartile, is the third quartile.
[0133] The interquartile range method is a commonly used statistical method for detecting outliers or extreme values in a dataset. This method does not rely on the assumption of normal distribution of the data, so it is particularly suitable for data with non-normal distribution or skewed distribution. In the verification of blood collection tube sample data in the blood database, the interquartile range method can effectively identify and remove outliers, thereby improving the accuracy and reliability of the data.
[0134] Further, based on the same inventive concept as the above blood collection tube sample data integration method, the present solution also proposes a blood collection tube sample data integration system, comprising:
[0135] A data collection module, which is used to generate a barcode based on the sample collection information during the blood collection process, and enter the collection information into the blood collection tube sample data;
[0136] A data processing module, the data processing module is used to number each sample data based on the barcode of the blood collection tube sample data, and use the number as the unique identifier of the sample data, use a fuzzy matching algorithm to associate blood collection tube sample data from different sources, test the blood collection tube samples, and merge the test results with the corresponding blood collection tube sample data, perform preliminary data cleaning processing on the collected blood collection tube sample data, remove duplicate and redundant blood collection tube sample data, and use a first-order linear interpolation method to fill in missing values in the blood collection tube sample data;
[0137] A data integration module is used to establish a blood database for storing blood collection tube sample data, enter the associated blood collection tube sample data into the blood database, verify the blood collection tube sample data in the blood database using the interquartile range method, detect and remove abnormal values in the blood collection tube sample data, and the user obtains the number of the corresponding blood collection tube sample data by scanning the barcode, and obtains the test result based on the number.
[0138] The data processing module specifically includes:
[0139] A data numbering unit, which is used to number each sample data based on the barcode of the blood collection tube sample data, and use the number as a unique identifier of the sample data;
[0140] A data association unit, the data association unit is used to associate blood collection tube sample data from different sources using a fuzzy matching algorithm;
[0141] A sample testing unit, which is used to test the blood collection tube sample and merge the test result with the corresponding blood collection tube sample data;
[0142] A data cleaning unit, the data cleaning unit is used to perform preliminary data cleaning processing on the collected blood collection tube sample data to remove duplicate and redundant blood collection tube sample data;
[0143] A data interpolation unit is used to fill in the missing values in the blood collection tube sample data using a first-order linear interpolation method.
[0144] The data integration module specifically includes:
[0145] A database establishment unit, which is used to establish a blood database for storing blood collection tube sample data and input the associated blood collection tube sample data into the blood database;
[0146] A data verification unit, which is used to verify the blood collection tube sample data in the blood database by using the interquartile range method, and detect and eliminate outliers in the blood collection tube sample data;
[0147] A user query unit, which is used for the user to obtain the number of the corresponding blood collection tube sample data by scanning the barcode and obtain the test result based on the number.
[0148] Furthermore, this solution also proposes a computer-readable storage medium, on which a computer-readable program is stored. When the computer-readable program is called, it executes the above-mentioned method for integrating blood collection tube sample data.
[0149] It can be understood that the storage medium can be a magnetic medium, such as a floppy disk, a hard disk, a magnetic tape; an optical medium such as a DVD; or a semiconductor medium such as a solid state disk (SSD), etc.
[0150] In summary, the advantages of the present invention are as follows: by setting up a data acquisition module, a data processing module and a data integration module, taking advantage of the advantages of a computer in terms of operation, storage and transmission, eliminating unnecessary working links, improving work efficiency, reducing the work burden, increasing operation standardization, all the information of this system can be accessed by all legal users on the network, and they can conveniently and quickly obtain the information within their authority scope to achieve resource sharing.
[0151] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for integrating blood collection tube sample data, characterized in that: include: During the blood collection process, a barcode is generated based on the sample collection information, and the collection information is entered into the blood collection tube sample data, wherein the collection information includes patient information, collection date and collection site; Number each sample data based on the barcode of the blood collection tube sample data, and use the number as the unique identifier of the sample data; The fuzzy matching algorithm is used to associate blood collection tube sample data from different sources; Testing the blood collection tube samples and merging the test results with the corresponding blood collection tube sample data; Perform preliminary data cleaning on the collected blood collection tube sample data to remove duplicate and redundant blood collection tube sample data; First-order linear interpolation was used to fill in missing values in blood collection tube sample data; Establishing a blood database for storing blood collection tube sample data, and entering the associated blood collection tube sample data into the blood database; The interquartile range method is used to verify the blood collection tube sample data in the blood database, and abnormal values in the blood collection tube sample data are detected and eliminated; The user obtains the number of the corresponding blood collection tube sample data by scanning the barcode, and obtains the test result based on the number; The preliminary data cleaning process of the collected blood collection tube sample data to remove duplicate and redundant blood collection tube sample data specifically includes: Obtain the same patient dataset; Determine whether the attribute values of records in the same patient data set are completely equal. If so, output the equal records as duplicate data and merge the equal records into one record. If not, do not output them. Perform multiple regression on each response variable in the blood collection tube sample data and the explanatory variables in the blood collection tube sample data; The fitted value of each response variable is obtained through the regression model, that is, the corresponding value on the regression line; Output the difference between the observed and fitted values of the response variable as residuals; Output includes the fitted value matrix and residual matrix of all response variables; Perform principal component analysis on the fitted value matrix to obtain the canonical eigenvector matrix; The original data matrix is used to obtain the sample plot order coordinates in the original variable space, and the obtained coordinates are set as the sample plot scores; Use the fitted value matrix to obtain the sample plot order coordinates in the explanatory variable space, and the obtained coordinates are set as sample plot constraints; Add the sample score and the sample constraint to get the redundancy score of each data; Determine whether the redundancy score of each data is lower than the preset redundancy threshold. If so, output the data as redundant data and delete the data record. If not, do not output it.
2. The blood collection tube sample data integration method according to claim 1, characterized in that: The method of associating blood collection tube sample data from different sources using a fuzzy matching algorithm specifically includes: creating a summary record of the basic standardized match, the summary record including the type of matched data that has been repaired; Select and map attributes based on the fuzzy matching that will occur; A fuzzy matching technique is selected for each attribute. Text information is matched based on keyboard distance and name variants, and numeric information is matched based on numeric similarity indicators. Preset weights for each attribute, and attributes with high weights have a greater impact on the overall matching confidence than attributes with low weights; Run the fuzzy matching algorithm and calculate the fuzzy matching confidence of blood collection tube sample data from different sources; Preset matching confidence threshold; Determine whether the fuzzy matching confidence of blood collection tube sample data from different sources is higher than a preset matching confidence threshold, if so, output the group of blood collection tube sample data from different sources as associated data, if not, do not output; The blood collection tube samples from the same patient at different times are matched and associated and output as a dataset of the same patient.
3. The blood collection tube sample data integration method according to claim 2, characterized in that: The method of using the first-order linear interpolation method to fill the missing values in the blood collection tube sample data specifically includes: Read blood collection tube sample data from the same patient dataset and mark known data points and missing data points; Select two known data points that are immediately adjacent to the missing data point as the basis for interpolation; Use the first-order linear interpolation formula to calculate the interpolation point values of missing data points; The calculated interpolation point values are used to replace the missing values in the same patient data set; The first-order linear interpolation formula is: , In the formula, are the collection date value and interpolation point value of the missing data point, respectively. are the collection date value and data value of the known data point immediately before the missing data point, are the collection date value and data value of the known data point immediately following the missing data point, respectively.
4. The blood collection tube sample data integration method according to claim 3, characterized in that: The step of establishing a blood database for storing blood collection tube sample data and entering the associated blood collection tube sample data into the blood database specifically includes: Construct an entity-relationship model to describe the relationship between blood collection tube sample data. In the blood database, entities include patient entities, blood collection tube sample entities, and test result entities, and determine the association relationship between each entity; Converting the entity-relationship model into a logical structure of a blood database, wherein the logical structure includes a table structure, a field type, and an index; Based on the logical structure design of the blood database, determine the storage structure, storage path and storage device of the database; The associated blood collection tube sample data is entered into the blood database.
5. The blood collection tube sample data integration method according to claim 4, characterized in that: The method of using the interquartile range method to verify the blood collection tube sample data in the blood database and detecting and removing abnormal values in the blood collection tube sample data specifically includes: Sort the blood collection tube sample data; Calculate the data points at the 25% position in the data set and output them as the first quartile; Calculate the data points at the 75% position in the data set and output them as the third quartile; The interquartile range formula was used to calculate the normal value range of sample data; Determine whether the blood collection tube sample data exceeds the normal value range of the sample data. If so, output the data as an abnormal value, remove the data in the database, and use the first-order linear interpolation method to fill the missing value. If not, do not output it; The interquartile range formula is: , In the formula, is the normal value of the sample data, is the first quartile, The third quartile.
6. A blood collection tube sample data integration system, used to implement the blood collection tube sample data integration method according to any one of claims 1 to 5, characterized in that: include: A data collection module, which is used to generate a barcode based on the sample collection information during the blood collection process, and enter the collection information into the blood collection tube sample data; A data processing module, the data processing module is used to number each sample data based on the barcode of the blood collection tube sample data, and use the number as the unique identifier of the sample data, use a fuzzy matching algorithm to associate blood collection tube sample data from different sources, test the blood collection tube samples, and merge the test results with the corresponding blood collection tube sample data, perform preliminary data cleaning processing on the collected blood collection tube sample data, remove duplicate and redundant blood collection tube sample data, and use a first-order linear interpolation method to fill in missing values in the blood collection tube sample data; A data integration module is used to establish a blood database for storing blood collection tube sample data, enter the associated blood collection tube sample data into the blood database, verify the blood collection tube sample data in the blood database using the interquartile range method, detect and remove abnormal values in the blood collection tube sample data, and the user obtains the number of the corresponding blood collection tube sample data by scanning the barcode, and obtains the test result based on the number.
7. The blood collection tube sample data integration system according to claim 6, characterized in that: The data processing module specifically includes: A data numbering unit, which is used to number each sample data based on the barcode of the blood collection tube sample data, and use the number as a unique identifier of the sample data; A data association unit, the data association unit is used to associate blood collection tube sample data from different sources using a fuzzy matching algorithm; A sample testing unit, which is used to test the blood collection tube sample and merge the test result with the corresponding blood collection tube sample data; A data cleaning unit, the data cleaning unit is used to perform preliminary data cleaning processing on the collected blood collection tube sample data to remove duplicate and redundant blood collection tube sample data; A data interpolation unit is used to fill in the missing values in the blood collection tube sample data using a first-order linear interpolation method.
8. The blood collection tube sample data integration system according to claim 7, characterized in that: The data integration module specifically includes: A database establishment unit, the database establishment unit is used to establish a blood database for storing blood collection tube sample data, and enter the associated blood collection tube sample data into the blood database; A data verification unit, the data verification unit is used to verify the blood collection tube sample data in the blood database using the interquartile range method, and detect and remove abnormal values in the blood collection tube sample data; A user query unit is used for a user to obtain a number corresponding to the blood collection tube sample data by scanning a barcode, and to obtain a test result based on the number.
Citation Information
Patent Citations
Method and system for predicting water lettuce invasion distribution areas
CN111178631A
Matching data from variant databases
WO2014182725A1