A Medical Data Synchronization Consistency Verification Method and System Based on Elasticsearch Database

By synchronizing medical data by diagnosis and synchronizing it to the Elasticsearch database and comparing the slice time based on the number of visits, the existing medical data synchronization consistency verification methods are solved, and efficient data consistency verification is achieved.

CN119127821BActive Publication Date: 2025-06-20PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411174432.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2025-06-20
Estimated Expiration
2044-08-26

AI Technical Summary

Technical Problem

The existing medical data synchronization consistency verification methods are inefficient when facing scenarios with complex or inconsistent structures after data fusion, and have significant resource consumption and time costs when processing large-scale data.

Method used

By synchronizing the medical data in the distributed file database to the Elasticsearch database according to the diagnosis times, abnormal data during the synchronization process is identified to form an abnormal data information table, and the slice time size is calculated based on the number of visits. The corresponding data is extracted from the Elasticsearch database and the distributed file database according to the time range of each slice for comparison, and inconsistent data are found.

Benefits of technology

It reduces the impact of deep traversal, improves comparison efficiency, does not need to generate additional comparison files, reduces the need for storage resources, and quickly obtains consistency verification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119127821B_ABST
    Figure CN119127821B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for medical data synchronization consistency verification based on an Elasticsearch database, belonging to the technical field of medical data processing, and solves the problems of high storage resources and time consumption for consistency verification in the prior art. The method includes the following steps: synchronizing the medical data in the distributed file database to the Elasticsearch database by visit; identifying abnormal data in the aggregation synchronization process to form an abnormal data information table; calculating the slice time size based on the number of visits to obtain the time range of each slice, and extracting the medical data corresponding to each slice in the Elasticsearch database and the distributed file database according to the time range of each slice for comparison to find inconsistent data; comparing the abnormal data information table with the inconsistent data to obtain the consistency verification result. It realizes fast consistency verification and traceability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical data, and in particular to a method and system for verifying the synchronization consistency of medical data in an Elasticsearch database. Background Art

[0002] With the rapid development of hospital informatization, each medical institution has accumulated a large amount of medical data and successively established a resource data center (Research Data Repository, RDR) for clinical research applications. In response to clinical research needs, the resource data center performs data governance processes such as post-structuring and processing and cleaning of medical data, stores the data in the distributed file storage database MongoDB, and at the same time, in order to achieve flexible and efficient data retrieval functions (including full-text retrieval, cascading retrieval, etc.), and then realize functions such as natural language query, knowledge graph, machine question answering, big data analysis, and big data mining, the data is synchronously stored in the search engine database Elasticsearch on the basis of MongoDB, and complex retrieval is realized by nested storage of the entire diagnosis data of the patient.

[0003] During the process of synchronizing data from MongoDB to Elasticsearch, data inconsistency often occurs. The degree of data consistency is an important indicator for evaluating data quality and determines the reliability of subsequent statistical analysis results. Therefore, how to efficiently detect the consistency between the nested data in Elasticsearch and the data in the upstream MongoDB is an important problem that urgently needs to be solved in the construction of a high-quality resource data center, and at the same time has high scientific research and application value.

[0004] To perform a full-scale comparison of the two databases, usually the first data to be compared is converted into the first structured data, the second data to be compared is converted into the second structured data, and then the first structured data and the second structured data are compared.

[0005] Existing data comparison methods are highly efficient when the data formats are unified, but their effectiveness is limited when faced with scenarios where the post-structure is complex or inconsistent after data fusion.

[0006] In addition, existing methods need to generate additional comparison files during the execution process, which not only increases the occupation of storage resources but also increases the time consumption of reading and writing operations. Especially when dealing with large-scale data, this resource consumption and time cost become particularly significant. Summary of the Invention

[0007] In view of the above analysis, the embodiments of the present invention aim to provide a method and system for verifying the consistency of medical data synchronization based on the Elasticsearch database to solve the problem of high resource and time consumption in verifying the consistency of existing medical data synchronization.

[0008] On the one hand, the embodiments of the present invention provide a method for verifying the consistency of medical data synchronization based on the Elasticsearch database, including the following steps:

[0009] Synchronize the medical data in the distributed file database to the Elasticsearch database by aggregation according to the number of clinic visits;

[0010] Identify the abnormal data during the aggregation synchronization process to form an abnormal data information table;

[0011] Calculate the size of the slice time based on the number of clinic visits to obtain the time range of each slice, and extract the medical data corresponding to each slice in the Elasticsearch database and the distributed file database according to the time range of each slice for comparison to find inconsistent data;

[0012] Compare the abnormal data information table with the inconsistent data to obtain the consistency verification result.

[0013] For a further improvement based on the above method, the following formula is used to calculate the size of the slice time:

[0014]

[0015] where W represents the size of the slice time, k represents the limit number of clinic visits for the slice, V i represents the number of clinic visits in the i-th unit time, T represents the size of the unit time, and n represents the number of unit times from the minimum clinic visit time to the maximum clinic visit time span of the data table.

[0016] For a further improvement based on the above method, identifying the abnormal data during the aggregation synchronization process to form an abnormal data information table includes:

[0017] For each piece of medical data after aggregation, determine whether there are format abnormalities, duplicate abnormalities, or encryption abnormalities. If so, add the data ID with abnormalities and the corresponding reasons for abnormalities to the abnormal data information table. If not, store the data in the Elasticsearch database.

[0018] For a further improvement based on the above method, the following method is used to determine whether there are format abnormalities:

[0019] Determine whether the content of the numeric type field is a number. If not, there is a format exception. If it is a number, determine whether the numeric format of the numeric type field conforms to the preset format. If not, there is a format exception;

[0020] Determine whether the content of the date type field conforms to the preset format. If not, there is a format exception. If it conforms, determine whether the date is legal. If not, there is a format exception;

[0021] Determine whether the content of the boolean type field is the preset content. If not, there is a format exception.

[0022] Based on the further improvement of the above method, the following method is used to determine whether there is an encryption exception:

[0023] Decrypt the encrypted field and determine whether there is an encryption flag in the decrypted content. If not, there is an encryption exception;

[0024] For the de-privacy fields of structured data, determine whether the field content is the content after de-privacy. If not, there is an encryption exception;

[0025] For unstructured data, search for whether there is privacy content in it. If so, there is an encryption exception.

[0026] Based on the further improvement of the above method, extract the medical data corresponding to each slice in the Elasticsearch database and the distributed file database according to the time range of each slice for comparison, and search for inconsistent data, including:

[0027] Search for the medical data of the target field in the current slice time range in the Elasticsearch database and the distributed file database respectively and sort them in the time order of the target field; obtain the first data and the second data;

[0028] Filter the first data and delete the data where the target field exceeds the current slice time range;

[0029] If the earliest times of the target fields of the first data and the second data are different, then use the largest earliest time as the start time, and use the data where the target field time is less than the start time as the inconsistent data;

[0030] Search for the data that does not exist in both the first data and the second data starting from the data where the target field is the start time in the first data and the second data as the inconsistent data.

[0031] Based on the further improvement of the above method, compare the abnormal data information table with the inconsistent data to obtain the consistency check result, including:

[0032] For each piece of inconsistent data, if the ID of the data exists in the abnormal data information table, the abnormal reason corresponding to the ID in the abnormal data information table is extracted as the reason for the verification inconsistency. If the ID of the data does not exist in the abnormal data information table, the reason for the verification inconsistency is a synchronization error; thus, the consistency verification result is obtained.

[0033] On the other hand, an embodiment of the present invention provides a medical data synchronization consistency verification system based on an Elasticsearch database, including the following modules:

[0034] The aggregation synchronization module is used to synchronize the medical data in the distributed file database to the Elasticsearch database by visit;

[0035] The abnormal data recognition module is used to recognize abnormal data during the aggregation synchronization process to form an abnormal data information table;

[0036] The consistency verification module is used to calculate the slice time size based on the number of visits to obtain the time range of each slice, extract the medical data corresponding to each slice in the Elasticsearch database and the distributed file database according to the time range of each slice for comparison, and find inconsistent data; compare the abnormal data information table with the inconsistent data to obtain the consistency verification result.

[0037] For a further improvement of the above system, the following formula is used to calculate the slice time size:

[0038]

[0039] Where, W represents the slice time size, k represents the number limit of visits for slicing, V i represents the number of visits in the i-th unit time, T represents the size of the unit time, and n represents the number of unit times from the minimum visit time to the maximum visit time span of the data table.

[0040] For a further improvement of the above system, extracting the medical data corresponding to each slice in the Elasticsearch database and the distributed file database according to the time range of each slice for comparison and finding inconsistent data includes:

[0041] Search for the medical data of the target field within the current slice time range in the Elasticsearch database and the distributed file database respectively and sort them in the time order of the target field; obtain the first data and the second data;

[0042] Filter the first data to delete the data where the target field exceeds the current slice time range;

[0043] If the earliest times of the target fields of the first data and the second data are different, then take the largest earliest time as the starting time, and take the data whose target field time is less than the starting time as inconsistent data;

[0044] In the first data and the second data, start searching from the data with the target field being the starting time for the data that does not exist in both the first data and the second data simultaneously as inconsistent data.

[0045] Compared with the prior art, the present invention aggregates and synchronizes medical data in a distributed file database by visit times to an Elasticsearch database, identifies abnormal numbers during the synchronization process to form an abnormal data information table. For the comparison of the entire amounts of the two databases, by calculating the slice time size based on the number of visits, corresponding data is extracted from the Elasticsearch database and the distributed file database respectively according to the time range of each slice for comparison, thereby reducing the impact of deep traversal, improving the comparison efficiency, and not requiring the generation of additional comparison files, reducing the need for storage resources. By searching for inconsistent data in the two databases, and then comparing the abnormal data information table with the inconsistent data, the consistency verification result can be quickly obtained.

[0046] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combined solutions. Other features and advantages of the present invention will be described in the subsequent specification, and some advantages can be made obvious from the specification, or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the content specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The drawings are only for the purpose of showing specific embodiments and are not considered as limiting the present invention. Throughout the drawings, the same reference signs represent the same components;

[0048] Figure 1 is a flowchart of the method for verifying the consistency of medical data synchronization based on the Elasticsearch database according to an embodiment of the present invention;

[0049] Figure 2 is a block diagram of the system for verifying the consistency of medical data synchronization based on the Elasticsearch database according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings, where the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, and are not used to limit the scope of the present invention.

[0051] To perform a full - consistency check between Elasticsearch and a distributed file database (such as MongoDB), it is necessary to traverse ElasticSearch. However, the following problems often occur during the traversal process: (1) Deep traversal needs to access each document recursively, which will result in high performance overhead; (2) Deep traversal needs to maintain a query context and a stack structure to track the traversal state, which will result in high memory consumption; (3) In a distributed environment, deep traversal needs to perform data interaction between multiple nodes, which will result in high network overhead; (4) Elasticsearch has a certain limit on the depth of recursive queries. If the level of deep traversal exceeds this limit, the query will fail or some data will be missed.

[0052] Based on the above problems, a specific embodiment of the present invention discloses a method for medical data synchronization consistency check based on the Elasticsearch database, as Figure 1 shown, including the following steps:

[0053] Synchronize the medical data in the distributed file database to the Elasticsearch database by aggregating according to the visit times;

[0054] Identify the abnormal data during the aggregation synchronization process to form an abnormal data information table;

[0055] Calculate the slice time size based on the number of visits to obtain the time range of each slice. Extract the medical data corresponding to each slice in the Elasticsearch database and the distributed file database according to the time range of each slice for comparison to find inconsistent data;

[0056] Compare the abnormal data information table with the inconsistent data to obtain the consistency check result.

[0057] It should be noted that the distributed file database can use MongoDB or other distributed file storage databases. During implementation, in order to support the aggregation query requirements, it is necessary to aggregate the data in each table in MongoDB according to the visit times and store it in ElasticSearch. The aggregation process needs to complete the data field mapping conversion and hierarchical conversion by reading the field mapping relationship from the data model.

[0058] One hospitalization of a patient is considered one visit, and one outpatient visit is also considered one visit. There may be multiple medical information such as doctor's orders within the same visit. In the distributed file database, in the doctor's order data table, the doctor's order data of the same patient for the same visit is stored separately. After the data in the distributed file database is aggregated by visit and synchronized to the Elasticsearch database, all the data of the same patient for the same visit is aggregated into one piece of data, that is, the data in the Elasticsearch database is stored in a nested manner. For example, for the doctor's order data table, the doctor's order data of the same patient for the same visit in the Elasticsearch database is one piece of data.

[0059] Compared with the prior art, the medical data synchronization consistency verification method for the Elasticsearch database provided in this embodiment synchronizes the medical data in the distributed file database to the Elasticsearch database by visit, identifies the abnormal data during the synchronization process to form an abnormal data information table. For the comparison of the full amount of data in the two databases, by calculating the slice time size based on the number of visits, corresponding data is extracted from the Elasticsearch database and the distributed file database respectively according to the time range of each slice for comparison, thereby reducing the impact of deep traversal, improving the comparison efficiency, and not requiring the generation of additional comparison files, reducing the need for storage resources. By finding the inconsistent data in the two databases, and then comparing the abnormal data information table with the inconsistent data, the consistency verification result can be obtained quickly.

[0060] During implementation, the aggregated data is detected to identify abnormal data to form abnormal data, and the abnormal data is stored in the abnormal data information table, while the normal data is synchronized to the Elasticsearch database.

[0061] Specifically, identifying the abnormal data during the aggregation and synchronization process to form an abnormal data information table includes:

[0062] For each piece of medical data after aggregation, it is judged whether there are format abnormalities, duplicate abnormalities or encryption abnormalities. If so, the data ID with abnormalities and the corresponding reasons for abnormalities are added to the abnormal data information table. If not, the data is stored in the Elasticsearch database.

[0063] Format exceptions generally refer to data that does not conform to the expected format or structure. For example, if a field is supposed to be in date format but the input data cannot be parsed as a date, this can be recognized as a format exception. Duplication exceptions involve data integrity and consistency checks. For example, if there should be only one admission record for a patient's visit but it is wrongly split into multiple records. Data security is an important part of medical data management. If it is found that the data is not encrypted during storage, the system will mark these data as security risk exceptions and corresponding security reinforcement measures need to be taken, and encryption exceptions are used to detect such exceptions.

[0064] Specifically, the following method is used to determine whether there is a format exception:

[0065] Judge whether the content of a numeric type field is a number. If not, there is a format exception. If it is a number, then judge whether the numeric format of the numeric type field conforms to the preset format. If not, there is a format exception;

[0066] Judge whether the content of a date type field conforms to the preset format. If not, there is a format exception. If it conforms, then judge whether the date is legal. If not, there is a format exception;

[0067] Judge whether the content of a boolean type field is the preset content. If not, there is a format exception.

[0068] During implementation, for a numeric type field, first judge whether the data content of this field is a number. For example, if a numeric type field is written as "not detected", "N / A", "not applicable", or a mixture of numeric data and text, such as "Patient's temperature: 38.5℃", then there is a format exception at this time. If the field content is a number, then judge whether the numeric format conforms to the preset format. For example, for some very large or very small numbers, the preset format is scientific notation, "1e5". Judge whether the number is in the preset format. If not, there is a format exception.

[0069] For the exception detection of date type fields, mainly judge whether the data conforms to the preset format. For example, the correct format is "YYYY-MM-DD HH:mm:SS", but in clinical data, different formats such as "DD / MM / YYYY", "MM-DD-YY", "YYYY / MM / DD" may be seen. Therefore, it does not conform to the preset format and there is a format exception. In addition, for a date that conforms to the preset format, it is also necessary to judge whether it is legal. For example, "2024-02-30", "2024-13-01" are illegal dates, so there is a format exception.

[0070] In addition, anomaly detection needs to be performed on Boolean type fields. The anomaly detection of Boolean type fields mainly determines whether the data is the preset content. For example, the preset content is "yes" or "no". If the data is not one of these two, there is a format anomaly.

[0071] Record the ID + anomaly reason in the anomaly data information table. Among them, the ID is composed of the combination of the current data table ID, patient ID, and data ID. For example, if the synchronized data table is the doctor's order data table, the ID is composed of the doctor's order data table ID, the patient ID to which the abnormal data belongs, and the ID of a specific piece of data with an anomaly.

[0072] In medical data, some business data can only have one record. For example, there can only be one admission record for one visit. For the admission record data table, if, after aggregating by the visit number, there should be only one admission record data for the same patient in the same visit. If there are more than one admission records, there is a duplicate anomaly. During implementation, for structured data, if there is a duplicate anomaly, the data with the largest timestamp is regarded as the normal data, and the other data is regarded as duplicate anomaly data. For unstructured data, by segmenting each piece of data and judging the number of segments, the piece of data with the largest number of segments is taken as the normal data, and the other data is regarded as duplicate anomaly data. Record the ID + anomaly reason in the anomaly data information table.

[0073] Specifically, the following method is used to determine whether there is an encryption anomaly:

[0074] Decrypt the encrypted field and determine whether there is an encryption flag in the decrypted content. If not, there is an encryption anomaly;

[0075] For the de-privatized fields of structured data, determine whether the field content is the de-privatized content. If not, there is an encryption anomaly;

[0076] For unstructured data, check whether there is any private content in it. If so, there is an encryption anomaly.

[0077] During implementation, for fields that require data encryption, domestic encryption algorithms are used for encryption. Since both before and after encryption are strings, it is impossible to directly identify whether the string has been encrypted. Therefore, an encryption flag (for example, adding the prefix SM4:) is added before the original string during encryption. Therefore, after decrypting with the decryption algorithm, determine whether this encryption mark exists. If not, it is judged as an encryption anomaly.

[0078] For fields that require privacy removal, they are divided into structured data and unstructured data according to the data type. For the privacy-removed fields of structured data, it is only necessary to determine whether the content of the field is a replacement string, such as whether it is "******". If not, there is an encryption anomaly. For unstructured data, the normal privacy-removal process is to split the unstructured data into fields, and then use string replacement to remove the privacy of the sensitive fields. Therefore, for unstructured data, check whether there is any private content. If there is, there is an encryption anomaly.

[0079] For example, an admission record is as follows:

[0080] Patient Name: Zhang San

[0081] Gender: Male

[0082] Age: 40 years old

[0083] ID Number: 123456789012345678

[0084] Contact Number: 13800138000

[0085] Medical History: Patient Zhang San visited on July 25, 2024, due to headache and fever for 3 days. Body temperature was 38.5°C, blood pressure was 130 / 85 mmHg, and the preliminary diagnosis was upper respiratory tract infection.

[0086] The result after normal privacy removal should be:

[0087] Patient Name: ******

[0088] Gender: Male

[0089] Age: 40 years old

[0090] ID Number: ******

[0091] Contact Number: ******

[0092] Medical History: Patient ******, visited on July 25, 2024, due to headache and fever for 3 days. Body temperature was 38.5°C, blood pressure was 130 / 85 mmHg, and the preliminary diagnosis was upper respiratory tract infection.

[0093] Therefore, check whether there are private contents such as "Zhang San", "123456789012345678", "13800138000", etc. If there are, there is an encryption anomaly. Record the ID + reason for the anomaly in the abnormal data information table.

[0094] Record the abnormal data information during each synchronization to facilitate judging whether there is a synchronization error during consistency.

[0095] Due to the high occupation and time consumption of existing data comparison and storage, the present invention calculates the slice time size based on the number of visits to obtain the time range of each slice, extracts the medical data corresponding to each slice from the Elasticsearch database and the distributed file database according to the time range of each slice for comparison, and searches for inconsistent data, thereby reducing the impact of deep traversal and improving efficiency.

[0096] During implementation, based on the limitation of the recursive depth of Elasticsearch, a suitable slice time is calculated. If the slice time is too long, it will cause instability in the deep recursion of Elasticsearch; if the slice time is too short, it will cause frequent queries to Elasticsearch, resulting in an increase in comparison time consumption. Therefore, for different numbers of table records, different slice time lengths should be calculated.

[0097] Specifically, the following formula is used to calculate the slice time size:

[0098]

[0099] Among them, W represents the slice time size, k represents the limit number of visits for the slice, V i represents the number of visits in the i-th unit time, T represents the size of the unit time, and n represents the number of unit times from the minimum visit time to the maximum visit time span of the data table.

[0100] By comparing the number of visits V at each time point i with a limit number k and multiplying by the corresponding time length T, the formula can comprehensively consider the impact of the change in the number of visits at different time points on the total weight. This method can better reflect the actual dynamic changes than the simple average speed.

[0101] During implementation, for example, if the unit time is 1 day, it can also be 1 week as the unit time, that is, T is 1 week. The specific unit time can be determined according to the total amount of specific clinic visit data.

[0102] For example, taking days as the time unit, the data in the outpatient visit table is shown in Table 1, and the visit time is from 2024.04.01 to 2024.04.18, then n is 18.

[0103] Table 1 Example Table of Outpatient Visit Data

[0104]

[0105]

[0106] According to the above formula, when k is taken as 10000, W≈20 days. Then, 20 days is used as a slice time size. The data for every 20 days is used as a slice.

[0107] By calculating the difference between the average number of visits within each time window and the target value k, multiplying it by the corresponding time interval, and finally summing to obtain an overall deviation. By adjusting the size of the sliced time window, this overall deviation is made as small as possible, thus achieving the goal of meeting the requirement that the number of visits within each sliced time window does not exceed k while making the sliced time window as large as possible.

[0108] Then, according to the time range of each slice, medical data corresponding to each slice is extracted from the Elasticsearch database and the distributed file database for comparison to find inconsistent data. Specifically including:

[0109] Search for medical data of the target field within the current slice time range in the Elasticsearch database and the distributed file database respectively and sort them in the time order of the target field; obtain the first data and the second data;

[0110] Filter the first data to delete data where the target field exceeds the current slice time range;

[0111] If the earliest times of the target fields of the first data and the second data are different, then take the maximum of the earliest times as the starting time, and take the data where the target field time is less than the starting time as inconsistent data;

[0112] In the first data and the second data, start searching from the data with the target field as the starting time to find the data that does not exist in both the first data and the second data at the same time as inconsistent data.

[0113] During implementation, for example, when the target field is a time field, this time field needs to be related to clinical operations (such as the registration time in the registration record form), and as the time field changes, the number of tables will show an upward trend. Search for medical data of the target field within the current slice time range in the Elasticsearch database and the distributed file database respectively. Take the IDs and the target fields of the data found in the two databases respectively and store them in memory in the time order of the target field, without generating additional comparison files, reducing the need for storage resources.

[0114] Since the data in the Elasticsearch database is stored in a nested manner, querying through the target field will result in data outside the slice time range, and the data outside the current slice time range needs to be removed. Taking the inpatient doctor's order data as an example for illustration:

[0115]

[0116]

[0117] Query the data with the execution time of the doctor's order between May 1, 2024 and May 20, 2024. The parameters for querying Elasticsearch are spliced as follows:

[0118]

[0119]

[0120] Among them, order_time refers to the start time of the inpatient doctor's order. Through the above query parameters, theoretically, the data with the doctor's order start time between May 1, 2024 and May 21, 2024 will be queried. However, due to the nested structure, data beyond this range may be queried. At this time, data filtering needs to be performed through time slicing to obtain accurate results. The specific logic is to judge the business time of the data queried from Elasticsearch. If it is within the current time slice time range, it is retained; otherwise, it is discarded and will not participate in subsequent comparisons.

[0121] Since the distributed file database is not nested in storage, therefore, the data obtained by querying is the data within the current slice time range. The ID and the target field are stored in memory in the time order of the target field, and no additional comparison file needs to be generated, reducing the need for storage resources. The ID and the target field of the medical data corresponding to the current slice obtained by querying the Elasticsearch database are used as the first data, and the ID and the target field of the medical data corresponding to the current slice obtained by querying the distributed file data are used as the second data.

[0122] If the earliest times of the target fields of the first data and the second data are different, it means there are inconsistent data. Take the maximum of the earliest times as the start time, and take the ID with the target field time less than the start time as the inconsistent data ID. Then, start searching from the data with the target field as the start time in the first data and the second data. If there is an ID that does not exist in both the first data and the second data at the same time, then that ID is the inconsistent data ID.

[0123] If the earliest times of the target fields of the first data and the second data are the same, also start searching from the data with the target field as the start time in the first data and the second data. If there is an ID that does not exist in both the first data and the second data at the same time, then that ID is the inconsistent data ID.

[0124] Data inconsistency may be caused by abnormal verification during the aggregation synchronization process, or there may be synchronization errors. Therefore, compare the abnormal data information table with the inconsistent data to obtain the consistency verification result, which specifically includes:

[0125] For each piece of inconsistent data, if the ID of the data exists in the abnormal data information table, the abnormal reason corresponding to the ID in the abnormal data information table is extracted as the reason for the inconsistent verification. If the ID of the data does not exist in the abnormal data information table, the reason for the inconsistent verification is a synchronization error; thus, the consistency verification result is obtained.

[0126] If the abnormal ID exists in the data abnormal data information table, it indicates that it is caused by abnormal verification. If it does not exist in the abnormal data information table, it is caused by a synchronization error. By tracing the reason for the inconsistency, it is convenient for developers to quickly troubleshoot errors and improve data quality.

[0127] A specific embodiment of the present invention discloses a medical data synchronization consistency verification system based on an Elasticsearch database, as Figure 2 shown, including the following modules:

[0128] The aggregation synchronization module is used to synchronize the medical data in the distributed file database to the Elasticsearch database by diagnosis times in an aggregated manner;

[0129] The abnormal identification module is used to identify abnormal data during the aggregation synchronization process to form an abnormal data information table;

[0130] The consistency verification module is used to calculate the slice time size based on the number of medical visits to obtain the time range of each slice, extract the medical data corresponding to each slice in the Elasticsearch database and the distributed file database according to the time range of each slice for comparison, and find inconsistent data; compare the abnormal data information table with the inconsistent data to obtain the consistency verification result.

[0131] Preferably, the following formula is used to calculate the slice time size:

[0132]

[0133] where W represents the slice time size, k represents the number limit of medical visits for slicing, V i represents the number of medical visits in the i-th unit time, T represents the size of the unit time, and n represents the number of unit times from the minimum medical visit time to the maximum medical visit time span of the data table.

[0134] Preferably, extracting the medical data corresponding to each slice in the Elasticsearch database and the distributed file database according to the time range of each slice for comparison and finding inconsistent data includes:

[0135] Search for medical data of the target field within the current slice time range in the Elasticsearch database and the distributed file database respectively, and sort them in the time order of the target field; obtain the first data and the second data;

[0136] Filter the first data, and delete the data where the target field exceeds the current slice time range;

[0137] If the earliest times of the target fields of the first data and the second data are different, then use the maximum of the earliest times as the starting time, and use the data where the target field time is less than the starting time as inconsistent data;

[0138] In the first data and the second data, start searching from the data with the target field as the starting time, and find the data that does not exist in both the first data and the second data at the same time as inconsistent data.

[0139] The above method embodiments and system embodiments are based on the same principle, and their relevant parts can be learned from each other and can achieve the same technical effects. For the specific implementation process, refer to the foregoing embodiments, and details are not described herein again.

[0140] Those skilled in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory or a random access memory, etc.

[0141] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. A medical data synchronization consistency verification method based on Elasticsearch database, characterized in that: The following steps are involved: Aggregate and synchronize medical data in the distributed file database into the Elasticsearch database by diagnosis times; Identify abnormal data in the aggregation synchronization process to form an abnormal data information table; The slice time size is calculated based on the number of visits to obtain the time range of each slice. According to the time range of each slice, the medical data corresponding to each slice is extracted from the Elasticsearch database and the distributed file database for comparison to find inconsistent data. Compare the abnormal data information table with the inconsistent data to obtain a consistency check result; The aggregation is to aggregate all the data of the same patient and the same consultation into one piece of data; The slice time size is calculated using the following formula: Where W represents the slice time size, k represents the number of visits to the slice, V i represents the number of visits per unit time, T represents the size of the unit time, and n represents the number of unit times spanning from the minimum visit time to the maximum visit time in the data table; Identify abnormal data during the aggregation synchronization process to form an abnormal data information table, including: For each piece of aggregated medical data, determine whether there is a format anomaly, duplication anomaly or encryption anomaly. If so, add the data ID with the anomaly and the corresponding anomaly reason to the abnormal data information table. If not, store the data in the Elasticsearch database.

2. According to claim 1, the medical data synchronization consistency verification method based on Elasticsearch database is characterized in that: Use the following methods to determine whether there is a format abnormality: Determine whether the content of the numeric type field is a numeric value. If not, there is a format abnormality. If it is a numeric value, determine whether the numeric format of the numeric type field conforms to the preset format. If not, there is a format abnormality. Determine whether the content of the date type field conforms to the preset format. If not, there is a format abnormality. If it conforms, determine whether the date is legal. If not, there is a format abnormality. Determine whether the content of the Boolean type field is the preset content. If not, there is a format abnormality.

3. According to claim 1, the medical data synchronization consistency verification method based on Elasticsearch database is characterized in that: Use the following methods to determine whether there is an encryption anomaly: Decrypt the encrypted field and determine whether there is an encryption flag in the decrypted content. If not, there is an encryption anomaly. For the de-privacy fields of structured data, determine whether the field content is the de-privacy content. If not, there is an encryption anomaly. For unstructured data, check whether there is any privacy content in it. If so, there is an encryption anomaly.

4. The medical data synchronization consistency verification method based on Elasticsearch database according to claim 1 is characterized in that: Extract the medical data corresponding to each slice from the Elasticsearch database and the distributed file database according to the time range of each slice, and compare them to find inconsistent data, including: Searching for medical data of a target field within a current slice time range in an Elasticsearch database and a distributed file database respectively and sorting the data in chronological order of the target field; obtaining first data and second data; Filter the first data and delete the data that exceeds the target field and the current slice time range; If the earliest time of the target field of the first data and the second data is different, the largest earliest time is used as the start time, and the data whose target field time is smaller than the start time is used as inconsistent data; In the first data and the second data, data that does not exist in both the first data and the second data is searched as inconsistent data starting from the data whose target field is the start time.

5. The medical data synchronization consistency verification method based on Elasticsearch database according to claim 1 is characterized in that: Compare the abnormal data information table with the inconsistent data to obtain consistency verification results, including: For each inconsistent data, if the ID of the data exists in the abnormal data information table, the abnormal reason corresponding to the ID in the abnormal data information table is extracted as the reason for the verification inconsistency. If the ID of the data does not exist in the abnormal data information table, the reason for the verification inconsistency is a synchronization error; thus obtaining the consistency verification result.

6. A medical data synchronization consistency verification system based on Elasticsearch database, characterized in that: Includes the following modules: Aggregation synchronization module, used to aggregate and synchronize medical data in the distributed file database into the Elasticsearch database by diagnosis number; An abnormality identification module is used to identify abnormal data in the aggregation synchronization process to form an abnormal data information table; The consistency check module is used to calculate the slice time size based on the number of visits to obtain the time range of each slice, extract the medical data corresponding to each slice from the Elasticsearch database and the distributed file database according to the time range of each slice, and compare them to find inconsistent data; compare the abnormal data information table with the inconsistent data to obtain the consistency check result; The aggregation is to aggregate all the data of the same patient and the same consultation into one piece of data; The slice time size is calculated using the following formula: Where W represents the slice time size, k represents the number of visits to the slice, V i represents the number of visits per unit time, T represents the size of the unit time, and n represents the number of unit times spanning from the minimum visit time to the maximum visit time in the data table; Identify abnormal data during the aggregation synchronization process to form an abnormal data information table, including: For each piece of aggregated medical data, determine whether there is a format anomaly, duplication anomaly or encryption anomaly. If so, add the data ID with the anomaly and the corresponding anomaly reason to the abnormal data information table. If not, store the data in the Elasticsearch database.

7. The medical data synchronization consistency verification system based on Elasticsearch database according to claim 6 is characterized in that: Extract the medical data corresponding to each slice from the Elasticsearch database and the distributed file database according to the time range of each slice, and compare them to find inconsistent data, including: Searching for medical data of a target field within a current slice time range in an Elasticsearch database and a distributed file database respectively and sorting the data in chronological order of the target field; obtaining first data and second data; Filter the first data and delete the data that exceeds the target field and the current slice time range; If the earliest time of the target field of the first data and the second data is different, the largest earliest time is used as the start time, and the data whose target field time is smaller than the start time is used as inconsistent data; In the first data and the second data, data that does not exist in both the first data and the second data is searched as inconsistent data starting from the data whose target field is the start time.

Citation Information

Patent Citations

  • Data query method and device for payment and settlement system

    CN111753015A

  • Synchronous data exception handling method, device and system

    CN117762672A