A data governance method, apparatus and device

By acquiring and restoring historical medical data, and training a data governance model based on patient identification and treatment sequence, the problem of data silos across different platforms was solved, achieving efficient data governance and the acquisition of high-quality medical data.

CN117271492BActive Publication Date: 2025-12-05LIANREN HEALTHCARE BIG DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311256617.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-26
Publication Date
2025-12-05
Estimated Expiration
2043-09-26

AI Technical Summary

Technical Problem

The significant differences in the medical data management systems of different third-party platforms result in low data governance efficiency and isolated medical governance data, failing to meet the demand for high-quality data.

Method used

By acquiring historical medical data from multiple third-party platforms, data repair and classification are performed. Training samples are divided based on patient identification information and the order of treatment to train a data governance model, thereby achieving data repair and table association processing and obtaining high-quality medical data.

Benefits of technology

It improved data governance efficiency, obtained high-quality medical data, and met the needs for medical data circulation and empowerment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117271492B_ABST
    Figure CN117271492B_ABST
Patent Text Reader

Abstract

The application discloses a kind of data governance method, device and equipment, the method includes: obtaining the historical diagnosis and treatment data generated by multiple third-party platforms;The historical diagnosis and treatment data meeting the preset condition is copied to the preset copy library, and the historical diagnosis and treatment data in the preset copy library is handled, and the historical repair diagnosis and treatment data is obtained;Based on patient identification information and the preset diagnosis and treatment sequence, the historical repair diagnosis and treatment data is divided into multiple training samples, and the target training sample set is obtained;The preset data governance model is trained based on the target training sample set, and the target data governance model is obtained;Based on the target data governance model, the data repair and data table association processing are carried out on the diagnosis and treatment data to be processed, and the target high-quality diagnosis and treatment data corresponding to the diagnosis and treatment data to be processed is obtained.The application can obtain high-quality diagnosis and treatment data, and improve the data governance efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of data processing, in particular to a data governance method, device and equipment. BACKGROUND

[0002] With the country increasingly emphasizing the position of data elements in the national development and security strategy, data circulation is the "hub" of data element generation to data value release. The urgent needs of data circulation and the characteristics of medical policies and medical industry provide a broad stage for data circulation related technologies. Limited by the professionalism and seriousness of the medical industry, higher requirements are put forward for diagnosis and treatment data standards and business data quality in order to achieve the purpose of data empowerment.

[0003] At present, since the diagnosis and treatment data management systems adopted by various third-party platforms are different, the diagnosis and treatment data generated by different third-party platforms are quite different, and the data execution standards are different. For different diagnosis and treatment data management systems, a special data governance model needs to be designed to complete the data governance task, so as to obtain high-quality diagnosis and treatment data.

[0004] However, this data governance method makes the medical governance data isolated from each other, and for different diagnosis and treatment data management systems, a special data governance model needs to be designed, which also has the technical problem of low data governance efficiency. SUMMARY

[0005] The present application provides a data governance method, device and equipment, which can obtain high-quality diagnosis and treatment data and improve the data governance efficiency.

[0006] According to a first aspect of the present application, a data governance method is provided, which comprises:

[0007] Obtaining historical diagnosis and treatment data generated by a plurality of third-party platforms; wherein the historical diagnosis and treatment data is provided by a plurality of diagnosis and treatment data management systems, each of the diagnosis and treatment data management systems provides a plurality of diagnosis and treatment process data tables, and the historical diagnosis and treatment data contains patient identification information;

[0008] Copying the historical diagnosis and treatment data meeting the preset condition to a preset copy library, and performing data repair processing on the historical diagnosis and treatment data in the preset copy library to obtain historical repair diagnosis and treatment data;

[0009] Based on the patient identification information and a preset diagnosis and treatment sequence, the historical repair diagnosis and treatment data is divided into a plurality of training samples to obtain a target training sample set; wherein the plurality of diagnosis and treatment process data tables in each training sample have an association relationship corresponding to the preset diagnosis and treatment sequence;

[0010] The target data governance model is obtained by training the pre-set data governance model based on the target training sample set.

[0011] Based on the target data governance model, data repair and data table association processing are performed on the treatment data to be processed to obtain the target high-quality treatment data corresponding to the treatment data to be processed; wherein, the treatment data to be processed is the current treatment data generated by the target third-party platform.

[0012] According to a second aspect of the present invention, a data governance apparatus is provided, the apparatus comprising:

[0013] The historical data acquisition module is used to acquire historical medical data generated by multiple third-party platforms; wherein, the historical medical data is provided by multiple medical data management systems, each of which provides multiple medical process data tables, and the historical medical data includes patient identification information;

[0014] The historical data copying module is used to copy historical medical data that meets preset conditions to a preset copying library, and to perform data repair processing on the historical medical data in the preset copying library to obtain historical repaired medical data.

[0015] The training sample set determination module is used to divide the historical repair treatment data into multiple training samples based on the patient identification information and the preset treatment sequence to obtain a target training sample set; wherein, multiple treatment process data tables in each training sample have an association relationship with the preset treatment sequence.

[0016] The governance model training module is used to train a preset data governance model based on the target training sample set to obtain the target data governance model.

[0017] The diagnosis and treatment data governance module is used to perform data repair and data table association processing on the diagnosis and treatment data to be processed based on the target data governance model, so as to obtain the target high-quality diagnosis and treatment data corresponding to the diagnosis and treatment data to be processed; wherein, the diagnosis and treatment data to be processed is the current diagnosis and treatment data generated by the target third-party platform.

[0018] According to a third aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0019] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data governance method according to any embodiment of the present invention.

[0020] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data governance method described in any embodiment of the present invention.

[0021] The technical solution of this invention involves acquiring historical medical data generated by multiple third-party platforms. This historical medical data is provided by various medical data management systems, each providing multiple medical process data tables. The historical medical data includes patient identification information. Then, historical medical data meeting preset conditions is copied to a preset copy library. Data repair processing is performed on the historical medical data in the preset copy library to obtain repaired historical medical data. Further, based on patient identification information and a preset treatment sequence, the repaired historical medical data is divided into multiple training samples to obtain a target training sample set. Each training sample contains multiple medical process data tables that have an association relationship corresponding to the preset treatment sequence. Further, a preset data governance model is trained based on the target training sample set to obtain a target data governance model. Then, based on the target data governance model, data repair and data table association processing are performed on the medical data to be processed to obtain target high-quality medical data corresponding to the medical data to be processed. The medical data to be processed is the current medical data generated by the target third-party platform. This application can obtain high-quality medical data and improves data governance efficiency.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of a data governance method provided in Embodiment 1 of the present invention;

[0025] Figure 2 This is a flowchart of a data governance method provided according to Embodiment 2 of the present invention;

[0026] Figure 3 This is a schematic diagram of the structure of a data governance device according to Embodiment 3 of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the data governance method of this invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Example 1

[0031] Figure 1 This is a flowchart of a data governance method provided in Embodiment 1 of the present invention. This embodiment is applicable to data governance of diagnostic and treatment data provided by any diagnostic and treatment data management system to obtain high-quality diagnostic and treatment data. This method can be executed by a data governance device, which can be implemented in hardware and / or software, and can be configured in a terminal and / or server. Figure 1 As shown, the method includes:

[0032] S110. Obtain historical medical data generated by multiple third-party platforms.

[0033] The third-party platforms include medical institutions such as hospitals, clinics, and health centers. Historical medical data encompasses all data generated during past treatments on these third-party platforms, including but not limited to treatment process data tables, images, and data logs. This historical medical data is provided by multiple medical data management systems, each offering multiple treatment process data tables. The historical medical data includes patient identification information.

[0034] In this embodiment, the medical data management system is a server that records and manages patient medical data. For a third-party platform, different departments may use different medical data management systems, and these different systems are developed by different manufacturers. For example, a hospital may include outpatient departments, laboratories, examination departments, radiology departments, pathology departments, vital signs departments, inpatient departments, nursing departments, and rehabilitation departments. The outpatient department may use medical data management system type A from manufacturer A, while the laboratory department may use medical data management system type B from manufacturer B. Therefore, the historical medical data obtained is provided by multiple medical data management systems.

[0035] In this embodiment, the patient's medical treatment process in a single department involves multiple steps, each corresponding to a different treatment process data table. Therefore, each treatment data management system provides multiple treatment process data tables. Patient identification information is a unique identifier for each patient, such as the patient's appointment sequence number, registration code, and unique identity identifier.

[0036] Specifically, to obtain a high-performance data governance model, it is necessary to train the pre-set data governance model based on a large amount of historical medical data. This allows the data governance model to understand the data flow relationships between the medical process data tables provided by various medical data management systems throughout a complete medical process, as well as the data characteristics of each medical process data table itself. Therefore, it is necessary to obtain historical medical data generated by multiple third-party platforms.

[0037] Specifically, after acquiring historical medical data from multiple third-party platforms, the historical medical data undergoes data preprocessing. The preprocessing includes at least one of the following steps: classifying and grouping the historical medical data based on the differences between various medical data management systems; performing intelligent restricted standardization and business-based standardization preprocessing on the historical medical data; and performing intelligent deduplication and normalization on value ranges, coded equivalent sets, and terminological attribute features across medical data management systems and versions.

[0038] S120. Copy the historical medical data that meets the preset conditions to the preset copy library, and perform data repair processing on the historical medical data in the preset copy library to obtain the historical repaired medical data.

[0039] Among them, the preset conditions are the pre-set prerequisites. The preset copy library is the pre-configured data storage structure, and the historical repaired treatment data is the treatment data obtained after repairing the historical treatment data.

[0040] Specifically, the preset conditions are: to evaluate the data quality of historical medical data from multiple evaluation dimensions, obtain the data quality score corresponding to the historical medical data, and the historical medical data with a data quality score greater than the preset threshold.

[0041] The evaluation dimensions include completeness, consistency, timeliness, effectiveness, uniqueness, and accuracy. Completeness assesses whether blank fields exist in historical medical data tables that should have them. Consistency assesses whether the content and titles of historical medical data tables do not match. Timeliness assesses whether historical medical data was generated at the expected time but was not. Effectiveness assesses whether useless data exists in historical medical data. Uniqueness assesses whether the same patient's medical data is repeated. Accuracy assesses whether erroneous data exists in historical medical data.

[0042] The data quality score is a numerical representation of the quality of historical medical data. For example, the data quality score is on a 100-point scale, with higher scores indicating higher quality and lower scores indicating lower quality. The preset threshold is a pre-defined threshold for the data quality score.

[0043] Optionally, the historical medical data can be evaluated for data quality from multiple evaluation dimensions to obtain the corresponding data quality score. Specifically, this includes: determining the data quality score to be applied for the historical medical data under each evaluation dimension; and determining the corresponding data quality score for the historical medical data based on the weighted sum of the data quality scores to be applied.

[0044] In this embodiment, the data quality score for each treatment process data table can be determined under the dimensions of completeness evaluation, consistency evaluation, timeliness evaluation, effectiveness evaluation, uniqueness evaluation, and accuracy evaluation. Based on this, six data quality scores are obtained for each treatment process data table. Furthermore, the data quality score corresponding to each treatment process data table is determined by the weighted sum of these six scores.

[0045] Furthermore, based on the data quality scores obtained for the treatment process data tables, treatment process data tables with data quality scores greater than a preset threshold are retained for subsequent processing. For example, if the preset threshold is 60 points, then treatment process data tables with data quality scores greater than 60 points are retained for subsequent processing. The purpose of this is that historical treatment data with very poor data quality is not only difficult to repair, but also degrades the performance of the data governance model; therefore, historical treatment data with very poor data quality can be filtered out.

[0046] Specifically, based on historical medical data that meets preset conditions, this data is copied to a preset copy library. This facilitates data repair processing of the historical medical data. Data repair processing includes at least one of the following steps: intelligently filling in missing data, detecting and correcting erroneous data, and standardizing non-standard data content. After performing these repair steps on the historical medical data, the repaired historical medical data is obtained.

[0047] S130. Based on patient identification information and preset treatment sequence, the historical repair treatment data is divided into multiple training samples to obtain the target training sample set.

[0048] The patient identification information consists of existing data from various treatment process tables within the historical treatment data, allowing direct access to the patient identification information for each patient. The training samples are ordered historical treatment sample sets obtained after summarizing and processing the previously disorganized historical treatment data. The target training sample set is a collection of multiple training samples.

[0049] The preset treatment sequence is a pre-defined arrangement of different departments. For example, a complete medical procedure may involve appointment registration, consultation, laboratory tests, diagnosis, treatment, and rehabilitation. Different departments correspond to different stages of the procedure; therefore, the departments at each stage can be arranged sequentially. An example preset treatment sequence could be: outpatient department, laboratory, examination department, radiology department, pathology department, physical examination department, inpatient department, nursing department, and rehabilitation department.

[0050] Specifically, determining the target training sample set includes the following steps:

[0051] S1301. Based on patient identification information, the historical repair and treatment data in the preset copy library are divided into multiple historical treatment sample groups.

[0052] In this embodiment, for each treatment process data table in the historical repair and treatment data in the preset copy library, the data related to the patient identification information in the treatment process data table is summarized according to the patient identification information in the treatment process data table, so as to divide the historical repair and treatment data into each historical treatment sample group corresponding to each patient identification information, thereby obtaining multiple historical treatment sample groups.

[0053] S1301. For each historical treatment sample group, perform sequential association processing on the data tables of each treatment process in the historical treatment sample group based on the preset treatment sequence to obtain the target historical treatment sample group corresponding to each historical treatment sample group.

[0054] In this embodiment, each historical treatment sample group includes treatment data corresponding to multiple treatment departments. Therefore, the treatment data in each historical treatment sample group are associated according to a preset order of treatment, so that the treatment data in each historical treatment sample group is an ordered data stream that reflects the actual treatment process, thereby obtaining the target historical treatment sample group corresponding to each historical treatment sample group.

[0055] S1301. Based on the historical diagnosis and treatment sample groups of each target, a target training sample set is formed.

[0056] In this embodiment, based on obtaining multiple target historical diagnosis and treatment sample groups, the entirety of each target historical diagnosis and treatment sample group is used as the target training sample set.

[0057] S140. Train the preset data governance model based on the target training sample set to obtain the target data governance model.

[0058] The preset data governance model is a pre-determined generative AI medical vertical model, and the model parameters in the preset data governance model are initial parameters.

[0059] In this embodiment, the target training sample set is input into a preset data governance model so that the target training samples can train the preset data governance model. This allows the trained target data governance model to recognize the data flow relationships between the data tables provided by various medical data management systems throughout a complete medical process, as well as the data characteristics of each data table itself.

[0060] S150. Based on the target data governance model, perform data repair and data table association processing on the treatment data to be processed to obtain the target high-quality treatment data corresponding to the treatment data to be processed.

[0061] The medical data to be processed refers to the current medical data generated by the target third-party platform. This current medical data can be understood as medical data that requires data governance.

[0062] In this embodiment, after obtaining the target data governance model, the target data governance model can be used to perform data repair and data table association processing on the medical data to be processed provided by any third-party platform. After the medical data to be processed is processed by the target data governance model, the target data governance model outputs high-quality target medical data corresponding to the medical data to be processed.

[0063] Specifically, the target high-quality medical data corresponding to the medical data to be processed is determined, including the following: based on the target data governance model, the medical data to be processed is subjected to error format correction and repair processing, missing data supplementation and repair processing, value range standardization and repair processing, and data table association processing to obtain the target high-quality medical data corresponding to the medical data to be processed.

[0064] In this embodiment, the target data governance model has learned the characteristics of clinical data provided by various mainstream clinical data management systems and the data table relationships between different mainstream clinical data management systems through a large amount of historical clinical data. Specifically, the mainstream clinical data management systems are those currently used by various third-party platforms. Therefore, after inputting the clinical data to be processed into the target data governance model, the model can identify which data has format errors, which has missing data, and which has non-standard value ranges. Based on this, the target data governance model performs data repair processing on these problematic data from at least one aspect: error format correction, missing data supplementation, and value range standardization. Furthermore, after repairing each clinical process data table, the model performs data table association processing. For example, based on the learned data flow relationships, the target data governance model concatenates the various clinical process data tables to form a closed-loop data flow that reflects the complete patient treatment process. Thus, the target high-quality clinical data corresponding to the clinical data to be processed is obtained.

[0065] The technical solution of this invention involves acquiring historical medical data generated by multiple third-party platforms. This historical medical data is provided by various medical data management systems, each providing multiple medical process data tables. The historical medical data includes patient identification information. Then, historical medical data meeting preset conditions is copied to a preset copy library. Data repair processing is performed on the historical medical data in the preset copy library to obtain repaired historical medical data. Further, based on patient identification information and a preset treatment sequence, the repaired historical medical data is divided into multiple training samples to obtain a target training sample set. Each training sample contains multiple medical process data tables that have an association relationship corresponding to the preset treatment sequence. Further still, a preset data governance model is trained based on the target training sample set to obtain a target data governance model. This target data governance model is then used to perform data repair and data table association processing on the medical data to be processed, resulting in target high-quality medical data corresponding to the medical data to be processed. The medical data to be processed is the current medical data generated by the target third-party platform. This application can obtain high-quality medical data and improves data governance efficiency.

[0066] Example 2

[0067] Figure 2 This is a flowchart of a data governance method provided in Embodiment 2 of the present invention. Based on the foregoing embodiments, it describes in detail how the trained target data governance model is applied to the target third-party platform. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0068] like Figure 2 As shown, the method includes:

[0069] S210. Obtain historical medical data generated by multiple third-party platforms.

[0070] S220. Copy the historical medical data that meets the preset conditions to the preset copy library, and perform data repair processing on the historical medical data in the preset copy library to obtain the historical repaired medical data.

[0071] S230. Based on patient identification information and preset treatment sequence, the historical repair treatment data is divided into multiple training samples to obtain the target training sample set.

[0072] S240. Train the preset data governance model based on the target training sample set to obtain the target data governance model.

[0073] S250. Apply the target data governance model to the target third-party platform, and extract the treatment data to be processed from all the treatment data of the target third-party platform based on the preset data extraction rules.

[0074] The target third-party platform refers to any medical institution such as a hospital, clinic, or health center. All medical data on the target third-party platform is provided by various target medical data management systems. The medical data to be processed is the data that will be processed based on the target data governance model.

[0075] Among them, the preset data extraction rules are the specific measures to be taken when extracting the medical data to be processed from all medical data.

[0076] Optionally, the data extraction rules include at least one of the following aspects: configuring the name of the data table to be extracted for the treatment process corresponding to the target treatment data management system; the time interval for extracting data to be processed from the target third-party platform; and the data extraction retry strategy to be adopted when the extraction of data to be processed from the target third-party platform fails.

[0077] Specifically, the trained target data governance model can be applied to any target third-party platform to perform data governance processing on the diagnostic and treatment data generated by that platform. Since some useless data exists within all the diagnostic and treatment data generated by the target third-party platform, in practical applications, it is necessary to extract the diagnostic and treatment data that requires data governance from the entire dataset. When extracting treatment data to be processed, data extraction is performed based on preset data extraction rules. The first aspect of the data extraction rules is to configure the name of the treatment process data table to be extracted, corresponding to the target treatment data management system. This determines the name of the treatment data table to be extracted from the target treatment data management system, and the specific treatment process data table to be extracted can be determined based on the data table name. The second aspect of the data extraction rules is the time interval for extracting treatment data from the target third-party platform. This can be understood as how often data is extracted from the target third-party platform. The third aspect of the data extraction rules is the data extraction retry strategy adopted when the extraction of treatment data from the target third-party platform fails. This can be understood as how often data is extracted from the target third-party platform if the extraction of treatment data fails. After a certain number of retries, the retry task ends and an error message is displayed.

[0078] S260. Based on the target medical data management system corresponding to the medical data to be processed, determine the synchronous replication scheme corresponding to the medical data to be processed.

[0079] The synchronous replication scheme refers to the specific measures taken during the process of replicating the treatment data to be processed to a pre-set governance replication library.

[0080] In this embodiment, a synchronization replication scheme corresponding to the target medical data management system can be pre-configured in the target data governance model. This synchronization replication scheme includes, but is not limited to, specifying the scripts to be replicated, the data source configuration of the replication repository, the data storage scheme, and the replication tasks. Based on the determination of the target medical data management system, a synchronization replication scheme for the medical data to be processed corresponding to that system can be determined.

[0081] S270. Based on the synchronous replication scheme, the diagnosis and treatment data to be processed is replicated to the preset treatment replication library, and the data replication progress is displayed in real time.

[0082] The preset governance replication library is a pre-configured data storage unit. Data replication progress is used to characterize the degree of completion of data replication. For example, data replication progress is the ratio of the amount of pending diagnosis and treatment data that has been replicated to the total amount of pending diagnosis and treatment data that needs to be replicated, which can be represented as a percentage.

[0083] In this embodiment, the medical data to be processed is copied to a preset management copy library according to the synchronous copy scheme. During the synchronous copying of the medical data to be processed, a data processing log is recorded and the data copying progress is displayed synchronously on the target display device.

[0084] S280. Based on the target data governance model, perform data repair and data table association processing on the treatment data to be processed to obtain the target high-quality treatment data corresponding to the treatment data to be processed.

[0085] S290. Evaluate the data quality of the target high-quality diagnosis and treatment data from multiple evaluation dimensions to obtain the target data quality score corresponding to the target high-quality diagnosis and treatment data.

[0086] Among them, the target data quality score is used to characterize the quality of the target high-quality diagnostic and treatment data.

[0087] The evaluation dimensions include completeness, consistency, timeliness, effectiveness, uniqueness, and accuracy.

[0088] In this embodiment, after obtaining the target high-quality medical data, the data quality is evaluated from the dimensions of completeness, consistency, timeliness, effectiveness, uniqueness, and accuracy, respectively, to obtain the target data quality score corresponding to the target high-quality medical data under each dimension.

[0089] The technical solution of this invention, after obtaining the target data governance model, applies the target data governance model to the target third-party platform. Based on preset data extraction rules, it extracts the treatment data to be processed from all the treatment data of the target third-party platform. Then, based on the target treatment data management system corresponding to the treatment data to be processed, it determines the corresponding synchronous replication scheme for the treatment data to be processed. Based on the synchronous replication scheme, the treatment data to be processed is replicated to the preset governance replication library, and the data replication progress is displayed in real time. By using data extraction rules, only the treatment data required for data governance is extracted, rather than performing data governance processing on all data of the target third-party platform, thereby improving data governance efficiency. Furthermore, replicating the treatment data to be processed to the preset governance replication library according to the synchronous replication scheme ensures the accuracy of data synchronous replication. In this embodiment, after obtaining the target high-quality treatment data, the data quality is evaluated from multiple evaluation dimensions to obtain the target data quality score corresponding to the target high-quality treatment data, thus achieving a quantitative evaluation of the target high-quality treatment data.

[0090] Example 3

[0091] Figure 3 This is a schematic diagram of the structure of a data governance device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: a historical data acquisition module 310, a historical data copying module 320, a training sample set determination module 330, a governance model training module 340, and a diagnosis and treatment data governance module 350.

[0092] The historical data acquisition module 310 is used to acquire historical medical data generated by multiple third-party platforms. The historical medical data is provided by multiple medical data management systems, each of which provides multiple medical process data tables. The historical medical data includes patient identification information.

[0093] The historical data copying module 320 is used to copy historical medical data that meets preset conditions to a preset copying library, and to perform data repair processing on the historical medical data in the preset copying library to obtain historical repaired medical data.

[0094] The training sample set determination module 330 is used to divide the historical repair treatment data into multiple training samples based on the patient identification information and the preset treatment sequence to obtain a target training sample set; wherein, multiple treatment process data tables in each training sample have an association relationship with the preset treatment sequence.

[0095] The governance model training module 340 is used to train the preset data governance model based on the target training sample set to obtain the target data governance model.

[0096] The diagnosis and treatment data governance module 350 is used to perform data repair and data table association processing on the diagnosis and treatment data to be processed based on the target data governance model, so as to obtain the target high-quality diagnosis and treatment data corresponding to the diagnosis and treatment data to be processed; wherein, the diagnosis and treatment data to be processed is the current diagnosis and treatment data generated by the target third-party platform.

[0097] The technical solution of this invention involves acquiring historical medical data generated by multiple third-party platforms. This historical medical data is provided by various medical data management systems, each providing multiple medical process data tables. The historical medical data includes patient identification information. Then, historical medical data meeting preset conditions is copied to a preset copy library. Data repair processing is performed on the historical medical data in the preset copy library to obtain repaired historical medical data. Further, based on patient identification information and a preset treatment sequence, the repaired historical medical data is divided into multiple training samples to obtain a target training sample set. Each training sample contains multiple medical process data tables that have an association relationship corresponding to the preset treatment sequence. Further still, a preset data governance model is trained based on the target training sample set to obtain a target data governance model. This target data governance model is then used to perform data repair and data table association processing on the medical data to be processed, resulting in target high-quality medical data corresponding to the medical data to be processed. The medical data to be processed is the current medical data generated by the target third-party platform. This application can obtain high-quality medical data and improves data governance efficiency.

[0098] Optionally, the historical data replication module 320 includes: a data determination unit that meets preset conditions, used to evaluate the data quality of the historical medical data from multiple evaluation dimensions to obtain a data quality score corresponding to the historical medical data, wherein the historical medical data with a data quality score greater than a preset threshold is historical medical data that meets the preset conditions.

[0099] Optionally, the multiple evaluation dimensions include completeness evaluation dimension, consistency evaluation dimension, timeliness evaluation dimension, effectiveness evaluation dimension, uniqueness evaluation dimension, and accuracy evaluation dimension; the data determination unit that meets preset conditions is specifically used to determine the data quality score to be applied corresponding to the historical medical data under each of the evaluation dimensions; and to determine the data quality score corresponding to the historical medical data based on the weighted sum of the data quality scores to be applied.

[0100] Optionally, the training sample set determination module 330 includes:

[0101] The sample group division unit is used to divide the historical repair and treatment data in the preset copy library into multiple historical treatment sample groups based on the patient identification information.

[0102] The data table association unit is used to perform sequential association processing on each treatment process data table in each historical treatment sample group based on a preset treatment sequence, so as to obtain the target historical treatment sample group corresponding to each historical treatment sample group.

[0103] The sample set determination unit is used to construct the target training sample set based on the historical diagnosis and treatment sample groups of each target.

[0104] Optionally, the data governance device further includes a data processing module, which includes:

[0105] The data extraction unit is used to apply the target data governance model to the target third-party platform and extract the data to be processed from all the diagnosis and treatment data of the target third-party platform based on preset data extraction rules; wherein, the all diagnosis and treatment data is provided by multiple target diagnosis and treatment data management systems.

[0106] The data replication scheme determination unit is used to determine the synchronous replication scheme corresponding to the medical data to be processed based on the target medical data management system corresponding to the medical data to be processed;

[0107] The data synchronization and replication unit is used to copy the treatment data to be processed to a preset treatment replication library based on the synchronization and replication scheme, and to display the data replication progress in real time.

[0108] Optionally, the diagnosis and treatment data governance module 350 is specifically used to: perform error format correction and repair processing, missing data supplementation and repair processing, value range standardization and repair processing, and data table association processing on the diagnosis and treatment data to be processed based on the target data governance model, so as to obtain the target high-quality diagnosis and treatment data corresponding to the diagnosis and treatment data to be processed.

[0109] Optionally, the data governance device also includes a target data quality evaluation module, which is specifically used to evaluate the target high-quality medical data from multiple evaluation dimensions to obtain the target data quality score corresponding to the target high-quality medical data.

[0110] The data governance apparatus provided in the embodiments of the present invention can execute the data governance method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0111] Example 4

[0112] Figure 4A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0113] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0114] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0115] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data governance methods.

[0116] In some embodiments, the data governance method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data governance method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data governance method by any other suitable means (e.g., by means of firmware).

[0117] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0118] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0119] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0120] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0121] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0122] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0123] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0124] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data governance method, characterized in that, include: The system acquires historical medical data generated by multiple third-party platforms; wherein the historical medical data is provided by multiple medical data management systems, each of which provides multiple medical process data tables, and the historical medical data includes patient identification information. Historical medical data that meets preset conditions is copied to a preset copy library, and the historical medical data in the preset copy library is subjected to data repair processing to obtain historical repaired medical data. Based on the patient identification information and the preset treatment sequence, the historical repair treatment data is divided into multiple training samples to obtain a target training sample set; wherein, multiple treatment process data tables in each training sample have an association relationship with the preset treatment sequence. The target data governance model is obtained by training the pre-set data governance model based on the target training sample set. Based on the target data governance model, data repair and data table association processing are performed on the treatment data to be processed to obtain the target high-quality treatment data corresponding to the treatment data to be processed; wherein, the treatment data to be processed is the current treatment data generated by the target third-party platform.

2. The method according to claim 1, characterized in that, The preset conditions are: The historical medical data is evaluated for data quality from multiple evaluation dimensions to obtain a data quality score corresponding to the historical medical data. The historical medical data with a data quality score greater than a preset threshold are included.

3. The method according to claim 2, characterized in that, The multiple evaluation dimensions include completeness evaluation dimension, consistency evaluation dimension, timeliness evaluation dimension, effectiveness evaluation dimension, uniqueness evaluation dimension, and accuracy evaluation dimension; The process of evaluating the historical medical data from multiple evaluation dimensions to obtain a data quality score corresponding to the historical medical data includes: Determine the data quality score of the historical medical data to be applied under each of the aforementioned evaluation dimensions; The data quality score corresponding to the historical medical data is determined based on the weighted sum of the quality scores of each of the data to be applied.

4. The method according to claim 1, characterized in that, Based on the patient identification information and the preset treatment sequence, the historical repair treatment data is divided into multiple training samples to obtain a target training sample set, including: Based on the patient identification information, the historical repair and treatment data in the preset copy library are divided into multiple historical treatment sample groups; For each historical treatment sample group, the treatment process data tables in the historical treatment sample group are sequentially associated based on a preset treatment sequence to obtain the target historical treatment sample group corresponding to each historical treatment sample group. Based on the historical diagnosis and treatment sample groups of each target, a target training sample set is constructed.

5. The method according to claim 1, characterized in that, Before performing data repair and data table association processing on the treatment data to be processed based on the target data governance model to obtain the target high-quality treatment data corresponding to the treatment data to be processed, the process also includes: The target data governance model is applied to the target third-party platform, and based on preset data extraction rules, the medical data to be processed is extracted from all the medical data of the target third-party platform; wherein, all the medical data is provided by multiple target medical data management systems; Based on the target medical data management system corresponding to the medical data to be processed, a synchronous replication scheme corresponding to the medical data to be processed is determined; Based on the synchronous replication scheme, the diagnosis and treatment data to be processed is copied to the preset treatment replication library, and the data replication progress is displayed in real time.

6. The method according to claim 5, characterized in that, The data extraction rules include at least one of the following aspects: Configure the name of the data table to be extracted for the treatment process corresponding to the target treatment data management system; The time interval for extracting data to be processed from the target third-party platform; When the extraction of data to be processed from the target third-party platform fails, a data extraction retry strategy is adopted.

7. The method according to claim 1, characterized in that, The process of performing data repair and data table association on the treatment data to be processed based on the target data governance model to obtain the target high-quality treatment data corresponding to the treatment data to be processed includes: Based on the target data governance model, the medical data to be processed is subjected to error format correction and repair, missing data supplementation and repair, value range standardization and repair, and data table association processing to obtain the target high-quality medical data corresponding to the medical data to be processed.

8. The method according to claim 7, characterized in that, After performing data repair and data table association processing on the treatment data to be processed based on the target data governance model to obtain the target high-quality treatment data corresponding to the treatment data to be processed, the process further includes: The target high-quality medical data is evaluated from multiple evaluation dimensions to obtain the target data quality score.

9. A data governance device, characterized in that, include: The historical data acquisition module is used to acquire historical medical data generated by multiple third-party platforms; wherein, the historical medical data is provided by multiple medical data management systems, each of which provides multiple medical process data tables, and the historical medical data includes patient identification information; The historical data copying module is used to copy historical medical data that meets preset conditions to a preset copying library, and to perform data repair processing on the historical medical data in the preset copying library to obtain historical repaired medical data. The training sample set determination module is used to divide the historical repair treatment data into multiple training samples based on the patient identification information and the preset treatment sequence to obtain a target training sample set; wherein, multiple treatment process data tables in each training sample have an association relationship with the preset treatment sequence. The governance model training module is used to train a preset data governance model based on the target training sample set to obtain the target data governance model. The diagnosis and treatment data governance module is used to perform data repair and data table association processing on the diagnosis and treatment data to be processed based on the target data governance model, so as to obtain the target high-quality diagnosis and treatment data corresponding to the diagnosis and treatment data to be processed; wherein, the diagnosis and treatment data to be processed is the current diagnosis and treatment data generated by the target third-party platform.

10. An electronic device, characterized in that, Electronic devices include: One or more processors; Storage device for storing one or more programs. When one or more programs are executed by one or more processors, the one or more processors implement the data governance method as claimed in any one of claims 1-8.

Citation Information

Patent Citations

  • Variable exception repair method and device, medium and computer program product

    CN113641525A

  • Text error correction model training method and device, electronic equipment and storage medium

    CN115759051A