Medical big data ES wide table generation method and device based on efficient dynamic data configuration
By adopting an efficient dynamic data configuration method in medical data systems, Elasticsearch wide tables are generated, and data silos and query efficiency in existing systems are solved, cross-institutional data sharing and efficient analysis are realized, maintenance costs are reduced, and multi-dimensional analysis is supported.
Patent Information
- Application Number
- CN202510089255.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing medical data systems have problems such as data silos, low query efficiency, and difficulty in realizing cross-institutional data sharing and analysis. The system design lacks flexibility and scalability, resulting in high maintenance costs and difficulty in supporting multi-dimensional analysis and intelligence.
Through the method of efficient dynamic data configuration, an Elasticsearch (ES) distributed cluster is built, data sources, tables and fields are dynamically configured, data is collected and processed, and ES wide tables are generated that meet the needs of medical big data to realize standardized data management and efficient query.
It significantly improves the performance of cross-institutional data sharing and high concurrent queries, reduces data processing complexity and maintenance costs, supports multi-dimensional analysis and in-depth mining, and realizes the in-depth value conversion of medical data.
Smart Images

Figure CN119513142B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical data processing, and in particular relates to a method and device for generating a medical big data ES wide table based on efficient dynamic data configuration. Background Art
[0002] In today's digital age, big data has gradually penetrated into all aspects of people's lives, and medical data is no exception. With the advancement of the informatization process of the medical industry, doctors, medical research experts and patients have an increasing demand for viewing, researching and statistical data. However, at this stage, the design of most hospitals and medical systems is still limited to a single hospital or local system. Relational database systems such as MySQL and Oracle are usually used for data storage, and multi-table connections and queries are implemented through SQL. This traditional method leads to low data processing efficiency and cannot effectively adapt to the complexity of medical data structures and the diversity of data sources.
[0003] Since these systems can usually only be used within the hospital, the underlying data structure, data query and statistical functions cannot be extended across hospitals or regions. This limitation leads to low reuse of data resources and complex data processing, making it difficult to support the sharing and efficient application of medical data. In addition, the system standards of different hospitals and regions are not unified, lacking universality and compatibility, resulting in the need to redevelop the system when it is used across regions, making it impossible to form a universal platform with high adaptability and scalability. This situation not only significantly increases the cost of system development and maintenance, but also hinders the realization of the potential value of medical data, making it difficult to provide consistent and effective data support for diverse medical scenarios.
[0004] In general, the current common solutions have the following shortcomings:
[0005] (1) System limitations: The system architecture is usually limited to a single hospital or local area, making it difficult to achieve data sharing and collaboration across hospitals and regions;
[0006] (2) Low data storage and query efficiency: Relying on traditional relational databases, multi-table association queries are inefficient and difficult to support medical data with complex structures and diverse sources;
[0007] (3) Lack of data standardization and universality: Systems in different hospitals and regions lack unified standards and need to be redeveloped when used across institutions, making it impossible to form a universal platform with scalability.
[0008] (4) Complex data processing process: The data processing process is complex and time-consuming, especially when it involves multi-table associations, complex queries, and large amounts of data. The system response speed is low, the database table pressure is too high, and it cannot support medical applications with high concurrency.
[0009] (5) Difficulty in supporting multi-dimensional analysis and intelligence: Existing systems are mostly limited to basic statistical analysis and lack support for multi-dimensional analysis and intelligent algorithms, making it difficult to meet the in-depth needs of medical research and diagnosis;
[0010] (6) High maintenance cost: Due to the lack of flexibility and scalability in the system design, each upgrade or adjustment requires high maintenance costs, making it difficult to quickly adapt to new changes in demand.
[0011] Therefore, there is an urgent need to establish a medical data management and analysis platform with flexible data structure, strong system compatibility and cross-regional applicability, so as to further improve the efficiency of medical information sharing, promote the development of precision medicine and realize broader data resource integration. Summary of the invention
[0012] In view of the above, the purpose of the present invention is to provide a method and device for generating a medical big data ES wide table based on efficient dynamic data configuration, by configuring the data source, tables and fields to be integrated through the management background, and by configuring the driving engine for dynamic data collection, processing, conversion and fusion, and finally forming an Elasticsearch (ES) wide table that meets the needs of medical big data, achieving efficient integration and management of large-scale medical data, providing powerful data retrieval and analysis capabilities, and significantly improving the ability of cross-institutional data sharing, fast query and accurate analysis.
[0013] In order to achieve the above-mentioned invention object, the technical solution provided by the present invention is as follows:
[0014] The embodiment of the present invention provides a method for generating a medical big data ES wide table based on efficient dynamic data configuration, comprising the following steps:
[0015] Based on the built ES distributed cluster, dynamic data configuration including data source configuration, table configuration, field configuration and ES configuration is performed;
[0016] According to the dynamic data configuration, the personnel master table is queried in the dimension of people and the source table is queried in the dimension of medical consultation from the specified table of the data source through multi-threading to collect data in the required fields;
[0017] Convert the collected data into a unified ES field structure type according to the data type to obtain standardized data;
[0018] Based on the standardized data, data is assembled according to the dimensions of people and medical treatment to generate ES personnel wide table data and ES medical treatment wide table data. The two ES wide table data are updated in real time to the ES distributed cluster to obtain ES personnel wide table and ES medical treatment wide table.
[0019] Preferably, the data source configuration includes: configuring the data source for data reading, and selecting the built ES distributed cluster as the target end data source;
[0020] Table configuration includes: configuring the database instance, table, and table information that the data source needs to read;
[0021] Field configuration includes: configuring the field data to be collected, field data type, corresponding ES field, whether to save as an array collection, and if it is a vertical table field, you also need to configure the query conditions;
[0022] ES configuration includes: ES wide tables that support creation and deletion, wide tables that specify ES data writing, and wide tables for ES data query.
[0023] Preferably, the method of performing a personnel master table query in the dimension of person and a source table query in the dimension of medical consultation from a designated table of a data source through multiple threads according to dynamic data configuration to collect data of required fields includes:
[0024] Read the data source configuration, table configuration, and field configuration, group them according to the query conditions of the data source, table, and field, and assemble each group into a query SQL. If it is not a medical record table, the ID number needs to be added to the query SQL. If it is a medical record table, the ID number and unique consultation number need to be added to the query SQL;
[0025] According to the patient's ID number, query the personnel master table in the dimension of people, obtain the patient's basic information, and obtain the data object for querying the personnel master table;
[0026] According to the query SQL and the data object of the personnel main table query, the source table query is performed based on the dimension of medical consultation. The thread pool is used to perform multi-threaded concurrent query on the query SQL to obtain the patient medical consultation information and obtain the data object of the source table query;
[0027] The data objects queried from the personnel main table and the data objects queried from the source table are integrated to obtain integrated data objects.
[0028] Preferably, when querying the personnel main table, the ID number of the data interruption is used as the interruption point. If the data information of the interruption point exists in the memory, the data information of the interruption point in the memory is directly read. If the data information of the interruption point does not exist in the memory, the database information is read and loaded into the memory. The data information of the interruption point is stored as a global variable to avoid repeated queries in the loop, and is locked to ensure thread safety.
[0029] Preferably, the personnel master table query, source table query and data integration are executed asynchronously, that is, only a fixed amount of personnel data is queried in each batch according to the configuration. When the personnel master table and source table data query are completed, the next batch of personnel master table query and source table query can be executed without waiting for the data integration to be completed, and this process is executed continuously.
[0030] Preferably, based on the standardized data, the ES personnel wide table data is written according to the dimension of people to obtain the ES personnel wide table, including:
[0031] Delete the historical records of the ES personnel wide table according to the personnel ID range in the standardized data;
[0032] Create a Personnel Map object based on the data object queried from the Personnel Master Table in the standardized data, traverse the integrated data objects, convert the currently traversed data according to the field type of the source table and the field type of the ES data, store all medical-related data of each patient in the Personnel Map object, and if there is a configuration set, store multiple pieces of similar data in a list set sorted by visit time, and then store it in the Personnel Map object, and integrate all patient records to obtain the Personnel Wide Table data object;
[0033] With the ID number as the primary key, the integrated personnel wide table data object is asynchronously inserted into the ES personnel index to obtain the ES personnel wide table, and insertion exception compensation is added for re-insertion. If the compensation fails, the ID number that failed to be inserted is saved in the table, waiting for subsequent reprocessing of the failed data.
[0034] Preferably, based on the standardized data, the ES visit wide table data is written according to the visit dimension to obtain the ES visit wide table, including:
[0035] Delete the historical records of ES visits in the wide table according to the range of personnel ID cards in the standardized data;
[0036] Create a medical consultation Map object based on the data object queried from the source table in the standardized data, traverse the integrated data objects, and obtain the value of the created medical consultation Map object according to the unique medical consultation number; if the value of the created medical consultation Map object cannot be obtained according to the unique medical consultation number, obtain the corresponding personnel basic information Map information according to the ID card number, and deeply clone the personnel basic information Map information into the value of the medical consultation Map object, that is, the inner layer Map of the medical consultation Map object; if the value of the created medical consultation Map object can be obtained according to the unique medical consultation number, take out the inner layer Map of the medical consultation Map object, and then convert the currently traversed data according to the field type of the source table and the field type of the ES data, and finally store the data in the inner layer Map corresponding to the unique medical consultation number; combine all the data of each medical consultation into a record and store it in the medical consultation Map object, and integrate all medical consultation records to obtain the medical consultation wide table data object;
[0037] Using the unique medical consultation number as the primary key, the integrated medical consultation wide table data object is asynchronously inserted into the ES medical consultation index to obtain the ES medical consultation wide table, and insertion exception compensation is added and inserted again. If the compensation fails, the ID number that failed to be inserted will be saved in the table, waiting for subsequent reprocessing of the failed data.
[0038] Preferably, the generation of the ES personnel wide table and the generation of the ES visit wide table are concurrently executed through a thread pool.
[0039] Preferably, the threads of data collection and data writing support concurrent execution until all patient data are integrated. At the same time, they support reading the abnormal patient ID numbers inserted in the table through the abnormal data processing thread, and compensating and reinserting the abnormal patient data according to the configured number of failures. If the specified number of failures is reached, the failure compensation is stopped and a failure alarm is sent, introducing manual processing.
[0040] To achieve the above-mentioned purpose of the invention, the embodiment of the present invention further provides a device for generating a medical big data ES wide table based on efficient dynamic data configuration, which is implemented by the above-mentioned method for generating a medical big data ES wide table based on efficient dynamic data configuration, and includes: a data configuration module, a data acquisition module, a data standardization module, and an ES wide table generation module;
[0041] The data configuration module is used to perform dynamic data configuration including data source configuration, table configuration, field configuration and ES configuration based on the constructed ES distributed cluster;
[0042] The data acquisition module is used to query the personnel master table in the dimension of people and query the source table in the dimension of medical consultation from the designated table of the data source through multithreading according to the dynamic data configuration, and collect data of required fields;
[0043] The data standardization module is used to convert the collected data into a unified ES field structure type according to the data type to obtain standardized data;
[0044] The ES wide table generation module is used to assemble data according to the dimensions of people and medical treatment based on standardized data, generate ES personnel wide table data and ES medical treatment wide table data, and update the two ES wide table data in real time to the ES distributed cluster to obtain the ES personnel wide table and ES medical treatment wide table.
[0045] The present invention dynamically and efficiently constructs medical data into a medical big data ES wide table, aiming to solve the multi-dimensional standardization, fusion, storage and efficient query requirements of medical big data, and realize the construction and optimization of ES wide table. Compared with the prior art, the present invention has at least the following beneficial effects:
[0046] (1) Improve data sharing capabilities: By generating ES personnel wide tables and ES medical wide tables, data is standardized and uniformly managed, eliminating data format and semantic differences between different systems, improving data consistency and operability, eliminating cross-hospital and cross-regional data islands, and providing support for medical data sharing. Data from different hospitals can be seamlessly connected to achieve cross-hospital and cross-regional sharing and collaboration. Standardized data structures can provide a reliable data foundation in the fields of collaborative diagnosis and treatment, cross-regional health record integration, and medical research, laying a solid foundation for smart medical care and public health research.
[0047] (2) Significantly optimize query performance: By utilizing Elasticsearch's distributed architecture and structured wide table storage, multi-source data can be aggregated into a comprehensive data view, enabling efficient query and analysis of massive medical data. The wide table structure reduces multi-table associations during querying, which can significantly improve query performance and meet the medical system's needs for real-time query and high-concurrency response, making it possible to quickly query patient health information and case history data, providing rapid support for medical staff's decision-making.
[0048] (3) Simplify the data integration process: Adopt a configuration-driven data integration method to avoid complex multi-table associations and manual data conversion. The configuration-based data collection and data writing processing method significantly reduces the workload of data cleaning and format conversion, reduces the complexity and time cost of the data processing process, and improves the efficiency of data integration.
[0049] (4) Reduce development and maintenance costs: The configuration-driven design model brings a high degree of flexibility and can adapt to changes in different data sources and fields. Whether it is a horizontal table or a vertical table, the required fields and data can be easily integrated through dynamic configuration, breaking the limitations of data structure, reducing repeated development needs and manual intervention, and achieving rapid system expansion and deployment. Through unified data standards and platforms, there is no need to redevelop the system for each hospital or region, reducing operation and maintenance and update costs. This approach is suitable for changing medical needs and system upgrades, and simplifies the developer's response process to different needs.
[0050] (5) Support for rapid fusion of large-scale data: Through thread pools and multi-threaded concurrent processing, the system can complete the collection, processing and storage of large-scale data in a short time. This solution supports efficient data fusion, can cope with the growing demand for data volume in medical institutions, help medical data platforms quickly accumulate and update information, and provide strong support for big data analysis.
[0051] (6) Enhance the adaptability and scalability of the system: The ES wide table design supports flexible expansion of data fields and data sources. It can meet the data needs brought about by business growth without making major adjustments to the underlying architecture, adapt to future data types and medical information system update requirements, and provide stable support for the long-term expansion of the system.
[0052] (7) Support for multi-dimensional analysis and deep mining: With the help of wide table data structure and efficient distributed query, the system can conduct comprehensive analysis of multi-dimensional medical data and support scenarios such as precision medicine and intelligent diagnosis. By deeply mining information such as patient health trends, disease predictions, and drug reactions, it enables clinical decision-making and scientific research analysis, and realizes the deep value transformation of medical data. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0054] Figure 1 It is a flowchart of a method for generating a medical big data ES wide table based on efficient dynamic data configuration provided by an embodiment of the present invention;
[0055] Figure 2 It is a structural diagram of a medical big data ES wide table generation device based on efficient dynamic data configuration provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0056] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0057] The inventive concept of the present invention is: in view of the various difficulties faced by most medical data systems in the prior art, such as data islands, low query efficiency, and difficulty in realizing cross-institutional data sharing and analysis, the embodiment of the present invention provides a method and device for generating a medical big data ES wide table based on efficient dynamic data configuration, and obtains medical data from different hospitals with different structures by dynamically configuring data, so as to realize cross-institutional data sharing and collaborative analysis. And by establishing unified data standards and specifications, the collected heterogeneous data is standardized and processed, so that the data format is consistent and the semantics are unified, thereby eliminating data barriers between different systems. On this basis, according to different business needs and data configurations, the standardized data is dynamically integrated according to the configuration and stored in the medical big data Elasticsearch (ES) wide table, which helps to break the medical data island phenomenon, realize the systematic and standardized management of data, and provide a solid data foundation for cross-hospital and cross-regional medical services, scientific research analysis and precision diagnosis and treatment.
[0058] First, the terms involved in the embodiments of the present invention are explained as follows:
[0059] (1) Elasticsearch (ES): Elasticsearch is a distributed search and analysis engine that excels at handling large-scale data query and analysis tasks. It is built on Apache Lucene and supports full-text search, data analysis, log management and other functions. Elasticsearch is often used to build real-time search systems and data analysis platforms, and is particularly suitable for application scenarios that require processing massive amounts of data.
[0060] (2) ES index: Similar to a wide table of a relational data set, it is also called an ES wide table. In Elasticsearch, an index is the basic unit for organizing and managing data, similar to the concept of a "database" in a relational database. Each index can contain different types of documents and create an inverted index for them to support fast full-text retrieval and data analysis.
[0061] (3) ThreadPoolExecutor: It is the core class for managing thread pools in Java. It allows the creation and management of multiple threads to execute tasks in parallel, thereby improving the performance and efficiency of the application. By configuring the core parameters of the thread pool (such as the number of core threads, the maximum number of threads, the task queue, etc.), ThreadPoolExecutor can flexibly control the thread life cycle, task scheduling, timeout processing, etc. The use of thread pools can reduce the overhead of thread creation and destruction, and is suitable for scenarios with a large number of concurrent tasks.
[0062] (4) CountDownLatch: It is a synchronization auxiliary class in Java that is used to coordinate the execution order of multiple threads. It allows one or more threads to wait for other threads to complete operations before continuing to execute. CountDownLatch specifies a counter when it is created. Multiple threads can reduce the counter value through the countDown() method. When the counter value is reduced to 0, the waiting thread can continue to run. This is very practical for scenarios where you need to wait for multiple tasks to complete before executing a certain operation, such as task batch processing, concurrent testing, etc.
[0063] (5) Data Source: Data Source is the source of data for an application, usually a database. By configuring the data source, the program can more conveniently manage connections to the database, especially in multi-user and high-concurrency situations, such as managing the creation, use, and recycling of connections. In Java, a data source can also refer to a "connection pool" that provides database connections to avoid reconnecting to the database for each operation, thereby improving system efficiency.
[0064] (6) JDBC: JDBC (Java DataBase Connectivity) is a standard interface in Java for connecting to and operating databases. Through JDBC, Java programs can connect to different types of databases (such as MySQL, Oracle, etc.), execute SQL statements, and return query results. The main components of JDBC include Connection (connecting to the database), Statement (executing SQL statements), and ResultSet (processing query results). JDBC helps Java programs communicate with databases and simplifies database operations.
[0065] (7) SQL: SQL (Structured Query Language) is a standard language for managing and querying databases. It is a powerful language that is widely used in relational databases and can perform various operations such as querying data, inserting data, updating data, deleting data, and creating or modifying database structures (such as tables and columns).
[0066] Figure 1 FIG. 1 is a flow chart of a method for generating a medical big data ES wide table based on efficient dynamic data configuration provided by an embodiment of the present invention. Figure 1 As shown, the embodiment provides a method for generating a medical big data ES wide table based on efficient dynamic data configuration, comprising the following steps:
[0067] S1, based on the built ES distributed cluster, dynamic data configuration including data source configuration, table configuration, field configuration and ES configuration is performed.
[0068] S1.1, environment setup.
[0069] Build an Elasticsearch (ES) distributed cluster to ensure high availability and query efficiency.
[0070] S1.2, data source configuration.
[0071] The user configures the data source for data reading on the management platform. Data is collected from these data sources, and the established ES distributed cluster is selected as the target data source.
[0072] S1.3, table configuration.
[0073] Configure the database instance, table, and table information that the data source needs to read. The table information includes whether it is a case table, whether it is a horizontal table or a vertical table, the table's sorting time field, etc., and bind the table to the data source to ensure that the source of the collected data is accurate.
[0074] S1.4, field configuration.
[0075] Configure the field data to be collected, field data type, corresponding ES field, whether to save as an array set, and if it is a vertical table field, you also need to configure the query conditions. For example, the source table field to be queried is test_rslt_vlu_d, the corresponding ES field to be stored is urine_wbc_vlu, and the search condition is index_cd = 'TS01000001'. Bind the fields and query conditions to the target table and ES field to ensure the consistency of the data structure and the target format.
[0076] S1.5, ES configuration.
[0077] Supports creation and deletion of ES wide tables (corresponding to ES indexes), specifies wide tables (also indexes) for ES data writing, and wide tables (also indexes) for ES data query, and can add field structures to ES indexes. ES indexes do not have to be consistent, because the ES structure definition cannot be modified. In order to achieve seamless modification, data can be inserted into the new ES index, and after the insertion, the query index points to the new index and the old index is deleted.
[0078] S2, according to the dynamic data configuration, uses multi-threading to query the personnel master table in the dimension of people and the source table in the dimension of medical treatment from the specified table of the data source to collect data in the required fields.
[0079] S2.1, SQL assembly.
[0080] Read the data source configuration, table configuration and field configuration, group them according to the query conditions of the data source, table and field, assemble each group into a query SQL, sort in ascending order according to the sorting time field of the table. If it is not a medical record table, the ID number needs to be added to the query SQL. If it is a medical record table, the ID number and unique consultation number need to be added to the query SQL.
[0081] S2.2, break point query.
[0082] When querying the personnel master table, the ID number of the data interruption is used as the interruption point. If the data information of the interruption point exists in the memory, the data information of the interruption point in the memory is directly read. If the data information of the interruption point does not exist in the memory, the database information is read and loaded into the memory. The data information of the interruption point is stored as a global variable to avoid repeated queries in the loop, and is locked to ensure thread safety.
[0083] S2.3, personnel master table query.
[0084] According to the patient's ID number, query the main table of personnel in the dimension of people, obtain the basic information of the patient, sort according to the patient's ID card, query fixed personnel data each time according to the configuration, such as ID cards greater than the currently queried, sort and query 1,000 patient information, and return the List of the main table query <Map<String,Object> >Data object, the key is the corresponding field name, and the value is the field value.
[0085] S2.4, source table query.
[0086] Assemble the user ID number into the query SQL to get the complete query SQL, use the ThreadPoolExecutor thread pool to query the query SQL concurrently, increase the concurrent query capability by configuring the number of core threads and the maximum number of threads, speed up data query efficiency, use JDBC query and configure the number of failed retries (the default is 3 times), and query only a fixed number of patients' outpatient, hospitalization, case, test, and other data in each batch according to the configuration (the default value is 1000), and return the List of source table queries. <Map<String,Object> >Data object.
[0087] S2.5, data collection is completed.
[0088] Use CountDownLatch (program counter) to wait for the thread pool to complete all SQL data queries, then integrate the data objects queried from the personnel main table with the data objects queried from the source table, and return the integrated List <List<Map<String,Object> >>Data object, the inner collection is the data collection of each combined SQL query, and the outer collection is the combined collection of data from all SQL queries.
[0089] It should be noted that the personnel master table query, source table query and data integration are executed asynchronously, that is, according to the configuration, only a fixed number of personnel data are queried in each batch. When the personnel master table and source table data query are completed, the next batch of personnel master table query and source table query can be executed without waiting for the data integration to be completed, and this process is executed continuously, thus forming an efficient data processing pipeline. Even if the processing time of a certain step is slightly longer, it will not block the entire process, thereby ensuring the overall high concurrency and response speed, and making full use of system resources.
[0090] S3, convert the collected data into a unified ES field structure type according to the data type to obtain standardized data.
[0091] The collected patient data in the integrated data object is converted into a unified ES field structure type according to the data type. For example, Timestamp and Date are uniformly converted into LocalDateTime.
[0092] S4, based on the standardized data, assemble the data according to the dimensions of people and medical treatment, generate ES personnel wide table data and ES medical treatment wide table data, update the two ES wide table data in real time to the ES distributed cluster, and obtain ES personnel wide table and ES medical treatment wide table.
[0093] Use the ThreadPoolExecutor thread pool concurrently to assemble and write large wide table data in the personnel and medical consultation dimensions, including the following two aspects.
[0094] S4.1, based on the standardized data, write the ES personnel wide table data according to the dimension of people to obtain the ES personnel wide table.
[0095] S4.1.1, personnel wide table data deletion: Delete the historical records of the ES personnel wide table according to the personnel ID card range in the standardized data. Because the source data may be deleted, if it is directly inserted, the originally deleted data will still be saved in the ES wide table (index). By deleting the historical records, dirty data can be avoided.
[0096] S4.1.2, personnel wide table data conversion and integration: Create a personnel Map object based on the data object of the personnel master table query in the standardized data, that is, the personnel Map <String,Map<String,Object> > object, the outer key is the user's ID number, and the value corresponding to the Map is the user's data. Traverse the integrated List <List<Map<String,Object> >>Data object, convert the currently traversed data according to the field type of the source table and the field type of the ES data, such as int (integer), long (long integer), double (double-precision floating point), boolean (Boolean), date (date type), etc., and store all medical-related data of each patient (including symptoms, tests, examinations, medical records, surgeries, drugs, medical history information, etc.) in the personnel Map object. If the field is configured with a collection, create a List to store the data. If a patient has multiple similar data, they will be sorted by the time of the visit and stored in the List list collection. All symptoms of the patient can be directly queried through the personnel, and the first, last, minimum, maximum and other data can be obtained as quickly as possible. Finally, the converted data is stored in the personnel Map object, one patient is integrated into one record, and the records of all patients are integrated to obtain the personnel wide table List <Map<String,Object> >Data object.
[0097] S4.1.3, personnel wide table data insertion: With ID number as primary key, insert the integrated personnel wide table List <Map<String,Object> >The data object is asynchronously inserted into the ES personnel index to obtain the ES personnel wide table, and insertion exception compensation is added for re-insertion. If the compensation fails, the ID card number that failed to be inserted is saved in the table, waiting for subsequent reprocessing of the failed data.
[0098] S4.2, based on the standardized data, the ES medical consultation wide table data is written according to the medical consultation dimension to obtain the ES medical consultation wide table.
[0099] S4.2.1, Deletion of data in the wide table of medical consultations: Delete the historical records of the ES wide table of medical consultations according to the range of the ID cards of the personnel in the standardized data. Because the source data may be deleted, if it is directly inserted, the originally deleted data will still be saved in the wide table (index) of ES medical consultations. By deleting the historical records, dirty data can be avoided.
[0100] S4.2.2, Medical consultation wide table data conversion and fusion: Create a medical consultation map object based on the data object of the source table query in the standardized data, that is, the medical consultation map <String,Map<String,Object> >Object, the outer key is the unique medical number, and the value is the data of that visit (corresponding to the inner Map). Traverse the integrated List <List<Map<String,Object> >>Data object, get the value of the created medical consultation Map object according to the unique medical consultation number. If the value of the created medical consultation Map object cannot be obtained according to the unique medical consultation number, then obtain the corresponding personnel basic information Map information according to the ID number, and deep clone the personnel basic information Map information into the value of the medical consultation Map object (corresponding inner map). If the value of the created medical consultation Map object can be obtained according to the unique medical consultation number, then take out the inner map of the medical consultation Map object, and then convert the currently traversed data according to the field type of the source table and the field type of the ES data, and finally store the data in the inner map corresponding to the unique medical consultation number. Combine all the data of each medical consultation into one record and store it in the medical consultation Map object, and integrate all medical consultation records to obtain the wide medical consultation table List. <Map<String,Object> >Data object.
[0101] S4.2.3, insert data into the wide table of medical consultations: use the unique medical consultation number as the primary key, and insert the integrated wide table List <Map<String,Object> >The data object is asynchronously inserted into the ES medical consultation index to obtain the ES medical consultation wide table, which reduces the waiting time for data insertion and increases insertion exception compensation for re-insertion. If the compensation fails, the ID card number that failed to be inserted will be saved in the table, waiting for subsequent reprocessing of the failed data.
[0102] S4.3, update breakpoints.
[0103] Wait for the personnel wide table and the medical consultation wide table threads to complete execution (the actual data insertion is an asynchronous operation and has not yet been completed, saving time for data insertion) to update the interruption point of data processing in the memory and database to ensure that the data can be synchronized from the interruption point when the process stops abnormally or restarts.
[0104] S4.4, data continues to be processed.
[0105] After the breakpoint update is completed, the next data collection will be carried out. The threads of data collection and data writing can be executed concurrently until all patient data are integrated. At the same time, the abnormal data processing thread will read the abnormal patient ID number inserted in the table, and re-insert the abnormal patient data according to the configured number of failures. If the specified number of failures is reached, the failure compensation will be stopped and a failure alarm will be sent, introducing manual processing.
[0106] S4.5, ES wide table data query.
[0107] The engine generates two ES large wide tables every day according to the dimensions of person and medical number. It uses the high availability, efficient query capabilities, rich query functions and powerful aggregation analysis capabilities of the ES distributed architecture to provide strong data support to the application side, fully unleash the value of massive data and solve the problem of data silos.
[0108] In summary, the medical big data ES wide table generation method based on efficient dynamic data configuration provided by the embodiment of the present invention realizes standardized data management and efficient query through Elasticsearch wide table in the field of medical big data based on efficient dynamic data configuration. This method simplifies the data fusion process by dynamically configuring the structure of data sources, fields and wide tables, and significantly improves the performance of cross-institutional data sharing and high-concurrency queries. Combined with the distributed architecture of Elasticsearch, the system has the ability to efficiently store and retrieve large-scale data, providing strong support for medical data sharing, scientific research analysis, and intelligent diagnosis, and realizing a highly adaptable, low-cost and high-performance medical big data management model, which has broad application prospects.
[0109] Based on the same inventive concept, Figure 2 As shown, an embodiment of the present invention also provides a medical big data ES wide table generation device 200 based on efficient dynamic data configuration, including: a data configuration module 210, a data acquisition module 220, a data standardization module 230, and an ES wide table generation module 240.
[0110] The data configuration module 210 is used to perform dynamic data configuration including data source configuration, table configuration, field configuration and ES configuration based on the constructed ES distributed cluster.
[0111] The data collection module 220 is used to query the personnel master table in the dimension of people and query the source table in the dimension of medical consultation from the designated table of the data source through multi-threading according to the dynamic data configuration, and collect data of required fields.
[0112] The data standardization module 230 is used to convert the collected data into a unified ES field structure type according to the data type to obtain standardized data.
[0113] The ES wide table generation module 240 is used to assemble data according to the dimensions of people and medical consultations based on standardized data, generate ES personnel wide table data and ES medical consultation wide table data, and update the two ES wide table data in real time to the ES distributed cluster to obtain the ES personnel wide table and ES medical consultation wide table.
[0114] It should be noted that the medical big data ES wide table generation device based on efficient dynamic data configuration provided in the embodiment of the present invention and the medical big data ES wide table generation method based on efficient dynamic data configuration belong to the same inventive concept. The specific implementation process is detailed in the embodiment of the medical big data ES wide table generation method based on efficient dynamic data configuration, which will not be repeated here.
[0115] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for generating a medical big data ES wide table based on efficient dynamic data configuration, characterized in that: The following steps are involved: Based on the built ES distributed cluster, dynamic data configuration including data source configuration, table configuration, field configuration and ES configuration is performed; According to the dynamic data configuration, the personnel master table is queried in the dimension of people and the source table is queried in the dimension of medical treatment from the designated table of the data source through multi-threading, and the data of the required fields are collected, including: reading the data source configuration, table configuration and field configuration, grouping according to the query conditions of the data source, table and field, and assembling each group into a query SQL. If it is not a medical record table, the ID number needs to be added to the query SQL. If it is a medical record table, the ID number and the unique medical treatment number need to be added to the query SQL; according to the patient's ID number, the personnel master table is queried in the dimension of people to obtain the patient's basic information and obtain the data object of the personnel master table query; according to the query SQL and the data object of the personnel master table query, the source table is queried in the dimension of medical treatment, and the query SQL is queried by multi-thread concurrently using the thread pool to obtain the patient's medical treatment information and obtain the data object of the source table query; the data object of the personnel master table query and the data object of the source table query are integrated to obtain the integrated data object; Convert the collected data into a unified ES field structure type according to the data type to obtain standardized data; Based on the standardized data, data is assembled according to the dimensions of people and medical treatment to generate ES personnel wide table data and ES medical treatment wide table data. The two ES wide table data are updated in real time to the ES distributed cluster to obtain ES personnel wide table and ES medical treatment wide table.
2. The method for generating a medical big data ES wide table based on efficient dynamic data configuration according to claim 1, characterized in that: Data source configuration includes: configuring the data source for data reading, and selecting the built ES distributed cluster as the target data source; Table configuration includes: configuring the database instance, table, and table information that the data source needs to read; Field configuration includes: configuring the field data to be collected, field data type, corresponding ES field, whether to save as an array collection, and if it is a vertical table field, you also need to configure the query conditions; ES configuration includes: ES wide tables that support creation and deletion, wide tables that specify ES data writing, and wide tables for ES data query.
3. The method for generating a medical big data ES wide table based on efficient dynamic data configuration according to claim 1, characterized in that: When querying the personnel master table, the ID number of the data interruption is used as the interruption point. If the data information of the interruption point exists in the memory, the data information of the interruption point in the memory is directly read. If the data information of the interruption point does not exist in the memory, the database information is read and loaded into the memory. The data information of the interruption point is stored as a global variable to avoid repeated queries in the loop, and is locked to ensure thread safety.
4. The method for generating a medical big data ES wide table based on efficient dynamic data configuration according to claim 1, characterized in that: Personnel master table query, source table query and data integration are executed asynchronously, that is, according to the configuration, only a fixed number of personnel data are queried in each batch. When the personnel master table and source table data query are completed, the next batch of personnel master table query and source table query can be executed without waiting for data integration to be completed, and this process is executed continuously.
5. The method for generating a medical big data ES wide table based on efficient dynamic data configuration according to claim 1, characterized in that: Based on the standardized data, the ES personnel wide table data is written according to the dimension of people to obtain the ES personnel wide table, including: Delete the historical records of the ES personnel wide table according to the personnel ID range in the standardized data; Create a Personnel Map object based on the data object queried from the Personnel Master Table in the standardized data, traverse the integrated data objects, convert the currently traversed data according to the field type of the source table and the field type of the ES data, store all medical-related data of each patient in the Personnel Map object, and if there is a configuration set, store multiple pieces of similar data in a list set sorted by visit time, and then store it in the Personnel Map object, and integrate all patient records to obtain the Personnel Wide Table data object; With the ID number as the primary key, the integrated personnel wide table data object is asynchronously inserted into the ES personnel index to obtain the ES personnel wide table, and insertion exception compensation is added for re-insertion. If the compensation fails, the ID number that failed to be inserted is saved in the table, waiting for subsequent reprocessing of the failed data.
6. The method for generating a medical big data ES wide table based on efficient dynamic data configuration according to claim 1, characterized in that: Based on the standardized data, the ES visit wide table data is written according to the visit dimension to obtain the ES visit wide table, including: Delete the historical records of ES visits in the wide table according to the range of personnel ID cards in the standardized data; Create a medical consultation Map object based on the data object queried from the source table in the standardized data, traverse the integrated data objects, and obtain the value of the created medical consultation Map object according to the unique medical consultation number; if the value of the created medical consultation Map object cannot be obtained according to the unique medical consultation number, obtain the corresponding personnel basic information Map information according to the ID card number, and deeply clone the personnel basic information Map information into the value of the medical consultation Map object, that is, the inner layer Map of the medical consultation Map object; if the value of the created medical consultation Map object can be obtained according to the unique medical consultation number, take out the inner layer Map of the medical consultation Map object, and then convert the currently traversed data according to the field type of the source table and the field type of the ES data, and finally store the data in the inner layer Map corresponding to the unique medical consultation number; combine all the data of each medical consultation into a record and store it in the medical consultation Map object, and integrate all medical consultation records to obtain the medical consultation wide table data object; Using the unique medical consultation number as the primary key, the integrated medical consultation wide table data object is asynchronously inserted into the ES medical consultation index to obtain the ES medical consultation wide table, and insertion exception compensation is added and inserted again. If the compensation fails, the ID number that failed to be inserted will be saved in the table, waiting for subsequent reprocessing of the failed data.
7. The method for generating a medical big data ES wide table based on efficient dynamic data configuration according to claim 1, characterized in that: The generation of the ES personnel wide table and the generation of the ES visit wide table are concurrently executed through the thread pool.
8. The method for generating a medical big data ES wide table based on efficient dynamic data configuration according to claim 1, characterized in that: The threads for data collection and data writing support concurrent execution until all patient data are integrated. At the same time, they support reading abnormal patient ID numbers inserted into the table through the abnormal data processing thread, and re-inserting the abnormal patient data based on the configured number of failures. If the specified number of failures is reached, the failure compensation is stopped and a failure alarm is sent, introducing manual processing.
9. A device for generating a medical big data ES wide table based on efficient dynamic data configuration, which is implemented by using the method for generating a medical big data ES wide table based on efficient dynamic data configuration according to any one of claims 1 to 8, characterized in that: include: Data configuration module, data collection module, data standardization module, and ES wide table generation module; The data configuration module is used to perform dynamic data configuration including data source configuration, table configuration, field configuration and ES configuration based on the constructed ES distributed cluster; The data acquisition module is used to query the personnel master table in the dimension of people and the source table in the dimension of medical treatment from the designated table of the data source through multi-threading according to the dynamic data configuration, and collect data of the required fields, including: reading the data source configuration, table configuration and field configuration, grouping according to the query conditions of the data source, table and field, assembling each group into a query SQL, if it is not a medical record table, the ID number needs to be added to the query SQL, if it is a medical record table, the ID number and the unique medical treatment number need to be added to the query SQL; query the personnel master table in the dimension of people according to the patient's ID number, obtain the patient's basic information, and obtain the data object of the personnel master table query; query the source table in the dimension of medical treatment according to the query SQL and the data object of the personnel master table query, use the thread pool to perform multi-threaded concurrent query on the query SQL, obtain the patient's medical treatment information, and obtain the data object of the source table query; integrate the data object of the personnel master table query and the data object of the source table query to obtain an integrated data object; The data standardization module is used to convert the collected data into a unified ES field structure type according to the data type to obtain standardized data; The ES wide table generation module is used to assemble data according to the dimensions of people and medical treatment based on standardized data, generate ES personnel wide table data and ES medical treatment wide table data, and update the two ES wide table data in real time to the ES distributed cluster to obtain the ES personnel wide table and ES medical treatment wide table.
Citation Information
Patent Citations
Label management method and device
CN115168361A
Query method and system based on big data medical general retrieval index construction
CN115563127A