Medical data multi-condition connection query optimization method and system based on data lake

By dividing the main tables and secondary tables in the medical database, and establishing a high-dimensional index for the secondary tables, and dividing hot tables and cold tables according to the query frequency, the problem of inefficiency and large space overhead of the medical data exploration framework based on the data lake is solved, and efficient medical data query and analysis is achieved.

CN120216550APending Publication Date: 2025-06-27HENAN UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510262574.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing medical data exploration framework based on data lakes has problems such as inefficient query and large space overhead when processing medical data multi-table connection queries.

Method used

By dividing the medical data table into main tables and secondary tables in the medical database, a query frequency table is created, a high-dimensional index is established for the secondary tables based on the medical query behavior characteristics, and the secondary table is divided into hot tables and cold tables according to the query frequency, and a preset query strategy is executed to optimize multi-condition join query.

Benefits of technology

It improves the efficiency of multi-condition connection query of medical data, reduces spatial overhead, and realizes rapid query and analysis of medical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216550A_ABST
    Figure CN120216550A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical big data processing, in particular to a medical data multi-condition connection query optimization method and system based on a data lake, and the method comprises the steps: dividing a medical data table into a main table and a secondary table in a medical database, and creating a query frequency table for recording the query frequency of the secondary table; establishing a high-dimensional index for the secondary table according to the medical query behavior characteristics, and uploading the secondary table to a data lake for storage according to an index value; the secondary table is divided into a hot table and a cold table according to the medical data query frequency, the hot table and the main table are fused, a fusion table and a fusion table name are generated according to the table name of the main table and the table name of the hot table, and the fusion table is transmitted to the data lake based on the fusion table name; and extracting a target primary table name and a target secondary table name in the multi-conditional connection query, executing a corresponding query strategy according to the target secondary table type, and obtaining the multi-conditional connection query medical data. According to the method, the space consumption and the time consumption can be well balanced, and the medical data multi-condition query access efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical big data processing, and particularly relates to a method and system for optimizing multi-condition join queries of medical data based on a data lake. Background Art

[0002] With the rapid development of information technology and various digital medical devices, and the increasing variety of medical devices, the data generated by various medical devices are recorded, making medical data more refined and accurate. Medical professionals rely on the data provided by various medical devices to comprehensively evaluate the health status of patients. By using data exploration techniques, extracting and integrating the information required by medical professionals from these data can effectively assist medical professionals in making treatment decisions. Among various data exploration methods, complex queries are a common method for medical data exploration, involving complex multi-condition filtering and multi-table join operations. Through the analysis of the complex query process, it is found that in the subqueries of complex queries, multi-condition queries occur frequently and take a long time.

[0003] Database technology was adopted earlier in the medical field to handle complex queries on structured medical data and provide data support for treatment decisions. It can meet the need for rapid data query in the case of less medical data. However, when it comes to analysis tasks involving a large amount of data, database technology may encounter limitations. To solve this problem, data warehouse technology is applied to manage large-scale structured medical data and utilize the computing power of distributed clusters to improve data query efficiency. Data warehouse technology provides improved performance for processing extensive data sets. In recent years, due to factors such as the increasing demand for efficient data query and the exponential growth of unstructured data in the healthcare field, data warehouse systems have encountered challenges in the storage of unstructured data.

[0004] To solve the problems of databases and data warehouses in medical data exploration, researchers have begun to apply data lake technology to medical data exploration. The data storage system based on the data lake has better data aggregation capabilities, can process tables with larger-scale data, and can also store and manage medical data with different structures. However, the current medical data exploration framework based on data lake technology faces the following two challenges in processing multi-table join queries of medical data: in the current framework, the solutions for high-frequency join queries in the medical field always may generate unnecessary space overhead; in the real scenario of medical field queries, medical-related personnel often need to quickly display query results, yet the existing framework has the problem of low efficiency in the commonly used join queries in the medical field, which may prevent medical-related personnel from efficiently analyzing the data generated by medical devices. Summary of the Invention

[0005] To this end, the present invention provides a method and system for optimizing multi-condition join queries of medical data based on a data lake, which solves the problems of low query efficiency and large query space overhead existing in the execution of multi-condition join queries in existing medical data lakes.

[0006] According to the design solution provided by the present invention, on the one hand, a method for optimizing multi-condition join queries of medical data based on a data lake is provided, including:

[0007] In a medical database, medical data tables are divided into main tables and secondary tables, and a query frequency table for recording the query frequencies of secondary tables is created. The medical database includes several medical data tables. The query frequency table includes a secondary table name field and a corresponding query frequency field. The main table is used to record medical entities, and the secondary table is used to record medical entity-related data. There is a one-to-many relationship between the data of the secondary table and the main table;

[0008] A high-dimensional index is established for the secondary table according to the characteristics of medical query behavior, and the secondary table is uploaded to the data lake for storage according to the index value;

[0009] The secondary table is divided into a hot table and a cold table according to the medical data query frequency. The hot table is fused with the main table, and a fusion table and a fusion table name are generated based on the main table name and the hot table name. The fusion table is transmitted to the data lake based on the fusion table name. The hot table is a secondary table with a medical data query frequency greater than the query threshold, and the cold table is a secondary table with a medical data query frequency less than or equal to the query threshold;

[0010] The target main table name and the target secondary table name in the multi-condition join query are extracted, the type of the target secondary table is judged according to the query frequency of the target secondary table name, and the preset corresponding query strategy is executed according to the target secondary table type to obtain the multi-condition join query medical data.

[0011] As the method for optimizing multi-condition join queries of medical data based on a data lake of the present invention, further, the medical entities in the main table are stored in a structured form, and each piece of data corresponds to a medical entity; the medical entity-related data in the secondary table is stored in a structured form or an unstructured form, and the secondary table uses historical query medical data or relevant personnel experience data as the join query condition.

[0012] As the method for optimizing multi-condition join queries of medical data based on a data lake of the present invention, further, establishing a high-dimensional index for the secondary table according to the characteristics of medical query behavior includes:

[0013] Randomly select a secondary table in the medical database as a common secondary table, and read the column names of the query conditions that have been recorded in the common secondary table;

[0014] Create a high-dimensional index for the secondary table based on the column names of the query conditions, and reduce the values of multiple columns to a single value through a dimensionality reduction method. Add the index values obtained after dimensionality reduction to the secondary table.

[0015] As the method for optimizing multi-condition join queries of medical data based on a data lake in the present invention, further, upload the secondary table to the data lake for storage according to the index value, including:

[0016] Read the data of the secondary table containing the index value, and sort the data of the secondary table in ascending order according to the size of the index value;

[0017] Upload the sorted data of the secondary table to the data lake with the primary key as the record key and the index column as the partition key.

[0018] As the method for optimizing multi-condition join queries of medical data based on a data lake in the present invention, further, divide the secondary table into a hot table and a cold table according to the query frequency of medical data, including:

[0019] Set a query threshold according to historical query records or the experience data of relevant personnel;

[0020] Use the secondary table name in the current query as the query condition, and obtain the query frequency value of the corresponding query frequency field in the query frequency table according to this query condition;

[0021] If the query frequency value is greater than the query threshold, recognize the secondary table in the current query as a hot table; otherwise, recognize the secondary table in the current query as a cold table.

[0022] As the method for optimizing multi-condition join queries of medical data based on a data lake in the present invention, further, fuse the hot table with the primary table, including:

[0023] Read the data of the primary table and the data of the secondary table;

[0024] Fuse the data of the primary table and the data of the secondary table into a data table structure, and then type the fused data table structure into the data lake with the primary key of the primary table as the record key and the index column of the secondary table as the partition. Among them, the table name of the data table structure is set according to the primary table name and the secondary table name.

[0025] As the method for optimizing multi-condition join queries of medical data based on a data lake in the present invention, further, execute a preset corresponding query strategy according to the target secondary table type, including:

[0026] If the target secondary table is a hot table and there is a fused table, the query strategy is set as: convert the multi-condition join query into a single-table condition query on the fused table, and execute the single-table condition query in the fused table;

[0027] When the target secondary table is a hot table but there is no fusion table, the query strategy is set as follows: perform a multi-condition join query, obtain the corresponding primary table data, and fuse the hot table with the corresponding primary table data;

[0028] If the target secondary table is a cold table, the query strategy is set as: perform a multi-condition join query in the medical database.

[0029] On the other hand, the present invention also provides a medical data multi-condition join query optimization system based on a data lake, including: a data partitioning module, an index building module, a data fusion module, and a join query module, where,

[0030] The data partitioning module is used to partition the medical data tables in the medical database into primary tables and secondary tables, and create a query frequency table for recording the query frequencies of the secondary tables. The medical database includes several medical data tables. The query frequency table includes a secondary table name field and a corresponding query frequency field. The primary table is used to record medical entities, and the secondary table is used to record medical entity-related data, and there is a one-to-many relationship between the data of the secondary table and the primary table;

[0031] The index building module is used to build a high-dimensional index for the secondary table according to the characteristics of medical query behaviors, and upload the secondary table to the data lake for storage according to the index values;

[0032] The data fusion module is used to partition the secondary tables into hot tables and cold tables according to the medical data query frequencies, fuse the hot tables with the primary tables, and generate a fusion table and a fusion table name based on the primary table name and the hot table name, and upload the fusion table to the data lake based on the fusion table name. The hot table is a secondary table with a medical data query frequency greater than the query threshold, and the cold table is a secondary table with a medical data query frequency less than or equal to the query threshold;

[0033] The join query module is used to extract the target primary table name and the target secondary table name in the multi-condition join query, judge the type of the target secondary table according to the query frequency of the target secondary table name, and execute the preset corresponding query strategy according to the target secondary table type and obtain the multi-condition join query medical data.

[0034] Advantages of the present invention:

[0035] Before writing the medical data tables into the data lake, the present invention classifies the data tables into primary tables and secondary tables, builds a high-dimensional index for the secondary tables, thereby accelerating the screening process speed in the join query; and uses the index values as the basis for physical data storage, accelerating the reading process speed in the join query, and executing different query strategies according to the classification of the secondary tables, achieving a balance between space consumption and time consumption, improving the multi-condition query access efficiency of medical data, and facilitating deployment and implementation in the field of medical data. Description of the Drawings

[0036] Figure 1 Schematic diagram of the optimization process for multi - conditional join query of medical data based on a data lake in the embodiment;

[0037] Figure 2 Schematic diagram of the architecture of the optimization system for multi - conditional join query of medical data in the embodiment. Detailed implementation manners

[0038] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and technical solutions.

[0039] In view of the problems of low query efficiency and large space overhead in the existing medical data lake when performing multi - conditional join queries, embodiments of the present invention, as shown in Figure 1 provide an optimization method for multi - conditional join query of medical data based on a data lake, including:

[0040] S101. Divide the medical data tables in the medical database into a main table and a secondary table, and create a query frequency table for recording the query frequencies of the secondary tables. The medical database includes several medical data tables. The query frequency table includes a secondary table name field and a corresponding query frequency field. The main table is used to record medical entities, and the secondary table is used to record data related to medical entities. There is a many - to - one relationship between the data of the secondary table and the main table.

[0041] Among them, the medical entities in the main table are stored in a structured form, and each piece of data corresponds to a medical entity; the data related to medical entities in the secondary table are stored in a structured form or an unstructured form, and the secondary table uses historical query medical data or relevant personnel experience data as join query conditions.

[0042] The medical data tables are classified into primary tables and secondary tables according to historical query information or driven by the experience of medical staff: In a join query, the data tables serving as the left table, such as the patient information table, the hospitalization information table, etc.; each piece of data of which is an entity, without loss during the join query process, and it is usually structured data. This type of data table can be regarded as the primary table; while the data tables serving as the right table in the join query, such as the vital sign information, the test data information, etc.; the information in the table has a many-to-many or many-to-one relationship with the primary table, mostly data generated by medical devices, mostly structured data, and some are unstructured data. This type of data table can be regarded as the secondary table; for the primary table, the built-in reading method of Spark can be used to process the structured data into the DataFrame format, and the primary key is used as the record key, and the foreign key is used as the partition key to store it in the data lake; for the secondary table belonging to structured data, the commonly used condition column names of its join query are recorded driven by historical query data or the experience of medical staff; for the secondary table belonging to unstructured data, it can also be converted into a unified DataFrame format through the built-in method of Spark to achieve structuring; create a query frequency table, which includes the secondary table name field and the query frequency field. Record the table name of the secondary table into the query frequency table, and the query frequency is recorded as 0.

[0043] S102. Establish a high-dimensional index for the secondary table according to the characteristics of medical query behavior, and upload the secondary table to the data lake for storage according to the index value.

[0044] Specifically, establishing a high-dimensional index for the secondary table according to the characteristics of medical query behavior can be designed to include:

[0045] Randomly select a secondary table in the medical database as the commonly used secondary table, and read the query condition column names that have been recorded in the commonly used secondary table;

[0046] Create a high-dimensional index for the secondary table according to the query condition column names, and reduce the values of multiple columns to a single value through the dimensionality reduction method, and add the obtained index value after dimensionality reduction to the secondary table.

[0047] In a medical database, there are usually multiple secondary tables that need to be processed separately. First, select a secondary table and read the column names of the common query conditions that have been recorded in this table. Create a high-dimensional index based on the column names of the query conditions, and use a dimensionality reduction method to reduce the numerical features of multiple columns into a single numerical value. For example, use dimensionality reduction methods such as Zorder. If the Zorder dimensionality reduction method is used, the index values obtained after dimensionality reduction will be calculated in a bit-interleaved manner, and this calculation method retains the characteristics of multi-dimensional data. Use spark to read the selected secondary table in Dataframe format, and add the index values obtained after dimensionality reduction of the data contained in the selected common query condition columns as a new column to the Dataframe generated by the secondary table, thereby generating a Dataframe containing index information.

[0048] Sort the secondary table in ascending order according to the added index values, and store it in physical space according to the index values.

[0049] After obtaining the Dataframe containing index information, use the built-in components of spark to sort it in ascending order according to the values in the added index column. Then, use the primary key as the record key and the index column as the partition key for the sorted Dataframe to cluster. Through this operation, points that are close in multi-dimensional space can be made adjacent in physical space, and there will be higher reading efficiency in subsequent queries.

[0050] S103. Divide the secondary tables into hot tables and cold tables according to the medical data query frequency, fuse the hot tables with the main table, and generate a fusion table and a fusion table name based on the main table name and the hot table name. Transmit the fusion table to the data lake based on the fusion table name. The hot table is a secondary table with a medical data query frequency greater than the query threshold, and the cold table is a secondary table with a medical data query frequency less than or equal to the query threshold.

[0051] Specifically, dividing the secondary tables into hot tables and cold tables according to the medical data query frequency can be designed to include:

[0052] Set the query threshold based on historical query records or relevant personnel's experience data;

[0053] Use the secondary table name in the current query as the query condition, and obtain the query frequency value of the corresponding query frequency field in the query frequency table according to this query condition;

[0054] If the query frequency value is greater than the query threshold, recognize the secondary table in the current query as a hot table; otherwise, recognize the secondary table in the current query as a cold table.

[0055] By reading the primary table data and secondary table data; integrating the primary table data and secondary table data into a data table structure, and after integration, using the primary key of the primary table as the record key and the index column of the secondary table as the partition to type the integrated data table structure into the data lake, where the table name of the data table structure is set according to the primary table name and secondary table name.

[0056] When performing a join query, the query frequency table will be updated. Each time the secondary table is queried, the corresponding query frequency value will be incremented by 1. Set a query threshold based on historical query records or driven by the experience of medical staff. When the system performs a join query, it will query the query frequency table with the secondary table name of the current query as the condition. If the query frequency value is greater than the threshold, it is considered a hot table, otherwise it is a cold table.

[0057] When performing a join query, if it is determined that the secondary table of the current query is a hot table and there is no fusion table. The primary table and secondary table of the current join query will be fused. The specific fusion process includes: using the spark component to unify the primary table and secondary table into the dataframe format, and then using the built-in Join method of spark to fuse the two dataframes into one dataframe. The newly generated dataframe is the union of the two data tables. Finally, using the primary key of the primary table as the record key and the index column of the secondary table as the partition key to put the newly generated dataframe into the lake, and the table name is in the form of the primary table name + secondary table name. The above process specifically refers to that the current join query only contains one primary table and one secondary table. If the join query contains one primary table and multiple secondary tables, it can be decomposed into a simple single-secondary-table join query form and the above process can be carried out in turn.

[0058] S104. Extract the target primary table name and target secondary table name in the multi-condition join query, judge the target secondary table type according to the query frequency of the target secondary table name, and execute the preset corresponding query strategy according to the target secondary table type to obtain the multi-condition join query medical data.

[0059] Specifically, the preset corresponding query strategy executed according to the target secondary table type can be designed to include:

[0060] If the target secondary table is a hot table and there is a fusion table, the query strategy is set as: transforming the multi-condition join query into a single-table conditional query on the fusion table and executing the single-table conditional query in the fusion table;

[0061] If the target secondary table is a hot table but there is no fusion table, the query strategy is set as: executing the multi-condition join query, obtaining the corresponding primary table data, and fusing the hot table with the corresponding primary table data;

[0062] If the target secondary table is a cold table, the query strategy is set as: executing the multi-condition join query in the medical database.

[0063] When performing a join query, the program will parse the multi-condition join query SQL statement to obtain the main table name and the secondary table name in the join query. Query the query frequency table with the secondary table name as the condition to judge the classification of the secondary table. At this time, three different situations may be obtained: (1) The secondary table is a hot table and there is a fusion table. The program will transform the original SQL statement into a single-table conditional query on the fusion table. The pre-created index value will improve the filtering speed of this query during the multi-dimensional query process. The physical storage format with the index column as the partition key will improve the reading speed of this query during the query process. (2) The secondary table is a hot table, but there is no fusion table. The system will first execute the original query statement, and then perform the preset fusion operation on the hot table and the main table after returning the query result; (3) The secondary table is a cold table. Execute the original query statement. The pre-created multi-dimensional index will improve the filtering speed of this query during the query process. The physical storage format with the index column as the partition key will improve the reading speed of this query during the query process.

[0064] Furthermore, based on the above method, an embodiment of the present invention also provides a multi-condition join query optimization system for medical data based on a data lake, including: a data partitioning module, an index building module, a data fusion module, and a join query module, where,

[0065] The data partitioning module is used to partition the medical data tables in the medical database into a main table and a secondary table, and create a query frequency table for recording the query frequencies of the secondary tables. The medical database includes several medical data tables. The query frequency table includes a secondary table name field and a corresponding query frequency field. The main table is used to record medical entities, and the secondary table is used to record medical entity-related data. There is a one-to-many relationship between the data of the secondary table and the main table;

[0066] The index building module is used to build a high-dimensional index for the secondary table according to the medical query behavior characteristics, and upload the secondary table to the data lake for storage according to the index value;

[0067] The data fusion module is used to divide the secondary tables into hot tables and cold tables according to the medical data query frequencies, fuse the hot tables with the main table, generate a fusion table and a fusion table name based on the main table name and the hot table name, and upload the fusion table to the data lake based on the fusion table name. The hot table is a secondary table with a medical data query frequency greater than the query threshold, and the cold table is a secondary table with a medical data query frequency less than or equal to the query threshold;

[0068] The join query module is used to extract the target main table name and the target secondary table name in the multi-condition join query, judge the target secondary table type according to the query frequency of the target secondary table name, and execute the preset corresponding query strategy according to the target secondary table type to obtain the multi-condition join query medical data.

[0069] As Figure 2 shown, the above-mentioned modules can be implemented through a data lake layer, an index construction area, a query decision area, and a multi-table join query area. Among them, the data lake layer uses APACHE spark to load structured medical data and uniformly converts it into a Dataframe data structure, which can be accessed through the rich APIs provided by the SparkSQL platform. Unstructured data can be converted into feature vectors through a deep learning model and then converted into a DataFrame data structure in the same way [kaige]. In the index construction layer, a high-dimensional index is established for these uniformly structured data and they are written into the Hudi table for persistent storage. In the query decision layer, the connection query method is determined according to the connection query frequency of the data tables: for the fusion table, the data tables for the connection query are fused into one data table for subsequent queries; for the data table connection, the optimized original connection query scheme is used.

[0070] In the process of writing medical data into the data lake in the solution of this case, the main table and the secondary table are distinguished in advance. For the secondary table, a high-dimensional index is created for multiple selected data columns, and the physical distribution of the data is determined according to the index, which can not only optimize the query efficiency of the cold table, but also further improve the query efficiency of the hot table. And by distinguishing between the hot table and the cold table, it is decided whether more storage space needs to be occupied to ensure the query efficiency of the hot table, which effectively makes the space overhead more reasonable.

[0071] Unless otherwise specifically stated, the relative steps, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0072] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0073] The units and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.

[0074] Those of ordinary skill in the art can understand that all or part of the steps in the above method can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as: read-only memory, magnetic disk or optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software function module. The present invention is not limited to any specific form of combination of hardware and software.

[0075] Finally, it should be noted that: the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be determined by the protection scope of the claims.

Claims

1. A medical data multi-condition join query optimization method based on a data lake, characterized in that: Include: In a medical database, a medical data table is divided into a main table and a secondary table, and a query frequency table for recording the query frequency of the secondary table is created, wherein the medical database comprises a plurality of medical data tables, the query frequency table comprises a secondary table name field and a corresponding query frequency field, the main table is used to record medical entities, the secondary table is used to record medical entity-related data, and there is a many-to-one relationship between the data of the secondary table and the main table; Create a high-dimensional index for the secondary table based on the characteristics of medical query behavior, and upload the secondary table to the data lake for storage based on the index value; According to the medical data query frequency, the secondary table is divided into a hot table and a cold table, the hot table is merged with the primary table, and a fusion table and a fusion table name are generated according to the primary table name and the hot table name, and the fusion table is transmitted to the data lake based on the fusion table name. The hot table is a secondary table whose medical data query frequency is greater than the query threshold, and the cold table is a secondary table whose medical data query frequency is less than or equal to the query threshold; Extract the target primary table name and target secondary table name in the multi-condition connection query, determine the target secondary table type based on the query frequency of the target secondary table name, execute the preset corresponding query strategy according to the target secondary table type and obtain the multi-condition connection query medical data.

2. The medical data multi-conditional join query optimization method based on data lake according to claim 1 is characterized in that: The medical entities in the main table are stored in a structured form, and each piece of data corresponds to a medical entity; the medical entity-related data in the secondary table are stored in a structured form or an unstructured form, and the secondary table uses historical query medical data or relevant personnel experience data as connection query conditions.

3. The medical data multi-conditional join query optimization method based on data lake according to claim 1 is characterized in that: A high-dimensional index is created for the secondary table based on the medical query behavior characteristics, including: A secondary table is randomly selected in the medical database as a common secondary table, and the query condition column name that has been recorded in the common secondary table is read; Create a high-dimensional index of the secondary table based on the query condition column name, reduce the values ​​of multiple columns to a single value through the dimensionality reduction method, and add the index value obtained after dimensionality reduction to the secondary table.

4. The medical data multi-conditional join query optimization method based on data lake according to claim 1 is characterized in that: Upload the secondary table to the data lake for storage based on the index value, including: Read the secondary table data containing the index value, and sort the secondary table data in ascending order according to the index value size; The secondary table data sorted in ascending order is uploaded to the data lake with the primary key as the record key and the index column as the partition key.

5. The medical data multi-conditional join query optimization method based on data lake according to claim 1 is characterized in that: The secondary tables are divided into hot tables and cold tables according to the frequency of medical data queries, including: Set query thresholds based on historical query records or relevant personnel experience data; Use the name of the secondary table in the current query as the query condition, and obtain the query frequency value of the corresponding query frequency field in the query frequency table according to the query condition; If the query frequency value is greater than the query threshold, the secondary table in the current query is identified as a hot table, otherwise, the secondary table in the current query is identified as a cold table.

6. The medical data multi-conditional join query optimization method based on data lake according to claim 1 is characterized in that: Merge the hot table with the main table, including: Read primary table data and secondary table data; The primary table data and the secondary table data are merged into a data table structure. After the fusion, the primary key of the primary table is used as the record key and the index column of the secondary table is used as the partition. The merged data table structure is entered into the data lake. The table name of the data table structure is set according to the table names of the primary table and the secondary table.

7. The medical data multi-conditional join query optimization method based on data lake according to claim 1 is characterized in that: Execute the preset corresponding query strategy according to the target secondary table type, including: If the target secondary table is a hot table and there is a fusion table, the query strategy is set to: convert the multi-condition join query into a single-table condition query on the fusion table, and execute the single-table condition query in the fusion table; If the target secondary table is a hot table but there is no fusion table, the query strategy is set to: perform a multi-condition join query, obtain the corresponding primary table data, and fuse the hot table with the corresponding primary table data; If the target secondary table is a cold table, the query strategy is set to: perform a multi-condition join query in the medical database.

8. A medical data multi-condition connection query optimization system based on a data lake, characterized in that: It includes: data partition module, index building module, data fusion module and connection query module, among which, A data partitioning module, used to partition a medical data table into a primary table and a secondary table in a medical database, and to create a query frequency table for recording the query frequency of the secondary table, wherein the medical database comprises a plurality of medical data tables, the query frequency table comprises a secondary table name field and a corresponding query frequency field, the primary table is used to record medical entities, the secondary table is used to record medical entity-related data, and there is a many-to-one relationship between the data of the secondary table and the primary table; The index building module is used to build a high-dimensional index for the secondary table based on the medical query behavior characteristics, and upload the secondary table to the data lake for storage according to the index value; A data fusion module is used to divide the secondary table into a hot table and a cold table according to the medical data query frequency, fuse the hot table with the primary table, and generate a fused table and a fused table name according to the primary table name and the hot table name, and transmit the fused table to the data lake based on the fused table name. The hot table is a secondary table whose medical data query frequency is greater than the query threshold, and the cold table is a secondary table whose medical data query frequency is less than or equal to the query threshold. The connection query module is used to extract the target main table name and target secondary table name in the multi-condition connection query, determine the target secondary table type according to the query frequency of the target secondary table name, execute the preset corresponding query strategy according to the target secondary table type and obtain the multi-condition connection query medical data.

9. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 can be implemented.