Data analysis method and system based on machine learning

Through machine learning, it integrates multi-source heterogeneous data, partitions and stores and analyzes user query requests, solves the problem of hospital data silos, improves data quality and query efficiency, and supports hospital management decisions.

CN120763259AActive Publication Date: 2025-10-10WUHAN YUANQI TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511247564.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-10-10
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

Hospital data is scattered across multiple systems, forming 'data silos'. This leads to difficulties in cross-departmental collaboration, poor accuracy and timeliness of data reporting, difficulty in database maintenance, and data storage methods that are difficult to meet the needs of fast query and complex analysis.

Method used

It adopts a data analysis method based on machine learning to acquire and fuse multi-source heterogeneous data, processes structured and unstructured data through API interfaces and natural language processing technology, performs weighted fusion based on data source weight coefficients, stores data in partitions and adjusts the partition granularity based on historical query frequency, parses user query requests, selects appropriate analysis models for analysis, and generates reports interactively through a graphical user interface.

Benefits of technology

Break down 'data silos', improve data quality and query efficiency, reduce database maintenance difficulty, enhance analysis efficiency, reduce human errors, provide in-depth insights, and support hospital management decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763259A_ABST
    Figure CN120763259A_ABST
Patent Text Reader

Abstract

The invention provides a data analysis method and system based on machine learning, and belongs to the technical field of data processing and analysis, and the method comprises the steps: obtaining, fusing and processing a multi-source heterogeneous data source of a target hospital; processing the multi-source heterogeneous data to obtain a target data set; based on a preset data model and a partition rule, storing the target data set to the plurality of storage partitions to obtain a plurality of data subsets; analyzing a filtering condition in the user query request, and matching a target data subset from the plurality of data subsets; according to analysis target parameters in the query request and partition attributes of the target data subset, selecting an analysis model from a preset model library, analyzing the target data subset and outputting an analysis result; and generating a corresponding report based on an analysis result, and interacting with the user through a graphical user interface. The hospital management level can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing and analysis, and in particular to a data analysis method and system based on machine learning. Background Art

[0002] The process of modern hospital informationization presents multiple challenges in data management. Hospital data is often scattered across multiple systems, such as HIS (Hospital Information System) and ERP (Enterprise Resource Planning System), creating "data silos." This not only hinders cross-departmental collaboration but also prevents the accounting department from timely access to patient billing information, impacting the accuracy and timeliness of data reporting. Furthermore, with the surge in hospital data volume, storage methods are unable to meet the demands of rapid query and complex analysis, making daily database maintenance difficult.

[0003] Therefore, there is an urgent need for a data analysis method and system based on machine learning to improve the level of hospital management. Summary of the Invention

[0004] In order to solve the above technical problems, the present application provides a data analysis method and system based on machine learning.

[0005] A first aspect of the embodiments of the present application provides a data analysis method based on machine learning, comprising:

[0006] Acquire and integrate the target hospital's multi-source heterogeneous data sources, including medical business system data, medical insurance settlement data, and consumables supply chain data;

[0007] Processing the multi-source heterogeneous data to obtain a target data set;

[0008] Based on a preset data model and partitioning rules, the target data set is stored in multiple storage partitions to obtain multiple data subsets; the partitioning rules support data division based on at least one of the following dimensions: time range, business department, and data type, and the partition granularity is adjusted based on historical query frequency;

[0009] Parsing the filter conditions in the user query request and matching the target data subset from the multiple data subsets;

[0010] Select an analysis model from a preset model library based on the analysis target parameters in the query request and the partition attributes of the target data subset, analyze the target data subset, and output the analysis results;

[0011] A corresponding report is generated based on the analysis results, and interacted with the user through a graphical user interface.

[0012] Optionally, the acquiring of multi-source heterogeneous data of the target hospital includes:

[0013] Obtain structured and unstructured data from the target hospital's internal systems and external data platforms through API interfaces;

[0014] Use natural language processing models to perform entity recognition and relationship extraction on unstructured data to generate structured processing data;

[0015] Performing data mode alignment and format unification processing on the structured data and the structured processed data;

[0016] The aligned data are weightedly fused based on the weight coefficients of each data source of the target hospital to generate multi-source heterogeneous data.

[0017] This application performs weighted fusion based on data source weight coefficients, which can reasonably fuse data from different data sources according to the importance and reliability of the data, thereby improving the quality of the fused data.

[0018] Optionally, the method for determining the weight coefficient of each data source includes:

[0019] Obtain evaluation index data for each data source of the target hospital, including: data update frequency index, historical accuracy index, and data provider reputation rating index;

[0020] The weight coefficient of each data source is calculated based on the evaluation index data and a preset first weight calculation formula.

[0021] This application calculates the weight coefficient based on evaluation index data such as data update frequency, historical accuracy and data provider credibility rating, which can objectively and comprehensively reflect the importance and reliability of each data source and provide a scientific basis for weighted fusion data.

[0022] Optionally, after calculating the weight coefficient of each data source according to the evaluation index data and a preset first weight calculation formula, the method further includes:

[0023] In response to the data update frequency indicator being less than a first update frequency threshold, reducing the weight coefficient of the data source by a first step;

[0024] In response to the data update frequency index being less than the second update frequency threshold and greater than or equal to the first update frequency threshold, the weight coefficient of the data source is reduced by a second step size; wherein the first update frequency threshold is less than the second update frequency threshold, and the first step size is greater than the second step size.

[0025] The application dynamically adjusts the weight coefficient of the data source according to the comparison result of the data update frequency index and different threshold values, so that the weight coefficient can be optimized according to the real-time situation of the data, and the accuracy and reliability of data fusion are further improved.

[0026] Optionally, the target data set is stored into a plurality of storage partitions based on the preset data model and partition rules, including:

[0027] The target data set is converted into a structured data set according to a preset data model;

[0028] The time partition granularity is determined based on the historical query frequency, and the structured data set is divided into a primary storage partition according to a time range dimension;

[0029] The data records in the primary storage partition are divided into secondary sub-partitions according to a business department dimension;

[0030] The data records in the secondary sub-partitions are divided into tertiary sub-partitions according to a data type dimension;

[0031] A partition metadata index table is established to record the time range, business department code and data type code of each partition;

[0032] The data subsets after the tertiary division are distributed and stored into corresponding physical storage nodes to generate the plurality of data subsets.

[0033] The application converts the target data set into a structured data set and performs tertiary partition division, establishes a partition metadata index table, realizes structured storage and fine management of data, facilitates quick positioning and query of data, and improves the efficiency of data access.

[0034] Optionally, the time partition granularity is determined based on the historical query frequency, including:

[0035] The query frequency of different time intervals in the historical query request is counted;

[0036] According to the mapping relationship between the query frequency and the time interval, a daily, monthly or annual partition granularity is selected.

[0037] Optionally, according to the mapping relationship between the query frequency and the time interval, a daily, monthly or annual partition granularity is selected, including:

[0038] When the time interval meets the first time interval condition and the corresponding query frequency is greater than the first query frequency threshold, a daily partition granularity is adopted;

[0039] When the time interval meets the second time interval condition and the corresponding query frequency is greater than the second query frequency threshold, a monthly partition granularity is adopted;

[0040] When the time interval meets the third time interval condition and the corresponding query frequency is greater than the third query frequency threshold, the grade partition granularity is adopted.

[0041] This application selects daily, monthly, or yearly partition granularity by clarifying different time interval conditions and query frequency thresholds, making the selection of partition granularity more accurate and scientific, and further improving the efficiency of data storage and query.

[0042] Optionally, parsing the filter condition in the user query request and matching the target data subset from the multiple data subsets includes:

[0043] Parse the filter conditions in the user's query request to obtain the query time interval, query department set and query data type code;

[0044] Query the partition metadata index table, filter the partitions that meet the filter conditions, and obtain the filtered partitions;

[0045] The screening conditions include: the partition time range overlaps with the query time interval, the partition business department code belongs to the query department set, and the partition data type code matches the query data type code;

[0046] The loaded data is determined based on the overlap between the time range of the filter partition and the query time interval, and the loaded data is merged into a target data subset.

[0047] This application parses the filtering conditions of user query requests, filters partitions that meet the conditions by querying the partition metadata index table, and determines the loaded data based on the overlap. It can accurately match the target data subset from multiple data subsets, reducing the data loading volume and query time.

[0048] Optionally, determining to load data based on the overlap between the time range of the filtered partition and the query time interval includes:

[0049] When the degree of overlap is less than a first overlap threshold, skipping the partition;

[0050] When the overlap is greater than or equal to a first overlap threshold and less than a second overlap threshold, data whose query time range overlaps with the partition time range is loaded;

[0051] When the degree of overlap is greater than or equal to a second overlap threshold, the data of the partition is loaded.

[0052] This application formulates different data loading strategies according to the overlap between the filter partition and the query time interval, which can load data according to the actual situation, avoid loading unnecessary data, and improve data processing efficiency and resource utilization.

[0053] A second aspect of the embodiments of the present application provides a data analysis system based on machine learning, comprising:

[0054] The data acquisition module is used to acquire and integrate the target hospital's multi-source heterogeneous data sources, including medical business system data, medical insurance settlement data, and consumables supply chain data;

[0055] A data processing module, configured to process the multi-source heterogeneous data to obtain a target data set;

[0056] A data partitioning module is configured to store the target data set into multiple storage partitions based on a preset data model and partitioning rules to obtain multiple data subsets; the partitioning rules support partitioning data by at least one of the following dimensions: time range, business department, and data type, and the partitioning granularity is adjusted based on historical query frequency;

[0057] A data matching module, configured to parse the filter conditions in the user query request and match a target data subset from the multiple data subsets;

[0058] A model analysis module is used to select an analysis model from a preset model library based on the analysis target parameters in the query request and the partition attributes of the target data subset, analyze the target data subset and output the results;

[0059] The data interaction module is used to generate a corresponding report based on the analysis results and interact with the user through a graphical user interface.

[0060] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the above-mentioned data analysis method based on machine learning when executing the computer program.

[0061] In a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned machine learning-based data analysis system are implemented.

[0062] The beneficial effects of a data analysis method and system based on machine learning provided by the embodiment of the present application are as follows: the present application breaks the "data island" by acquiring and fusing multi-source heterogeneous data, and processes the data at the same time to improve data quality; secondly, the multi-source heterogeneous data is partitioned and stored according to preset rules, and the partition granularity is adjusted according to the historical query frequency, which can effectively improve data query efficiency and reduce the difficulty of database maintenance; different data loading strategies are formulated according to the overlap between the screening partition and the query time interval, which can load data according to the actual situation, avoid loading unnecessary data, and improve data processing efficiency and resource utilization. Secondly, according to the user query request, the data subset is intelligently matched and the analysis model is selected, which greatly improves the analysis efficiency and reduces human errors compared with the traditional manual code generation report; wherein, the present application selects candidate models by parsing the user analysis task type and data characteristics, and then optimizes the candidate models through the hyperparameter optimization algorithm, and uses the optimized model as the final analysis model, wherein the hyperparameter optimization algorithm uses the particle swarm optimization algorithm to search for the optimal hyperparameters, which can converge to the global optimal solution faster and reduce computing resource consumption. The weight of individual experience and social collaboration is adjusted according to the task type to adapt to the optimization needs of different data characteristics. Dedicated fitness functions are designed for different analysis tasks to ensure that optimization objectives align with business requirements. Dual control of the number of iterations and the fitness change threshold prevents over-optimization and balances model performance with computational cost. The number of particles is dynamically adjusted based on data volume and computing resources, ensuring optimization results while avoiding resource waste. Finally, reports are generated based on the analysis results and interacted with through a graphical user interface. This not only visualizes the data but also provides deep insights, helping hospital management quickly obtain key information, providing strong support for hospital decision-making, and improving hospital management. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 A flowchart of a data analysis method based on machine learning provided in one embodiment of the present application;

[0064] Figure 2 A flowchart of a data analysis method based on machine learning provided in one embodiment of the present application;

[0065] Figure 3 A structural block diagram of a data analysis system based on machine learning provided in one embodiment of the present application;

[0066] Figure 4 A schematic block diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0067] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0068] To make the purpose, technical solutions and advantages of this application clearer, Figure 1-3 The following description will be given through specific examples.

[0069] Please refer to Figure 1 , Figure 1 A flow chart of a data analysis method based on machine learning provided in one embodiment of the present application, the method comprising:

[0070] S101: Acquire and integrate the target hospital's multi-source heterogeneous data sources, including medical business system data, medical insurance settlement data, and consumables supply chain data;

[0071] In this example, ETL tools are used to collect data from the target hospital's medical business system, medical insurance settlement system, and consumables supply chain system through API interfaces, direct database connections, and file transfers. Natural language processing techniques are used to extract key information from the unstructured medical record expense descriptions in the medical business system data. Data cleaning is used to remove duplicate and erroneous data. Data mapping and conversion techniques are used to unify the different formats of medical insurance settlement data and consumables supply chain data into a standard format. Data fusion is then performed based on data lineage and metadata information to generate the fused data, i.e., multi-source heterogeneous data.

[0072] S102: Process multi-source heterogeneous data to obtain a target data set;

[0073] In this embodiment, the data volume of the fused multi-source heterogeneous data is reduced through data sampling, and feature engineering techniques, including but not limited to feature selection, feature extraction, and feature construction, are used to screen out features that are strongly relevant to data analysis; the isolation forest algorithm is used to identify and process outliers in the data, and finally the target data set is obtained.

[0074] S103: Based on a preset data model and partitioning rules, the target data set is stored in multiple storage partitions to obtain multiple data subsets; the partitioning rules support dividing data based on at least one of time range, business department, and data type, and the partition granularity is adjusted based on historical query frequency;

[0075] In this embodiment, based on a preset data model, the target data set is stored in multiple storage partitions according to at least one dimension of time range (for example, divided by day, week, month, quarter, year), business department (different clinical departments, administrative departments, etc.), and data type (income data, expenditure data, cost data, etc.), using partitioning techniques such as hash partitioning, range partitioning, and list partitioning to obtain multiple data subsets; at the same time, a historical query frequency analysis module is established to analyze the user's query operations in the past period of time, count the query frequency and response time under different dimensions and different granularities, and adjust the partition granularity. If the frequency of weekly query data is high within a certain time period, the granularity of the weekly partition is increased accordingly.

[0076] In this embodiment, a distributed storage network is constructed in combination with blockchain technology to generate a unique hash value for each data subset, thereby realizing encrypted storage and version traceability of data. At the same time, smart contracts are introduced to ensure dynamic management of data access rights; a spatiotemporal index structure is established to support rapid location and query of data.

[0077] S104: parsing the filter conditions in the user query request and matching the target data subset from multiple data subsets;

[0078] In this embodiment, regular expressions, semantic parsing and other technologies can be used to parse the filter conditions in the user query request and build a query condition tree; traverse the partition attribute information of multiple data subsets, and through the condition matching algorithm, quickly locate and match the target data subset from multiple data subsets.

[0079] S105: Select an analysis model from a preset model library based on the analysis target parameters in the query request and the partition attributes of the target data subset, analyze the target data subset, and output the analysis results;

[0080] In this embodiment, the preset model library includes a variety of basic machine learning models, such as long short-term memory network models for trend prediction; decision tree models and random forest models for classification; clustering algorithm models for cluster analysis; difference detection models for data comparison, etc.

[0081] In this embodiment, an analysis model corresponding to the analysis target parameters in the query request is selected. An ensemble learning algorithm is then used to comprehensively analyze the target data subset. Model fusion is then used to achieve various tasks, including prediction, classification, and association analysis. For prediction tasks, an ensemble learning approach is employed, training a long-short-term memory network model as the first-layer model. A linear regression model is then used as the second-layer model, fusing the output of the first-layer model to achieve more accurate data trend prediction. For classification tasks, a voting mechanism based on a random forest model is used to determine the classification result. For association analysis, a mining algorithm is used to discover association rules between different data items.

[0082] Secondly, this embodiment can also use a multi-objective optimization framework to establish a mapping relationship between user needs and objective function parameters, and dynamically adjust the objective function of the analysis model according to user needs. In addition, when the user pursues computational efficiency, the model with low computational complexity is given priority, and the objective function is adjusted to balance accuracy and computational efficiency. Through optimization algorithms such as particle swarm optimization algorithm and genetic algorithm, the parameters of the analysis model are adjusted to optimize accuracy, interpretability and computational efficiency at the same time. In addition, this embodiment can also use causal inference technology to construct a data causal relationship network based on a subset of target data, and combine causal inference methods such as propensity score matching and instrumental variable method to explore the causal relationship between data and provide users with analysis results with more decision-making value.

[0083] S106: Generate a corresponding report based on the analysis results, and interact with the user through a graphical user interface.

[0084] In this embodiment, based on the analysis results, template engine technology is used in combination with user preference settings to generate a corresponding report, where the user preference settings include: report format and display focus, etc.; the report content includes but is not limited to text descriptions, statistical charts and data tables, etc., where statistical charts include: bar charts, line charts and pie charts, etc.; data visualization interaction functions are provided to users through a graphical user interface, where visualization functions include: data filtering, chart zooming, drilling to view detailed data, etc.; users are also supported to set parameters through the interface, re-initiate query requests or adjust analysis models to achieve in-depth interaction with users.

[0085] From the above, we can conclude that this application breaks the "data island" by acquiring and integrating multi-source heterogeneous data, and at the same time processes the data to improve data quality; partitions and stores data according to preset rules, and adjusts the partition granularity according to the historical query frequency, which can effectively improve data query efficiency and reduce the difficulty of database maintenance. Secondly, it intelligently matches data subsets and selects analysis models based on user query requests. Compared with the traditional manual code writing to generate reports, it greatly improves analysis efficiency and reduces human errors; generates reports based on analysis results and interacts through a graphical user interface, which not only realizes the visualization of data, but also provides deep insights, helping hospital management to quickly obtain key information, providing strong support for hospital decision-making, and improving the level of hospital management.

[0086] In one embodiment of the present application, the data analysis method based on machine learning further includes:

[0087] A real-time anomaly detection module based on autoencoders and Transformers is built to continuously monitor data trends. When abnormal fluctuations are detected, an early warning mechanism is automatically triggered, and the root cause of the anomaly is located through the attention mechanism.

[0088] Build a financial digital twin model of the target hospital, map the financial status of the target hospital, and help users conduct risk assessment and strategic planning by simulating different business scenarios.

[0089] In one embodiment of the present application, obtaining multi-source heterogeneous data of a target hospital includes:

[0090] Obtain structured and unstructured data from the target hospital's internal systems and external data platforms through API interfaces;

[0091] Use natural language processing models to perform entity recognition and relationship extraction on unstructured data to generate structured processing data;

[0092] Perform data schema alignment and format unification on structured data and structured processed data;

[0093] The aligned data are weightedly fused based on the weight coefficients of each data source of the target hospital to generate multi-source heterogeneous data.

[0094] In this embodiment, data from the target hospital's internal systems and external data platforms are obtained through various types of API interfaces. For internal systems such as medical business systems and medical insurance settlement systems, RESTful APIs or SOAP APIs are used to obtain structured data such as patient medical expenses details and medical insurance reimbursement records in a regular or real-time manner in accordance with the data transmission protocol agreed upon by the system. For external data platforms, such as government fiscal policy release platforms and industry data statistics platforms, OAuth-authenticated API interfaces are used. After obtaining permissions, structured and unstructured data such as macro-fiscal policy documents and industry financial indicators are pulled. At the same time, an API interface monitoring module is set up to monitor the interface call status, response time, and data return integrity in real time. When an interface exception occurs, the backup interface is automatically switched or a retry operation is performed, and an exception log is recorded.

[0095] In this embodiment, a natural language processing model based on deep learning is used to process unstructured data. For medical record expense description documents in the medical business system, word segmentation technology is first used to segment the text into word sequences, and then input into the BERT pre-trained language model for feature extraction to capture the contextual semantic information of the words. The features are then bidirectionally learned through the long and short-term memory network. Finally, the CRF layer is used for entity recognition and relationship extraction to extract entities such as drug names, inspection items, and expense amounts, as well as the expense association relationships between them, to generate structured processing data. For unstructured texts such as government fiscal policy documents, named entity recognition technology is used to extract key entities such as subsidy objects and subsidy standards involved in the policy, and relationship extraction technology is used to determine the logical relationship between policy clauses to form structured data.

[0096] In this embodiment, a data model mapping table is established to analyze the field meanings, data types, and semantic relationships between structured data and structured processed data. For data with the same meaning but different field names, such as "patient ID" in the internal system and "medical record number" in the external platform, unified naming is performed through the mapping table. For inconsistent data types, such as date fields stored in "YYYY-MM-DD" and "DD / MM / YYYY" formats in different systems, a data conversion function is used to unify them into a standard date format. This embodiment uses data standardization technology to normalize or standardize numerical data so that the data is in the same numerical range. For categorical data, one-hot encoding or label encoding is used for conversion.

[0097] In this embodiment, the weight coefficients of each data source of the target hospital are determined through a combination of subjective and objective methods. Subjectively, experts and business personnel are organized to score the data sources based on their importance in data analysis. Objectively, the information entropy of each data source is calculated using the entropy weight method. The lower the information entropy, the greater the amount of information provided by the data source and the higher the weight. Combining the subjective and objective weights, the hierarchical analysis method is used to determine the final weight coefficients of each data source. Based on the determined weight coefficients, the aligned data is weighted and fused. For data with the same attributes, such as medical income data from different data sources, a weighted average is calculated according to the weight ratio. For data with complementary attributes, such as detailed expense data from an internal system and industry cost reference data from an external platform, they are directly integrated to generate multi-source heterogeneous data. During the data fusion process, a data conflict monitoring mechanism is established. When data from different data sources conflict, the real-time data from the internal system is prioritized according to preset priority rules, and the conflict resolution process is recorded.

[0098] In one possible embodiment, the formula for weighted fusion of the aligned data based on the weight coefficients of each data source of the target hospital is:

[0099]

[0100] in, For multi-source heterogeneous data, For the The weight coefficient of each data source, For the After the data sources are aligned, is element-wise multiplication; is the feature confidence mask matrix, which is used to adjust the fusion weight of each eigenvalue; when the confidence ≥ θ (confidence threshold), ; When confidence < θ (confidence threshold), .

[0101] In one embodiment of the present application, a method for determining the weight coefficient of each data source includes:

[0102] Obtain evaluation indicator data for each data source of the target hospital, including: data update frequency index, historical accuracy index, and data provider credibility index;

[0103] The weight coefficient of each data source is calculated based on the evaluation index data and the preset first weight calculation formula.

[0104] In this example, data update logs are obtained from each data source system through an API interface. The timestamp of each data update is recorded, the time interval between two consecutive updates is calculated, and the average update frequency within a certain period of time is calculated. Data sources that update in real time are given the highest update frequency score. Data sources that update on a daily, weekly, or monthly basis are scored accordingly.

[0105] In this embodiment, a certain percentage of data samples are regularly extracted from each data source and compared with a manually verified standard dataset. For numerical data, absolute and relative errors are calculated; for categorical data, metrics such as precision, recall, and F1 value are calculated. Based on the comparison results, the average accuracy of each data source over a historical period is calculated as a historical accuracy indicator.

[0106] We collect relevant information about the data providers' qualifications, including their industry certifications, years of operation, and historical data quality feedback. We then use a reputation assessment model to quantify the reputation of data providers. Data provided by government departments and authoritative industry organizations will be assigned a higher reputation rating. For commercial data providers, we assess their reputation based on a comprehensive assessment of factors such as service quality and historical data accuracy.

[0107] The data update frequency indicator is converted into a value between 0 and 1. The higher the update frequency, the closer the value is to 1. For example, a data source with real-time updates is scored as 1, a data source with daily updates is scored as 0.8, a data source with weekly updates is scored as 0.5, and a data source with monthly updates is scored as 0.2.

[0108] For the historical accuracy indicator, the calculated accuracy value is directly used as the score. For example, if the historical accuracy of a data source is 95%, its score is 0.95.

[0109] The data provider's reputation rating indicator is mapped to a value between 0 and 1 based on the rating results of the reputation assessment model. For example, a data source with an "excellent" reputation rating is scored 0.9, a "good" rating is scored 0.7, a "fair" rating is scored 0.5, and a "poor" rating is scored 0.3.

[0110] This embodiment also monitors changes in the evaluation indicator data for each data source in real time. When a data source's data update frequency decreases significantly, its historical accuracy fluctuates, or the data provider's credibility changes, the recalculation of the weight coefficient is triggered. Weight parameters are adjusted regularly (e.g., quarterly) based on data usage feedback and changing business needs. For example, as hospital informatization progresses, the accuracy and real-time performance of internal data sources continue to improve, and the weight coefficient of internal data sources can be increased.

[0111] In a possible embodiment, the first weight calculation formula is:

[0112]

[0113] in, is the first adjustment coefficient, the initial value of which is set by experts and then automatically optimized based on model feedback; This is the accuracy amplification factor to ensure that high-accuracy data receives higher weight; The second adjustment coefficient is initially set by experts and is automatically optimized based on model feedback. For the Update frequency indicators of each data source; For the Historical accuracy metrics for each data source; For the Reliability rating indicators for each data source; The minimum credit rating benchmark value; The third adjustment coefficient is initially set by experts and is automatically optimized based on model feedback. Prevent high-frequency data from monopolizing weights.

[0114] In one embodiment of the present application, after calculating the weight coefficient of each data source according to the evaluation index data and the preset first weight calculation formula, the method further includes:

[0115] In response to the data update frequency indicator being less than a first update frequency threshold, reducing the weight coefficient of the data source by a first step;

[0116] In response to the data update frequency index being less than the second update frequency threshold and greater than or equal to the first update frequency threshold, the weight coefficient of the data source is reduced by a second step size; wherein the first update frequency threshold is less than the second update frequency threshold, and the first step size is greater than the second step size.

[0117] In this embodiment, based on the statistical analysis of the update frequency of historical data, a time series prediction model is established to predict the update frequency trend of each data source in the future; the first update frequency threshold and the second update frequency threshold are determined in combination with the business needs and the timeliness requirements of the data analysis task. For example, for medical insurance settlement data with high real-time requirements, the first update frequency threshold is set to update once every hour, and the second update frequency threshold is set to update once every day. This embodiment can also regularly collect user feedback and business department needs, evaluate the rationality of the current threshold setting, and adjust the threshold size according to the feedback results. When the data analysis task has higher requirements for data real-time performance, the update frequency threshold is lowered accordingly.

[0118] In this embodiment, when the data update frequency indicator of a data source falls below the first update frequency threshold, a rapid weight adjustment mechanism is initiated, reducing the weight coefficient of the data source by the first step length. For example, if the first step length is 0.1 and the original weight coefficient of the data source is 0.6, the adjusted weight coefficient becomes 0.5. A weight adjustment log is also recorded, including information such as the adjustment time, the weight values ​​before and after the adjustment, and the update frequency value that triggered the adjustment.

[0119] When the data update frequency index is less than the second update frequency threshold and greater than or equal to the first update frequency threshold, the slow weight adjustment mechanism is activated to reduce the weight coefficient of the data source by the second step size. For example, if the second step size is 0.05 and the original weight coefficient of the data source is 0.6, the adjusted weight coefficient becomes 0.55.

[0120] In this embodiment, in order to prevent the weight coefficient from being excessively reduced, a lower limit value of the weight coefficient is set. When the adjusted weight coefficient is lower than the lower limit, it is fixed to the lower limit value to ensure that even if the data source is updated very slowly, a certain reference value can still be retained.

[0121] This embodiment also establishes a weight adjustment buffer period, during which changes in the data source's update frequency are continuously monitored. If the data source returns to normal update frequency within the buffer period, its weight coefficient is gradually restored; if the update frequency still does not meet the standard, the adjusted weight coefficient is maintained.

[0122] In the specific implementation of this embodiment, when the weight coefficient of a data source is adjusted due to an update frequency issue, the weight coefficients of other data sources are recalculated to keep the sum of the weight coefficients of all data sources equal to 1. The weight coefficients of other data sources are adjusted using a proportional scaling method to ensure that the adjusted weight distribution is reasonable.

[0123] In the specific implementation of this embodiment, the data fusion quality and analysis result accuracy before and after the weight adjustment are compared. If the analysis result quality is found to be degraded after the adjustment, the update frequency threshold and step size parameters are adjusted in a timely manner to ensure the effectiveness of the weight adjustment strategy.

[0124] In one embodiment of the present application, based on a preset data model and partitioning rules, the target data set is stored in multiple storage partitions, including:

[0125] Convert the target dataset into a structured dataset according to the preset data model;

[0126] Determine the time partition granularity based on historical query frequency, and divide the structured data set into primary storage partitions based on the time range dimension;

[0127] Divide the data records in the primary storage partition into secondary sub-partitions based on the business department dimension;

[0128] According to the data type dimension, the data records in the secondary sub-partition are divided into three-level sub-partitioning;

[0129] A partition metadata index table is established to record the time range, business department code and data type code of each partition;

[0130] The data subsets after three-level division are distributed and stored in corresponding physical storage nodes to generate multiple data subsets.

[0131] In this embodiment, the target data set is structured and converted according to a preset data model. For the nested expense detail data in the medical business system, the multi-layer nested JSON or XML format data is expanded into a two-dimensional table structure through data flattening processing; for the redundant fields existing in the medical insurance settlement data, the data standardization technology is used to remove duplicate data storage, and the data structure is ensured to meet the standard of the preset model. In the conversion process, the data quality checking tool is used to check the field type consistency, primary key uniqueness, and foreign key association accuracy, etc. If the data format is found to be incorrect or incomplete, the data cleaning process is automatically triggered to modify or mark the abnormal data.

[0132] In this embodiment, the time range involved in the user query operation in the past 12 months is counted, and the query frequency proportion of different time granularity (day, week, month, quarter, year) is calculated. For example, if it is found that the frequency of quarterly query reaches 40%, which is significantly higher than other granularities, then the quarter is taken as the initial time partition granularity. Combined with the data growth trend prediction, the exponential smoothing method or linear regression model is used to estimate the future data volume. If it is predicted that the data volume will increase significantly in a certain period of time, the time partition granularity is refined in advance, such as dividing the quarterly partition into monthly partition, to avoid the influence of the large data volume in a single partition on the query performance.

[0133] According to the determined time partition granularity, the structured data set is divided into primary storage partitions according to the time range dimension. A unique time identifier is assigned to each primary partition, such as "2025Q1" and "2025Q2", and a time index is established in the partition to speed up the time range query operation.

[0134] In this embodiment, a business department coding system is established to uniformly code the target hospital's clinical departments (e.g., cardiology, respiratory medicine), medical technology departments (e.g., radiology, laboratory), and administrative departments (e.g., finance, medical services). The coding rules adopt a hierarchical structure, such as "01-cardiology-0101-outpatient" and "01-cardiology-0102-inpatient," to ensure clear classification of department data. In specific implementation, each data record within the primary storage partition is traversed, the business department information fields contained in the data are extracted, and the data is divided into corresponding secondary sub-partitions according to the coding system. For cross-departmental data, such as joint medical project expense data, data is divided according to the primary responsible department, and auxiliary identification fields are added to the data records to indicate the other departments involved. If this embodiment finds that the data volume of certain departments is too high, resulting in reduced query performance, the secondary sub-partitions of these departments are further subdivided, such as by tertiary sub-partitioning according to the different wards or project groups under the department.

[0135] In this example, data type classification standards are defined to categorize data into income (medical income, medical insurance income, and fiscal subsidy income), expenditure (personnel expenditure, consumables expenditure, and equipment procurement expenditure), and cost (direct cost and indirect cost). Each data type is assigned a unique code, such as "01 - Medical Income" or "02 - Personnel Expenditure."

[0136] Data records within the second-level subpartitions are divided into third-level subpartitions based on the data type field. A data type mapping table is established. When data type field values ​​are ambiguous or non-standard, they are mapped to the correct data type partition through semantic analysis and rule matching. Differentiated storage strategies are employed for different data types. For frequently queried revenue data, a columnar storage format is used to improve aggregate query efficiency. For less frequently queried historical cost data, a row-based storage format is used to save storage space.

[0137] In this embodiment, the partition metadata index table is established, including fields such as time range (e.g., start time and end time), business department code, data type code, partition physical storage path, and data record quantity. This embodiment can automatically update statistical information such as the number of data records in the index table by scanning the data in each partition.

[0138] In this embodiment, a composite index can be created for the index table, using the time range, business department code, and data type code as index keys to accelerate the location of the target partition based on user query conditions. Furthermore, distributed hash table (DHT) technology is used to store index table data across multiple nodes, improving index table query and maintenance performance.

[0139] This application can provide a visual management interface for the index table, allowing administrators to manually modify or add partition metadata information, such as adjusting the partition time range, modifying the department code correspondence, etc., and automatically synchronize to each storage node after the operation is completed.

[0140] In this embodiment, a distributed file system or distributed database is used as the storage medium. Data subsets that have undergone three-level partitioning are stored in corresponding physical storage nodes based on the metadata information of each partition. During the storage process, data redundancy technologies, such as replication or erasure coding, are utilized to ensure data reliability. If a storage node fails, data can be retrieved from other replica nodes. The load and storage space utilization of each storage node are dynamically monitored. When a node is overloaded or lacks storage space, a data migration mechanism is automatically triggered to migrate some data subsets to less loaded nodes, achieving balanced distribution of storage resources.

[0141] In one embodiment of the present application, determining the time partition granularity based on the historical query frequency includes:

[0142] Collect statistics on the query frequency of historical query requests in different time intervals;

[0143] Based on the mapping relationship between query frequency and time interval, select the day-level, month-level, or year-level partition granularity.

[0144] In this embodiment, a query log collection system is established to collect query requests initiated by users in real time, and record information such as the query time, the time interval in the query conditions (such as "2025-01-01 to 2025-01-31"), the amount of data involved in the query, and the query response time.

[0145] Clean and standardize historical query logs, remove invalid query records (such as requests with incorrect query parameters), unify the representation format of time intervals, and convert time ranges in different formats (such as "the first quarter of 2025" to "2025-01-01 to 2025-03-31").

[0146] This embodiment extracts key features from the cleaned query logs, where the key features include time interval length, query frequency, query popularity (calculated based on query response time and resource consumption), etc., to form a structured historical query data set.

[0147] In this embodiment, a time interval bucket division rule is established to divide the time axis into buckets of different lengths, such as daily buckets (each bucket represents a day), weekly buckets (each bucket represents a week), monthly buckets (each bucket represents a month), quarterly buckets (each bucket represents a quarter), and annual buckets (each bucket represents a year).

[0148] Count the query frequency within each time bucket and calculate the query frequency distribution at different time granularities (daily, weekly, monthly, quarterly, and annual). For example, count the number of queries with "daily" as the time interval and the number of queries with "monthly" as the time interval in the past year.

[0149] Construct a query frequency-time interval matrix, with the horizontal axis representing time granularity (day, week, month, quarter, year) and the vertical axis representing query frequency. Visualize the mapping between the two using heat maps or scatter plots. Use association rule mining algorithms to discover strong associations between different time granularities and query frequency.

[0150] In this embodiment, we calculate the query frequency weights for each time granularity and establish a partition granularity selection decision tree model. This model takes into account factors such as the query frequency weights for each time granularity, the predicted data volume, and storage cost, and outputs the optimal partition granularity. The decision tree's partitioning rules are based on a cost-benefit analysis, comprehensively considering the balance between improved query performance and increased storage overhead.

[0151] Set a partition granularity selection threshold. When the query frequency weight of a certain time granularity exceeds the threshold, and the average query response time at this time granularity meets the business needs, this time granularity is preferentially selected as the partition granularity. If the weights of multiple time granularities are close, the middle granularity is selected as a compromise. This embodiment tracks the growth of the data volume and the query frequency change trend of each time partition in real time. When it is found that the data volume of a certain time partition exceeds the preset threshold, or the query frequency ratio of a certain time granularity changes significantly, the partition granularity adjustment process is triggered. This embodiment can also predict the query frequency change trend of each time granularity in the future based on the time series prediction model, and adjust the partition granularity in advance to adapt to the change of the query pattern. For example, if it is predicted that the monthly query frequency will increase significantly in the future, some quarterly partitions will be split into monthly partitions.

[0152] In one embodiment of the present application, based on the mapping relationship between query frequency and time interval, the daily, monthly, or grade-level partitioning granularity is selected, including:

[0153] When the time interval meets the first time interval condition and the corresponding query frequency is greater than the first query frequency threshold, the day-level partition granularity is adopted;

[0154] When the time interval meets the second time interval condition and the corresponding query frequency is greater than the second query frequency threshold, the monthly partition granularity is adopted;

[0155] When the time interval meets the third time interval condition and the corresponding query frequency is greater than the third query frequency threshold, the grade partition granularity is adopted.

[0156] In this embodiment, the first time interval condition is that when the length of the time interval is less than or equal to a preset first length threshold, the first time interval condition is determined to be met. For example, the first length threshold is 5 days, the query condition is "2025-01-01 to 2025-01-05", the length of the time interval is 5 days, and the first time interval condition is met.

[0157] The second time interval condition is that when the length of the time interval is greater than the first length threshold and less than or equal to a preset second length threshold, the second time interval condition is determined to be met. For example, the second length threshold is 90 days, the query condition is "2025-01-01 to 2025-03-31", the length of the time interval is 90 days, and the second time interval condition is met.

[0158] The third time interval condition is that when the length of the time interval is greater than the second length threshold, the third time interval condition is determined to be met. For example, the query condition is "2025-01-01 to 2025-12-31", the length of the time interval is 365 days, and the third time interval condition is met.

[0159] In this embodiment, according to the seasonal fluctuation characteristics of hospital data, the time interval threshold can be reduced during the business peak period to increase the proportion of fine-grained partitions; during the business trough period, the threshold can be relaxed to reduce the number of partitions and reduce storage costs.

[0160] In this embodiment, based on the statistical distribution characteristics of historical query data, the query frequency threshold is determined using a box plot or a 3σ principle. For example, the median and interquartile range of all time interval query frequencies are calculated, and the first query frequency threshold is set to the median plus 1.5 times the interquartile range.

[0161] In this embodiment, an association model between the query frequency threshold and the data volume is established. When the data volume at a certain time granularity exceeds a preset capacity threshold, the query frequency threshold of the time granularity is increased accordingly to avoid excessive partitions due to excessive data volume.

[0162] In this embodiment, when a time interval meets multiple time interval conditions at the same time, a priority rule is used to determine the final partition granularity. The priority order is: daily partition granularity > monthly partition granularity > annual partition granularity. For example, if the length of a time interval is 30 days and meets both the first and second time interval conditions, and the corresponding query frequencies exceed their respective thresholds, the daily partition granularity is preferred.

[0163] In this embodiment, the priority rule can also be adjusted; when the storage resources are sufficient, the priority of fine-grained partitions is increased; when the storage resources are tight, the priority of fine-grained partitions is reduced.

[0164] In this embodiment, when the user query involves multiple time granularities, both daily and monthly data are queried simultaneously, and a hybrid partitioning strategy is adopted, i.e., fine-grained partitioning is used for the time range of high-frequency queries, and coarse-grained partitioning is used for the time range of low-frequency queries. For example, daily partitioning is used for the data of the last month, and monthly partitioning is used for historical data. When a significant change in the mapping relationship between query frequency and time interval is detected (for example, the query frequency of a certain time granularity suddenly decreases by 80%), a partition granularity reevaluation process is triggered, and the partitioning strategy is adjusted according to the new mapping relationship.

[0165] In an embodiment of the present application, the filtering conditions in the user query request are parsed, and the target data subset is matched from multiple data subsets, including:

[0166] The filtering conditions in the user query request are parsed to obtain the query time interval, the query department set, and the query data type code;

[0167] The partition metadata index table is queried to filter the partitions that meet the filtering conditions, and the filtered partitions are obtained;

[0168] The filtering conditions include: the partition time range overlaps with the query time interval, the partition business department code belongs to the query department set, and the partition data type code matches the query data type code;

[0169] According to the overlap degree of the time range of the filtered partition and the query time interval, the loaded data is determined, and the loaded data is merged into the target data subset.

[0170] In this embodiment, the natural language processing technology is combined with regular expressions to parse the user query request. For structured query statements, the time interval, department name, and data type keywords are directly extracted; for natural language description query requests, such as "view the medical income data of the cardiology department in the first half of 2025", the named entity recognition model is used to extract the key entities "the first half of 2025", "cardiology department", and "medical income", and then the natural language is converted into structured query conditions through semantic analysis to obtain the accurate query time interval, query department set (the department code set corresponding to "cardiology department"), and query data type code (the code corresponding to "medical income").

[0171] In this embodiment, when the parsed keywords have semantic ambiguity or spelling errors, synonym libraries and similar word matching algorithms are used for correction and supplementation. For example, if the user inputs "cardiology lesson", it is automatically matched to "cardiology department"; if the query data type keyword is incomplete, such as only "income" is input, it is inferred and supplemented to "medical income" according to the context and historical query mode, ensuring the accuracy of the parsed results.

[0172] In this embodiment, a multi-level index structure is established for the partition metadata index table. In addition to the existing composite index of time range, business department code, and data type code, a prefix index for time range and an inverted index for business department code are also established for high-frequency query scenarios. During queries, the prefix index is preferentially used to quickly locate the partitions with a rough time range, and the inverted index is then used to filter out the partitions corresponding to the department. Finally, the composite index is used for precise matching, improving the query efficiency of the index table.

[0173] In this embodiment, an interval tree data structure is used to organize and manage partitioned time ranges for time range screening. Each partition's time range is used as a node in the interval tree. During a query, the interval tree is used to quickly identify partitions that overlap with the query time range. This significantly reduces time complexity compared to traditional linear scanning methods.

[0174] When filtering business department codes, bitwise operations are used to accelerate the matching process when the query department set contains multiple departments. Each business department code is assigned a unique binary bit, and the query department set is converted into a binary mask. Bitwise AND operations are then used to quickly determine whether a partitioned business department code belongs to the query department set, improving screening efficiency.

[0175] In data type code matching, a hierarchical mapping relationship is established for data type codes. When the query data type code is a parent node code, the partitions corresponding to all its child node codes are automatically matched. For example, if the query data type code is "income category," partitions containing all income subtypes such as "medical income" and "medical insurance income" are automatically filtered out, enhancing filtering flexibility.

[0176] In this embodiment, a dynamic data loading strategy is employed based on the degree of overlap between the time range of the filtered partition and the query time interval. When the overlap exceeds a second overlap threshold, the entire partition data is directly loaded. When the overlap is lower, a paged loading or on-demand loading approach is used to load only the data that overlaps with the query time interval, reducing unnecessary data transfer and memory usage.

[0177] In one embodiment of the present application, determining to load data based on the overlap between the time range of the filter partition and the query time interval includes:

[0178] When the overlap is less than the first overlap threshold, the partition is skipped;

[0179] When the overlap is greater than or equal to the first overlap threshold and less than the second overlap threshold, data that overlaps the query time range and the partition time range is loaded;

[0180] When the degree of overlap is greater than or equal to a second overlap threshold, the data of the partition is loaded.

[0181] In this embodiment, based on historical query data, the relationship between data loading efficiency and query response time at different overlap levels is analyzed, and an overlap-performance indicator prediction model is established. With the goal of minimizing query response time and data loading resource consumption, the optimal first and second overlap thresholds are calculated. For example, based on an analysis of the past 1,000 queries, data loading efficiency begins to improve significantly when the overlap reaches 30, and approaches saturation at 70%. Therefore, the first overlap threshold can be preliminarily set to 30%, and the second overlap threshold to 70%. This embodiment also allows for manual adjustment of the overlap thresholds.

[0182] When the amount of partitioned data is large, the first overlap threshold can be increased to reduce unnecessary data loading. When the amount of data is small, the first overlap threshold can be lowered to avoid data incompleteness caused by excessive skipping of partitions. Furthermore, the threshold is dynamically adjusted based on the frequency of data updates. For data that is updated in real time, the threshold is lowered to ensure that the latest data is loaded promptly.

[0183] In this embodiment, when the overlap falls below a first overlap threshold, the partition is skipped and its information is recorded in a "low-correlation partition list." If the partition is skipped multiple times in subsequent queries due to low overlap, the system automatically archives or compresses the partition's data to reduce storage resource usage. Furthermore, the "low-correlation partition list" is reviewed periodically (e.g., monthly). If the partition data changes (e.g., due to data updates), its overlap with the query time interval is reassessed to avoid misjudgments.

[0184] When the overlap is greater than or equal to the first overlap threshold and less than the second overlap threshold, incremental data loading is employed. Using the partition's time index, data segments where the query time range overlaps with the partition's time range are quickly located and only that portion of data is loaded. The loaded data is formatted and preprocessed to meet the input requirements of the subsequent analysis model. To increase loading speed, this embodiment employs multithreaded or distributed loading techniques to load overlapping data from multiple partitions in parallel.

[0185] When the overlap is greater than or equal to the second overlap threshold, all the data of the partition is loaded.

[0186] In this embodiment, a hash check code is generated for the loaded data. After the data is loaded, the integrity of the data during the transmission and loading process is verified by comparing the hash value. If the hash values ​​are inconsistent, data reloading is automatically triggered, and an error log is recorded, including the error partition information, error time, error cause, etc., to facilitate subsequent troubleshooting. For loaded overlapping data, the integrity of its boundary data is checked. For example, when loading data of the overlapping parts of two partitions, ensure that the data at the overlapping boundary is neither loaded repeatedly nor omitted. This embodiment can also select appropriate analysis models and parameters based on the amount of loaded data and the characteristics of the data. For example, when the amount of loaded data is small, a model with lower computational complexity is preferred; when the amount of data is large, a distributed computing framework is used to run the analysis model to ensure that the analysis task is completed efficiently.

[0187] In one embodiment of the present application, processing multi-source heterogeneous data to obtain a target data set includes:

[0188] Detecting and correcting outliers through an isolation forest algorithm to generate first data;

[0189] Identifying and eliminating noise data in the first data by a clustering algorithm;

[0190] Use statistical methods to process missing values ​​in the data after noise removal to generate cleaned data;

[0191] Based on the hospital operation analysis goals, construct feature datasets from clean data;

[0192] Features are selected based on feature importance and correlation analysis to generate the target dataset.

[0193] In this embodiment, numerical data in multi-source heterogeneous data is input into the isolation forest algorithm model. During the model training phase, multiple isolation trees are constructed by randomly selecting sample data, and the number of trees and subsample size are set. Each isolation tree starts from the root node, randomly selects a feature and a split point of the feature, and divides the sample data into left and right child nodes until each leaf node contains only one sample or reaches a preset tree depth. The anomaly score is calculated based on the path length of the sample in the tree. If the anomaly score of a sample is higher than the preset anomaly threshold, it is determined to be an outlier. For outliers, a correction method based on statistical distribution is used to calculate the mean and standard deviation of the feature data, and the outlier is replaced with the nearest normal value within the range of ±3 times the standard deviation of the mean to generate the first data.

[0194] A clustering algorithm is used to identify noise data in the first data. A neighborhood radius is set to the minimum number of samples. Each sample point in the dataset is traversed and the number of samples within the neighborhood radius is calculated. If the number of neighboring samples for a sample point is less than the minimum number of samples and the sample point does not belong to any established clusters, it is considered noise data. The identified noise data is removed from the first data to reduce its interference with subsequent analysis.

[0195] In this embodiment, missing values ​​are handled using statistical methods for the noise-removed data. For missing values ​​in numerical data, if the missing ratio is lower than the first missing range, the mean or median is used for filling. If the missing ratio is within the second missing range, a multiple imputation method is used to construct a regression model based on other relevant features to predict missing values. If the missing ratio is within the third missing range, the feature column is directly deleted. For missing values ​​in categorical data, the mode is used for filling. The third missing range is larger than the second missing range, and the second missing range is larger than the first missing range.

[0196] In this example, based on hospital operational analysis objectives, such as medical revenue forecasting and cost control analysis, relevant fields are extracted from cleaned data to construct a feature dataset. For example, to analyze factors influencing medical revenue, fields such as patient age, department visited, treatment program, medical insurance type, and length of stay are extracted as features. For cost control analysis, fields such as consumable name, purchase quantity, purchase price, department used, and duration of use are extracted. Furthermore, some features are derived to enrich the feature dimensionality, such as calculating new features such as per capita medical expenses and consumable utilization rate.

[0197] The Random Forest algorithm calculates feature importance by randomly partitioning the training and test sets multiple times. The contribution of each feature to the model's accuracy is calculated, with higher contributions indicating greater importance. The Pearson correlation coefficient is used to analyze the correlation between features. A correlation threshold is set to eliminate redundant features with excessively high correlations. Based on the feature importance and correlation analysis results, the most valuable features for the analysis are selected to generate the target dataset.

[0198] In one embodiment of the present application, the feature data set includes: features reflecting the cost-effectiveness ratio of the department, features reflecting the difference between medical insurance payment and actual charges, and features reflecting the correlation between consumables use and income;

[0199] in,

[0200] By comparing the payment details in the medical insurance settlement data with the charge details in the medical business system, the difference characteristics between medical insurance payment and actual charges are constructed;

[0201] By associating consumables usage records in consumables supply chain data with revenue data in the medical business system, we can build a correlation feature between consumables usage and revenue.

[0202] Calculate the cost-benefit ratio characteristics of each business department, including the proportion of personnel costs, equipment depreciation rate and income profit margin.

[0203] In this embodiment, payment details in the medical insurance settlement data and charge details in the medical service system are cleaned and formatted uniformly to ensure consistency in key information such as patient identification and treatment item codes. Using a data matching algorithm, the patient's unique identifier (such as their ID number or medical insurance card number) and the treatment item code serve as matching keys to link medical insurance payment details with corresponding medical charge details.

[0204] For each successfully matched record, we calculate metrics like the difference between the medical insurance payment and the actual billing amount, as well as the discrepancy rate. For multiple treatment records of the same patient over different periods of time, we aggregate and compile them by treatment cycle (e.g., monthly or quarterly), obtaining features like the total difference between medical insurance payments and actual billing within each cycle and the average discrepancy rate.

[0205] Using time series analysis, we analyze the changing trends in the discrepancy between medical insurance payments and actual charges over time, extracting trend and seasonal characteristics to enrich the feature dimensions of the discrepancy. We also establish an anomaly detection mechanism. When the discrepancy rate within a certain time period is greater than the discrepancy rate, the data is marked as abnormal and manually reviewed to ensure the accuracy of the feature data.

[0206] In this embodiment, consumable usage records extracted from the consumable supply chain data include information such as consumable name, quantity used, department used, and time of use. Revenue data extracted from the medical business system includes information such as revenue item, revenue amount, corresponding department, and time. Using a data association algorithm, using department and time as the association dimensions, consumable usage records are associated with revenue data to create a consumable usage-revenue association table.

[0207] Using association rule mining algorithms, we can mine association rules between different combinations of consumables and specific revenue items. For example, we can discover that when a combination of consumables A and B is used, revenue from a certain type of surgery increases significantly. These rules can then be converted into feature representations for subsequent analysis and decision support.

[0208] In this embodiment, personnel cost percentage calculation involves obtaining personnel salary, benefits, training, and other expense data for each business department from the hospital's human resources management system and financial accounting system as personnel costs. Furthermore, total revenue data for each department is obtained. The personnel cost percentage characteristic is calculated using the formula "Personnel cost percentage = personnel cost / total department revenue × 100%." ​​To more accurately reflect the personnel cost structure, personnel types (e.g., doctors, nurses, administrative staff) are further subdivided, and the cost percentages for each type are calculated separately.

[0209] In this embodiment, the equipment depreciation rate is calculated: according to the purchase amount, the expected service life and the used life of the equipment in each department of the hospital fixed asset management system, the linear depreciation method is used to calculate the equipment depreciation amount. The formula is "annual depreciation amount = (purchase amount - expected net salvage value) / expected service life", wherein the expected net salvage value can be set as a fixed proportion according to the type of equipment. The equipment depreciation rate of each department is calculated, and the formula is "equipment depreciation rate = depreciation amount / equipment purchase amount x 100%".

[0210] In this embodiment, the income profit rate is calculated: the net profit and total income data of each department are obtained from the financial data report, and the income profit rate characteristics are calculated by the formula "income profit rate = net profit / department total income x 100%". According to the business characteristics and market environment of the department, the income profit rate is compared and analyzed horizontally and vertically. The horizontal comparison is made with similar departments of other hospitals in the same industry, and the vertical analysis is made on the income profit rate change trend of the department in different time periods. The key factors affecting the income profit rate are found out, and these analysis results are converted into auxiliary features for more comprehensive evaluation of the cost-benefit ratio of the department.

[0211] In an embodiment of the present application, according to the analysis target parameters in the query request and the partition attributes of the target data subset, an analysis model is selected from a preset model library, including:

[0212] The analysis task type parameters in the query request are parsed, including task type, analysis index and business constraint conditions;

[0213] Based on the analysis task type parameters and the key data characteristics of the target data subset, a candidate model set is selected from the preset model library; the key data characteristics include at least one of data volume, feature dimension, time series characteristics and label distribution;

[0214] The performance of the candidate model is evaluated using sample data of the target data subset, and the evaluation index is set according to the task type;

[0215] If any key evaluation index is lower than a preset threshold, the candidate model is optimized by a hyperparameter optimization algorithm, and the optimized model is used as the final analysis model.

[0216] In this embodiment, natural language processing technology is used to perform semantic analysis on the query request to identify the analysis target parameters therein. For structured query statements, the prediction type, analysis index and business constraint conditions are directly extracted; wherein the prediction type includes: income prediction and cost prediction; the analysis index includes: profit rate and growth rate; the business constraint conditions include: time window and department range;

[0217] For queries described in natural language, a named entity recognition model is used to extract key entities and intents and convert them into structured analysis target parameters.

[0218] Query requests are divided into different analysis types based on the combination of analysis target parameters:

[0219] Forecasting: includes time series forecasting, regression forecasting, etc. The corresponding forecasting type parameters are "revenue forecasting", "cost forecasting", etc.

[0220] Classification: including risk level classification, medical insurance payment type classification, etc. The corresponding prediction type parameters are "high-risk identification" and "medical insurance type classification".

[0221] Clustering: including patient group clustering, department performance clustering, etc. The corresponding prediction type parameters are "patient clustering" and "department classification".

[0222] Anomaly detection: This includes medical insurance fraud detection and financial data anomaly identification. The corresponding prediction type parameters are "abnormal transaction detection" and "cost anomaly identification."

[0223] This embodiment screens candidate models by analyzing the user's analysis task type and data characteristics, avoiding manual trial and error and improving model adaptability. The performance of candidate models is evaluated using sample data to ensure the selected model's effectiveness on real-world data and reduce model bias. When model performance falls short of expectations, hyperparameter optimization is automatically triggered to improve model accuracy and generalization capabilities to meet the needs of different business scenarios. Dedicated evaluation metrics and fitness functions are set for different analysis tasks to ensure targeted model optimization.

[0224] In one embodiment of the present application, screening a candidate model set based on analysis task type parameters and key data characteristics includes:

[0225] If the analysis task type parameter is prediction and the key data characteristics include time series characteristics, the time series prediction model in the model library is filtered;

[0226] If the analysis task type parameter is classification and the feature dimension in the key data characteristics is greater than the first preset dimension (50 dimensions), the classification model that supports high-dimensional features in the model library is selected;

[0227] If the analysis task type parameter is clustering and the data volume in the key data characteristic is greater than the first sample quantity threshold (1 million samples), the distributed clustering model in the model library is screened;

[0228] If the analysis task type parameter is anomaly detection and the label distribution in the key data features is unbalanced, the anomaly detection model suitable for unbalanced data in the model library is selected.

[0229] In one embodiment of the present application, the performance evaluation indicators are set according to the task type, including:

[0230] When the task type is prediction, the evaluation indicator is root mean square error or mean absolute percentage error;

[0231] When the task type is classification, the evaluation indicator is F1 score or area under the ROC curve;

[0232] When the task type is clustering, the evaluation indicator is the silhouette coefficient;

[0233] When the task type is anomaly detection, the evaluation indicator is the area under the precision-recall curve or the Matthews correlation coefficient.

[0234] In one embodiment of the present application, the candidate model is optimized by a hyperparameter optimization algorithm, specifically a particle swarm optimization algorithm, including the following steps:

[0235] Determine the initialization position range of each particle based on the hyperparameter search space of the candidate model;

[0236] Determine the number of particles based on the data volume of the target data subset and the computational resource constraints;

[0237] Set the inertia weight. The inertia weight can be linearly decreased based on the number of iterations or fixed to a preset empirical value.

[0238] Setting the learning factor (cognitive factors) and (social factor), learning factor is set based on experience or adjusted based on the type of analysis task;

[0239] Randomly initialize the position vector and velocity vector of each particle, where the position vector represents a set of candidate hyperparameter combinations;

[0240] The fitness value of each particle is calculated based on the validation set of the target data subset. The fitness function is set according to the analysis task type:

[0241] Prediction tasks: fitness value = 1 / (1 + root mean square error);

[0242] Classification tasks: fitness value = F1 score;

[0243] Clustering task: fitness value = silhouette coefficient;

[0244] Anomaly detection tasks: Fitness value = area under the precision-recall curve or F1 score;

[0245] For each particle, the fitness value of its current position is compared with the fitness value of its historical optimal position. If the current position is better, the historical optimal position is updated to the global optimal position.

[0246] Update the individual optimal position and global optimal position of each particle;

[0247] Update particle speed and position according to the following formula:

[0248]

[0249] in, For particles In the number of iterations Velocity vector at time ; For particles In the number of iterations The position vector at , that is, the hyperparameter combination; 、 is the learning factor, which is used to control the weight of individual experience and social cooperation; 、 is a random number; For particles The best historical position; is the global optimal position of the entire particle swarm; is the inertia weight;

[0250] Iterative calculation is performed based on the number of particles, inertia weight, learning factor and speed until a preset number of iterations is reached or the difference between multiple consecutive fitness values ​​is less than a first fitness threshold;

[0251] The particle position corresponding to the global optimal position is used as the hyperparameter configuration of the candidate model.

[0252] This embodiment uses a particle swarm optimization algorithm to search for optimal hyperparameters. Compared with traditional grid search or random search, it can converge to the global optimal solution faster and reduce computing resource consumption. Secondly, a linear decrease strategy is used to balance global search and local development capabilities to improve optimization efficiency. The weights of individual experience and social collaboration are adjusted according to the task type to adapt to the optimization needs of different data characteristics. Dedicated fitness functions are designed for different analysis tasks to ensure that the optimization goals are consistent with business needs. Through dual control of the number of iterations and the fitness change threshold, over-optimization is avoided and the model performance and computing cost are balanced. The number of particles is dynamically adjusted according to the amount of data and computing resources to avoid resource waste while ensuring the optimization effect. It is suitable for medical big data scenarios.

[0253] In one specific implementation, APIs are used to acquire heterogeneous data from multiple sources, including hospital medical business systems, medical insurance settlement systems, consumables supply chains, and government financial platforms. Models are applied to unstructured medical record expense descriptions for entity recognition and relationship extraction, extracting key information such as drug costs and examination item costs and converting it into structured data. All acquired data is then aligned and formatted.

[0254] When determining the weight coefficients for each data source, we obtain the data update frequency index, historical accuracy index, and data provider credibility index. We use a formula to calculate the weight coefficients for each data source, and then perform weighted fusion based on this to generate multi-source heterogeneous data.

[0255] The isolation forest algorithm was used to detect and correct outliers, followed by a clustering algorithm to identify and eliminate noise data, handle missing values, and generate clean data. Based on operational analysis goals such as controlling consumable costs and predicting medical revenue, a feature dataset was constructed from the clean data. Finally, the random forest algorithm and Pearson correlation coefficient were used to select features and generate the target dataset.

[0256] Convert the target dataset into a structured dataset using a star schema. Determine the time partition granularity based on historical query frequency. If statistics show that the quarterly query frequency reached 40% over the past year and the data volume growth trend is stable, use quarterly as the time partition granularity and divide the data into primary storage partitions.

[0257] By business department, we code departments like cardiology and respiratory medicine, and divide data records within the primary partition into corresponding secondary sub-partitions. By data type, we divide data like income and expenditure into tertiary sub-partitions. We create a partition metadata index table to record information such as the time range, department code, and data type code for each partition, and distribute the data subsets to the corresponding physical storage nodes.

[0258] The user sends a query request "Predict the medical revenue of cardiology in the first quarter of 2025, and analyze the correlation between consumables usage and revenue." The system parses the analysis target parameters as time series forecast (forecast type), medical revenue (analysis indicator), and cardiology in the first quarter of 2025 (business constraints).

[0259] Based on this, a set of candidate models was screened from the preset model library, and the long short-term memory network model was selected for time series prediction; the random forest model was selected for analyzing the correlation between consumables usage and income.

[0260] Candidate models were evaluated for performance using sample data from the target data subset covering cardiology departments from 2024 to 2025. Root mean square error and mean absolute percentage error were used for time series prediction, while feature importance and correlation coefficients were used to assess the correlation between consumable usage and income. If the evaluation metrics were greater than or equal to the preset thresholds, the long-short-term memory network model was selected for time series prediction; the random forest model was selected for analyzing the correlation between consumable usage and income, without optimization.

[0261] The target data subset was analyzed using the selected model. The long short-term memory network model predicted that the cardiology medical revenue in the first quarter of 2025 would be 8 million yuan. The random forest model analysis concluded that the correlation between the usage of a certain type of high-value consumables and surgical revenue was 0.7.

[0262] Based on the analysis results, a template engine is used to generate a report containing text descriptions, revenue forecast line charts, bar charts showing the correlation between consumables and revenue, and other content. The report is presented to users through a graphical user interface, supporting interactive operations such as data filtering and chart zooming.

[0263] refer to Figure 2 , a flow chart of a data analysis method based on machine learning; this method includes the following functions:

[0264] HIS (hospital information system, responsible for diagnosis and treatment / billing, etc.), LIS (laboratory information system, responsible for laboratory data), PACS (imaging system, responsible for medical imaging) and other medical business systems are used as data sources. With the help of the Kettle tool in the ETL platform, data scattered in various systems are extracted according to the rules set by the collection and scheduling. Preliminary conversion is also performed to ensure that the data format and specifications are adapted to the subsequent processes.

[0265] Unify the acquired data, sort out the data relationships and structures based on the preset data modeling, and make the messy data regular and usable;

[0266] Load the data that has completed data modeling into the data warehouse that has completed physical modeling (designing the storage structure of the data warehouse).

[0267] Match relevant analysis models and corresponding data based on customer query requests;

[0268] The corresponding data is input into the analysis model to obtain the analysis results. The analysis model can explore the deep value of the data, predict future conditions, revenue trends and risk indicators, etc., and monitor key performance indicators.

[0269] The analysis results (for example, department daily revenue and revenue trends) are displayed intuitively on a large visual screen, allowing managers to understand hospital operations in real time.

[0270] Corresponding to the data analysis method based on machine learning in the above embodiment, Figure 3 This is a structural block diagram of a data analysis system based on machine learning provided in one embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown. Figure 3 The data analysis system 20 based on machine learning includes: a data acquisition module 21, a data processing module 22, a data partitioning module 23, a data matching module 24, a model analysis module 25 and a data interaction module 26.

[0271] The data acquisition module 21 is used to acquire and integrate the multi-source heterogeneous data sources of the target hospital, including medical business system data, medical insurance settlement data, and consumables supply chain data;

[0272] The data processing module 22 is used to process multi-source heterogeneous data to obtain a target data set;

[0273] The data partitioning module 23 is used to store the target data set into multiple storage partitions based on a preset data model and partitioning rules to obtain multiple data subsets. The partitioning rules support dividing data by at least one of the following dimensions: time range, business department, and data type, and the partition granularity is adjusted according to the historical query frequency.

[0274] The data matching module 24 is used to parse the filter conditions in the user query request and match the target data subset from multiple data subsets;

[0275] Model analysis module 25, configured to select an analysis model from a preset model library based on the analysis target parameters in the query request and the partition attributes of the target data subset, analyze the target data subset, and output the results;

[0276] The data interaction module 26 is used to generate corresponding reports based on the analysis results and interact with the user through a graphical user interface.

[0277] See also Figure 4 , Figure 4 This is a schematic block diagram of an electronic device provided in one embodiment of the present application. Figure 4 The electronic device 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. The processor 301 is configured to call the program instructions to execute the functions of the modules in the above-mentioned device embodiments, such as Figure 3 The functions of the data acquisition module 21, the data processing module 22, the data partition module 23, the data matching module 24, the model analysis module 25 and the data interaction module 26 are shown.

[0278] It should be understood that, in the embodiments of the present application, the processor 301 can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0279] The input device 302 can include a touchpad, a fingerprint acquisition sensor (for acquiring fingerprint information and direction information of a fingerprint of a user), a microphone, etc., and the output device 303 can include a display (LCD, etc.), a speaker, etc.

[0280] The memory 304 can include read-only memory and random access memory, and provide instructions and data for the processor 301. A part of the memory 304 can also include non-volatile random access memory. For example, the memory 304 can also store device type information.

[0281] In specific implementations, the processor 301, the input device 302 and the output device 303 described in the embodiments of the present application can perform the implementation manner described in any embodiment of the machine learning-based digital operation analysis method provided by the embodiments of the present application, and can also perform the implementation manner of the electronic device described in the embodiments of the present application, which will not be described here.

[0282] In another embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, all or part of the process of the method in the above embodiment is implemented. The computer program can also be used to instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above method embodiments are implemented. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium.

[0283] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the aforementioned embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the computer-readable storage medium can include both an internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or is about to be output.

[0284] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0285] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the electronic devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0286] In several embodiments provided in the present application, it should be understood that the disclosed electronic device and method can be implemented in other manners. For example, the embodiments of the apparatus described above are merely schematic; the division of the units is merely logical function division; an actual implementation can be divided into different units depending on actual conditions; or a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, or can be in electrical, mechanical or other forms.

[0287] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.

[0288] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.

[0289] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto; any skilled person in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be encompassed in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data analysis method based on machine learning, characterized in that: include: Acquire and integrate the target hospital's multi-source heterogeneous data sources, including medical business system data, medical insurance settlement data, and consumables supply chain data; Processing the multi-source heterogeneous data to obtain a target data set; Based on a preset data model and partitioning rules, the target data set is stored in multiple storage partitions to obtain multiple data subsets; the partitioning rules support data division based on at least one of the following dimensions: time range, business department, and data type, and the partition granularity is adjusted based on historical query frequency; Parsing the filter conditions in the user query request and matching the target data subset from the multiple data subsets; Select an analysis model from a preset model library based on the analysis target parameters in the query request and the partition attributes of the target data subset, analyze the target data subset, and output the analysis results; A corresponding report is generated based on the analysis results, and interacted with the user through a graphical user interface.

2. The data analysis method based on machine learning according to claim 1, characterized in that: The acquisition of multi-source heterogeneous data of the target hospital includes: Obtain structured and unstructured data from the target hospital's internal systems and external data platforms through API interfaces; Use natural language processing models to perform entity recognition and relationship extraction on unstructured data to generate structured processing data; Performing data mode alignment and format unification processing on the structured data and the structured processed data; The aligned data are weightedly fused based on the weight coefficients of each data source of the target hospital to generate multi-source heterogeneous data.

3. The data analysis method based on machine learning according to claim 2, characterized in that: The method for determining the weight coefficient of each data source includes: Obtain evaluation index data for each data source of the target hospital, including: data update frequency index, historical accuracy index, and data provider reputation rating index; The weight coefficient of each data source is calculated based on the evaluation index data and a preset first weight calculation formula.

4. The data analysis method based on machine learning according to claim 3, characterized in that: After calculating the weight coefficient of each data source according to the evaluation index data and the preset first weight calculation formula, the method further includes: In response to the data update frequency indicator being less than a first update frequency threshold, reducing the weight coefficient of the data source by a first step; In response to the data update frequency index being less than the second update frequency threshold and greater than or equal to the first update frequency threshold, the weight coefficient of the data source is reduced by a second step size; wherein the first update frequency threshold is less than the second update frequency threshold, and the first step size is greater than the second step size.

5. The data analysis method based on machine learning according to claim 1, characterized in that: Storing the target data set into multiple storage partitions based on a preset data model and partitioning rules includes: Convert the target dataset into a structured dataset according to the preset data model; Determine the time partition granularity based on the historical query frequency, and divide the structured data set into primary storage partitions according to the time range dimension; Divide the data records in the primary storage partition into secondary sub-partitions according to the business department dimension; Divide the data records in the second-level sub-partition into three-level sub-partitions according to the data type dimension; Establish a partition metadata index table to record the time range, business department code and data type code of each partition; The data subsets that have completed the three-level division are distributed and stored in corresponding physical storage nodes to generate the multiple data subsets.

6. The data analysis method based on machine learning according to claim 5, characterized in that: Determining the time partition granularity based on the historical query frequency includes: Collect statistics on the query frequency of historical query requests in different time intervals; Based on the mapping relationship between query frequency and time interval, select the day-level, month-level, or year-level partition granularity.

7. The data analysis method based on machine learning according to claim 6, characterized in that: The selection of daily, monthly or grade-level partitioning granularity based on the mapping relationship between query frequency and time interval includes: When the time interval meets the first time interval condition and the corresponding query frequency is greater than the first query frequency threshold, the day-level partition granularity is adopted; When the time interval meets the second time interval condition and the corresponding query frequency is greater than the second query frequency threshold, the monthly partition granularity is adopted; When the time interval meets the third time interval condition and the corresponding query frequency is greater than the third query frequency threshold, the grade partition granularity is adopted.

8. The data analysis method based on machine learning according to claim 1, characterized in that: The parsing of the filter conditions in the user query request and matching the target data subset from the multiple data subsets includes: Parse the filter conditions in the user's query request to obtain the query time interval, query department set and query data type code; Query the partition metadata index table, filter the partitions that meet the filter conditions, and obtain the filtered partitions; The screening conditions include: the partition time range overlaps with the query time interval, the partition business department code belongs to the query department set, and the partition data type code matches the query data type code; The loaded data is determined based on the overlap between the time range of the filter partition and the query time interval, and the loaded data is merged into a target data subset.

9. The data analysis method based on machine learning according to claim 8, characterized in that: Determining the loading data based on the overlap between the time range of the filtered partition and the query time interval includes: When the degree of overlap is less than a first overlap threshold, skipping the partition; When the overlap is greater than or equal to a first overlap threshold and less than a second overlap threshold, data whose query time range overlaps with the partition time range is loaded; When the degree of overlap is greater than or equal to a second overlap threshold, the data of the partition is loaded.

10. A data analysis system based on machine learning, characterized in that: The method according to any one of claims 1 to 9 is implemented, wherein the system comprises: The data acquisition module is used to acquire and integrate the target hospital's multi-source heterogeneous data sources, including medical business system data, medical insurance settlement data, and consumables supply chain data; A data processing module, configured to process the multi-source heterogeneous data to obtain a target data set; A data partitioning module is configured to store the target data set into multiple storage partitions based on a preset data model and partitioning rules to obtain multiple data subsets; the partitioning rules support partitioning data by at least one of the following dimensions: time range, business department, and data type, and the partitioning granularity is adjusted based on historical query frequency; A data matching module, configured to parse the filter conditions in the user query request and match a target data subset from the multiple data subsets; A model analysis module is used to select an analysis model from a preset model library based on the analysis target parameters in the query request and the partition attributes of the target data subset, analyze the target data subset and output the results; The data interaction module is used to generate a corresponding report based on the analysis results and interact with the user through a graphical user interface.

Citation Information

Patent Citations

  • Data query method and device, storage medium and computer equipment

    CN114880329A

  • Hospital quality monitoring data analysis and fine management system and method

    CN117038025A

  • Reconciliation data processing method and device, equipment and storage medium

    CN117056340A

  • Data analysis processing system based on big data

    CN117851490A

  • Data query method and device

    CN118035289A