Case resource integration data system based on big data analysis

By constructing a case resource integration system based on big data analysis, the problems of data dispersion, inconsistent formats, and limited risk analysis in case data management have been solved. This has enabled high-quality sharing of case data and accurate risk analysis, supporting precision medicine and efficient scientific research.

CN121191671APending Publication Date: 2025-12-23BEIJING YOUAN HOSPITAL CAPITAL MEDICAL UNIV +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511284247.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing case data management suffers from problems such as data fragmentation, inconsistent formats, limited risk analysis, and a disconnect between updates and clinical applications, making it difficult to achieve cross-system sharing and in-depth data analysis, and failing to meet the needs of precision medicine and efficient scientific research.

Method used

A case resource integration system based on big data analysis is constructed. Through multi-source data acquisition, desensitization processing, standardization transformation, and shallow and deep analysis, predictive diagnostic reports are generated, realizing standardized integration of the entire data chain and multi-dimensional risk analysis. Combined with shallow threshold comparison and deep trend modeling, multi-level alarm signals and contingency plan matching are generated.

Benefits of technology

It enables high-quality sharing and reuse of case data, improves the accuracy and foresight of risk analysis, quickly identifies abnormal explicit indicators and captures the risk of hidden trend deterioration, and supports precision medicine and efficient scientific research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121191671A_ABST
    Figure CN121191671A_ABST
Patent Text Reader

Abstract

The invention discloses a case resource integration data system based on big data analysis, and belongs to the technical field of medical data. The method comprises the following steps: acquiring hospital case data and corresponding disease type data to construct a resource integration range, acquiring a personal case information set provided by medical consultation of a patient, performing sensitive data extraction on the personal case information set to obtain a dynamic case parameter set, and sending the dynamic case parameter set to a data risk analysis module; the multi-source data acquisition module processes the personal case information set as follows; according to the method, a data integration-risk analysis-clinical intervention closed-loop system is constructed, preorder data standardization integration guarantees analysis reliability, accurate risk analysis provides a direction for intervention, multi-level alarm and pre-plan matching is achieved through linkage of the preorder data standardization integration and the accurate risk analysis, prediction diagnosis reports and intervention suggestions are automatically generated, invalid operations are reduced, the clinical decision-making efficiency is improved, and the system is suitable for large-scale popularization and application. And meanwhile, through dynamic threshold updating and system self-iteration optimization, the adaptability and practicability of the system are continuously enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data technology, specifically to a case resource integration data system based on big data analysis. Background Technology

[0002] Medical case resources are the core foundation of medical clinical diagnosis and treatment and scientific research. They include multi-source data such as hospital electronic medical records, laboratory reports, imaging diagnoses, pathology results, and personal symptom information provided by patients during medical consultations. They cover key content such as disease classification, clinical indicators, and treatment records. The massive amount of data is scattered across different business systems, with diverse formats and dynamic updates. Its effective integration is of great significance for disease diagnosis, optimization of treatment plans, and medical research. In the traditional model, case data mostly relies on manual processing, making it difficult to achieve cross-system sharing and in-depth analysis. However, with the development of big data technology, building a standardized and intelligent case resource integration system has become a key requirement for improving medical quality and scientific research efficiency.

[0003] The existing case data management has significant shortcomings: First, the data is scattered and inconsistent in format. Multi-source data such as electronic medical records and laboratory reports lack standardized conversion, making integration difficult and prone to duplicate entry or logical errors. Second, risk analysis is limited in scope, relying heavily on single-indicator threshold comparisons and ignoring the dynamic trends of indicators and the linkage between multiple parameters, which can easily lead to false alarms or omissions. Third, data updates are disconnected from clinical applications. Static thresholds are not included in new case data, and there is a lack of closed-loop management from data collection and analysis to intervention, making it difficult to meet the actual needs of precision medicine and efficient scientific research.

[0004] To address the aforementioned technical shortcomings, a solution is proposed. Summary of the Invention

[0005] The purpose of this invention is to provide a case resource integration data system based on big data analysis to solve the problems mentioned above.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a case resource integration data system based on big data analysis, which acquires hospital case data and corresponding disease type data to construct the resource integration scope, including the following steps;

[0007] The resource management platform acquires new and old data from the resource integration scope and builds temporary cache databases and historical databases.

[0008] Multi-source data acquisition module: Acquires the personal medical information set provided by patients during medical consultations, extracts sensitive data from the personal medical information set, obtains a dynamic medical parameter set, and sends it to the data risk analysis module;

[0009] Data Risk Analysis Module: Constructs shallow analysis unit and deep analysis unit. The shallow analysis unit obtains dynamic case parameter set, decomposes the indicators to obtain related dataset, and performs shallow analysis with threshold and historical database to generate normal log, fluctuation abnormal log, indicator abnormal signal and shallow risk value and send them to the comprehensive evaluation and analysis module.

[0010] The deep analysis unit acquires the associated dataset and the rate of change of indicators △Zt, selects the time series data of key indicators in the historical database to construct a trend curve, generates a normal trend signal, a deteriorating trend signal and deep risk nodes, and sends them to the comprehensive evaluation and analysis module.

[0011] Comprehensive assessment and analysis module: acquires historical data and several pre-stored symptom data, constructs several corresponding symptom treatment plans, obtains comparison results and matches them with the treatment plans, generates a predictive diagnosis report and feeds it back to the resource management platform.

[0012] Furthermore, the multi-source data acquisition module processes the individual medical record information set as follows:

[0013] The system acquires a personal medical record set, which includes hospital electronic medical records, laboratory reports, imaging diagnoses, pathology results, and basic information about symptoms provided by patients through consultations. Sensitive data is anonymized; miscellaneous items caused by entry / registration errors are removed; data of different formats are standardized and converted; and the units and timestamps of laboratory indicators are unified.

[0014] The processed data is spatiotemporally labeled according to patient name, disease type, and collection time to construct a dynamic case parameter set, which is then stored in batches to a temporary cache database. The historical database is synchronously updated with data of similar cases from previous periods.

[0015] Furthermore, the data risk analysis module processes the dynamic case parameter set as follows:

[0016] The dynamic case parameter set is classified into primary categories according to disease type, and then further broken down into clinical indicator types, namely laboratory indicators, imaging indicators, and pathological indicators. Numerical and textual data of each indicator are extracted, and the textual indicators are structured. A three-level indicator matrix of disease type, indicator type, and specific indicator is constructed, and patient basic information is associated to form an associated dataset.

[0017] Furthermore, the analysis process of the shallow analysis unit is as follows:

[0018] Retrieve normal and warning thresholds for indicators of the same disease and stage from the historical database, and compare the associated dataset of the current case with the normal thresholds: if the associated dataset is within the normal threshold and the fluctuation range is less than the warning threshold, generate a normal log; if the associated dataset is within the normal threshold but the fluctuation range is greater than or equal to the warning threshold, generate an abnormal fluctuation log; if the associated dataset exceeds the normal threshold and the fluctuation range is greater than or equal to the warning threshold, generate an abnormal indicator signal; at the same time, obtain the deviation value between the indicator compared in the associated dataset and the normal threshold of the indicator in the historical data, and mark it as a shallow risk value.

[0019] Furthermore, the analysis process of the deep analysis unit is as follows:

[0020] Obtain the associated dataset, extract the recording period of the disease in the associated dataset, construct several period nodes according to the recording period from the onset of the disease to the current time, obtain the numerical change of the same indicator under different period nodes, and mark it as the current indicator reference value. Retrieve the past records of the same type of data in the historical database and mark them as past indicator reference values. Obtain the indicator value difference between the current indicator reference value and the past indicator reference value. Finally, with the time interval of the period nodes as the denominator and the numerical difference as the numerator, obtain the indicator change rate ΔZt per unit time.

[0021] Furthermore, based on the historical database, key indicator time-series data from the past three months were selected, and a trend curve was constructed using a time series analysis model. The data collection timestamp was used as the X-axis, and the indicator value was used as the Y-axis. The rate of change of the indicator ΔZt over two consecutive periods was compared, and the recovery trend curve of similar cases in the historical database was retrieved as a reference baseline. If the current curve fits the baseline by ≥85%, a normal trend signal is generated; if the fit is <85% and the rate of change of the indicator ΔZt >0.2 / week, a trend deterioration signal is generated. The time point and indicator value corresponding to the inflection point of the curve are extracted and marked as deep risk nodes.

[0022] Furthermore, the processing procedure of the comprehensive evaluation and analysis module is as follows:

[0023] Based on historical databases, pre-stored disease management plans are retrieved. These plans are categorized into three triggering conditions: routine care, intensive monitoring, and interventional treatment. The shallow risk values ​​and deep risk nodes generated by the multi-source data acquisition module are matched against the triggering conditions of the disease management plans: if only normal logs and trends are present, the routine care plan is matched; if abnormal logs or trends are present, the intensive monitoring plan is matched; if abnormal indicators or worsening trends are present, the interventional treatment plan is matched. The matching results are integrated with the patient's basic information to generate a predictive diagnostic report containing risk level and recommended measures.

[0024] Furthermore, based on the risk level of the predicted diagnostic report, multi-level alarm signals are generated. Level 1 alarm corresponds to the routine nursing plan, and weekly data summaries are pushed to the medical staff terminal; Level 2 alarm corresponds to the enhanced monitoring plan, increasing the data collection frequency to once a day and pop-up reminders; Level 3 alarm corresponds to the intervention and treatment plan, and immediately sends audible and visual alarms to the responsible physician's terminal; for cases with three consecutive Level 1 alarms, the data archiving cycle is automatically extended to once a month; for cases with Level 3 alarms, an examination checklist is generated simultaneously.

[0025] Furthermore, the data in the temporary cache database is quality-checked every morning at midnight to remove duplicate entries and data with logical errors; the verified data is archived to the historical database by disease type and indicator type, and the number of cases and time range in the database are updated; based on the new data, the threshold of each indicator is updated using the moving average method, and the baseline curve of the trend analysis model is updated synchronously to ensure the timeliness of historical data.

[0026] Furthermore, the accuracy rate of predictive diagnostic reports based on clinical feedback is collected monthly. If the accuracy rate is less than 90%, the data preprocessing stage is reviewed, and the desensitization and standardization parameters are adjusted. For abnormal signals with frequent false alarms, the rationality of the threshold is analyzed using a confusion matrix, and the warning threshold is corrected. Based on the case data of newly added diseases, the disease classification and contingency plan library are expanded. The optimized parameters and newly added contingency plans are stored in the system configuration library to achieve self-iterative optimization.

[0027] The beneficial effects of this invention are:

[0028] 1. This invention achieves standardized integration of case data across the entire data chain. Through multi-source data anonymization type extraction, noise reduction, and standardized conversion, sensitive information and erroneous miscellaneous data are removed, and indicator units and timestamps are unified to construct a structured dynamic case parameter set. This is used to solve the problems of traditional data being scattered and having chaotic formats, providing a high-quality data foundation for subsequent analysis and improving data sharing and reuse capabilities.

[0029] 2. This invention improves the accuracy and foresight of risk analysis by combining shallow threshold comparison with deep trend modeling. Through multi-dimensional output of normal / abnormal logs, risk signals, and trend change rates, it can quickly identify abnormal explicit indicators and capture the risk of hidden trend deterioration, overcoming the limitations of single parameter analysis and providing a scientific basis for clinical risk assessment. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a flowchart of the system of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Example 1: Please refer to Figure 1 As shown, this embodiment is a case resource integration data system based on big data analysis, which acquires hospital case data and corresponding disease type data to construct the resource integration scope, including the following steps;

[0034] Multi-source data acquisition module: Acquires the personal medical record information set provided by patients during medical consultations, extracts sensitive data from the personal medical record information set, obtains a dynamic medical record parameter set, and sends it to the data risk analysis module; the processing procedure of the personal medical record information set by the multi-source data acquisition module is as follows:

[0035] The system acquires a personal medical record set, which includes hospital electronic medical records, laboratory reports, imaging diagnoses, pathology results, and basic information about the patient's condition provided through consultations. Sensitive data is anonymized by removing identifying information such as names and ID numbers; miscellaneous items caused by entry / registration errors are removed; data of different formats is standardized and converted to unify the units of measurement for laboratory indicators and timestamps.

[0036] The processed data is spatiotemporally labeled according to patient name, disease type, and collection time to construct a dynamic case parameter set, which is then stored in batches to a temporary cache database. The historical database is updated synchronously with data of similar cases from previous periods. It should be noted that data quality is ensured through desensitization, noise reduction, and standardization, laying the foundation for subsequent analysis.

[0037] The data risk analysis module processes the dynamic case parameter set as follows:

[0038] The dynamic case parameter set is obtained and classified into primary categories according to disease type, including hepatitis B cirrhosis and autoimmune liver disease. Under the disease type classification, clinical indicators are further broken down into laboratory indicators, imaging indicators, and pathological indicators. Laboratory indicators include ALT and HBV-DNA; imaging indicators include liver capsule smoothness; and pathological indicators include inflammation grade.

[0039] Numerical and textual data for each indicator were extracted. Textual indicators underwent structured transformation; for example, "unsmooth liver capsule" was converted into a quantitative score. It should be noted that a 5-point scale was used: 1 point (smooth), 2 points (relatively smooth), 3 points (less smooth), 4 points (rough), and 5 points (unsmooth). This standard is derived from the Radiology Department's "Structured Standard for Abdominal Ultrasound Reports." Verification through correlation analysis of several imaging reports and pathology results showed a Kappa value consistency coefficient of 0.87. A three-level indicator matrix of disease type, indicator type, and specific indicator was constructed, linking patient basic information with age and gender to form an associated dataset. It should be noted that hierarchical decomposition facilitates data structuring and precise analysis.

[0040] The shallow analysis unit acquires the dynamic case parameter set, decomposes the indicators to obtain the associated dataset, and performs shallow analysis with the threshold and historical database to generate normal logs, fluctuation abnormal logs, indicator abnormal signals, and shallow risk values, which are then sent to the comprehensive assessment and analysis module. The analysis process of the shallow analysis unit is as follows:

[0041] Normal and warning thresholds for indicators of the same disease and stage were retrieved from the historical database. Normal thresholds included ALT in hepatitis B cirrhosis, with a normal range of 0-40 U / L, derived from statistical analysis of cases of the same disease and stage in the historical database. These thresholds were used to determine whether the current case's ALT level was within the physiologically normal range. The warning threshold could be set at ALT > 80 U / L. Based on statistical analysis of ALT thresholds for liver function damage in hepatitis B cirrhosis patients in the historical database, when ALT exceeds 80 U / L, the risk of liver damage progression increases by 3.2 times. Based on historical case follow-up data, this threshold was set as the warning threshold. This distinguishes between thresholds requiring attention and thresholds requiring intervention for indicator fluctuations. The current case's associated dataset was compared with the normal thresholds.

[0042] If the associated dataset is within the normal ALT threshold of 0-40 U / L and the ALT fluctuation is less than the warning threshold of 80 U / L, a routine log is generated. The shallow analysis unit compares the current indicator with the normal and warning thresholds. After confirming that there are no fluctuations exceeding the range, the log is automatically generated and fed back to the resource management platform. The resource management platform automatically pushes the log to the medical staff terminal and simultaneously builds a routine case set in the historical database. Medical staff receive data summaries once a week without additional intervention. After three consecutive routine logs are generated, the system automatically extends the data collection cycle to once every three days.

[0043] If the associated dataset is within the normal threshold but the fluctuation range is greater than or equal to the warning threshold, an abnormal fluctuation log is generated. The generated normal log is fed back to the resource management platform, which automatically pushes the log to the medical staff terminal with a pop-up reminder of the text message "Abnormal indicator fluctuation (not exceeding the normal range)". Medical staff can view the indicator baseline comparison chart through the system within 2 hours to confirm whether the temporary fluctuation is caused by medication or diet. If the fluctuation continues, the system automatically increases the detection frequency to once a day.

[0044] If the associated dataset exceeds the normal threshold and the fluctuation range is greater than or equal to the warning threshold, an abnormal indicator signal is generated. The generated routine log is fed back to the resource management platform. The resource management platform sends an audible and visual alarm to the responsible physician, and the terminal displays the text message "Abnormal indicator, high risk of liver damage". The physician retrieves historical test data and imaging reports within 30 minutes to assess whether to start liver protection treatment. At the same time, an examination list is generated: 13 liver function tests and abdominal ultrasound examination.

[0045] Simultaneously, the deviation values ​​of the indicators in the associated dataset compared with the normal threshold values ​​of the indicators in historical data are obtained and marked as shallow risk values. It should be noted that: the shallow risk value is calculated by the deviation value of the indicators in the associated dataset compared with the historical normal threshold values. The formula is: Shallow Risk Value = (Current Indicator Value - Upper Limit of Normal Threshold) / Upper Limit of Normal Threshold (if the indicator is abnormal); or = (Current Indicator Value - Historical Average) / Historical Average (if the indicator is normal but fluctuates; it quantifies the degree to which the indicator deviates from the normal range, and the larger the value, the higher the risk; it is displayed in the predictive diagnosis report as a score of 0-10 (≥6 points require intervention), and the physician adjusts the monitoring frequency based on the risk value (e.g., if the risk value is 8 points, it is monitored daily); explicit risks are quickly identified through threshold comparison, and preliminary signals are generated.

[0046] The deep analysis unit acquires the associated dataset and the rate of change of indicators ΔZt, selects key indicator time-series data from the historical database to construct trend curves, generates normal trend signals, deteriorating trend signals, and deep risk nodes, and sends them to the comprehensive evaluation and analysis module. The analysis process of the deep analysis unit is as follows:

[0047] Obtain the associated dataset, extract the recording period of the disease in the associated dataset, and construct several period nodes according to the recording period from the onset of the disease to the current time. Obtain the numerical changes of the same indicator at different period nodes and mark them as the current indicator reference value. The current indicator reference value includes key clinical indicators in the case data, pH value and exudate volume in wound monitoring. Retrieve past records of similar data in the historical database and mark them as past indicator reference values. Past indicator reference values ​​include historical indicator records of the past N days, where N can be 30. Based on the timeliness of short-term fluctuations in clinical indicators, select historical data of the past 3 days to balance the data volume and recent correlation, and refer to the indicator changes in the historical database. The average cycle is determined to be 28 days. The difference between the current reference value and the previous reference value is obtained. Finally, the time interval of the cycle node is used as the denominator and the numerical difference is used as the numerator to obtain the rate of change of the indicator per unit time, ΔZt. It should be noted that: for example, in wound monitoring, the rate of change of exudate volume ΔZt is generated by the ratio of the difference of exudate volume at consecutive cycle nodes to the time interval; the rate of change of the indicator ΔZt in case data is calculated by the time series difference of key indicators in the past 3 months and the corresponding time interval. It should be noted that: the time interval of the cycle node is divided equally according to the disease recording cycle, with each cycle node being 7 days; because the changes of indicators in chronic diseases such as hepatitis B and cirrhosis show regularity on a weekly basis, the 7-day interval can accurately capture trend changes, and the minimum cycle of indicator fluctuation in historical follow-up data of 6.8 days is used as a reference.

[0048] Based on historical database data, key indicators from the past three months were selected, and a time series analysis model was used to construct trend curves. The data collection timestamp was used as the X-axis, and the indicator values ​​as the Y-axis. The rate of change (ΔZt) of the indicators over two consecutive periods was compared. The recovery trend curves of similar cases in the historical database were used as a reference baseline. It should be noted that the recovery trend curves were analyzed using the indicator trend curves of several recovered hepatitis B cirrhosis patients from the historical database. If the current curve's fit to the recovery baseline is ≥85%, the probability of clinical recovery meeting expectations is 92%. Therefore, this is used as the normal threshold to determine whether the current case's indicator trend is consistent with the recovery standard trend. The rate of change (ΔZt) is based on the temporal characteristics of indicator deterioration in hepatitis B cirrhosis patients in historical data. When the weekly change rate of the key indicator ALT exceeds 0.2 U / L, the risk of disease progression increases significantly, with a positive predictive value of 89%. Therefore, this is set as the threshold for judging trend deterioration.

[0049] If the current curve fits the baseline by ≥85%, a normal trend signal is generated; the normal trend signal is fed back to the resource management platform, which marks "the trend meets the recovery expectation" in the prediction and diagnosis report; the original treatment plan is maintained, and patient education information and dietary precautions are pushed; the trend curve is updated every 2 weeks and included in the historical database;

[0050] If the fit is less than 85% and the rate of change of the index ΔZt is greater than 0.2 / week, a trend deterioration signal is generated. The trend deterioration signal is fed back to the resource management platform, which triggers a level-two alarm and pushes a risk warning to the department director. The medical team conducts a multidisciplinary consultation within 4 hours, adjusts the treatment plan, and increases the dosage of antiviral drugs. The changes in the index are monitored daily until the signal is eliminated.

[0051] Extract the time points and indicator values ​​corresponding to the inflection points of the curve and mark them as deep risk nodes. It should be noted that trend modeling is used to capture changes in hidden risks.

[0052] Example 2:

[0053] The comprehensive assessment and analysis module acquires historical data and several pre-stored disease information, constructs several corresponding treatment plans for each disease, matches the comparison results with the treatment plans, generates a predictive diagnosis report, and feeds it back to the resource management platform. The processing procedure of the comprehensive assessment and analysis module is as follows:

[0054] Based on historical databases, pre-stored disease management plans are retrieved. These plans are categorized into three levels of triggering conditions: routine care, intensive monitoring, and interventional treatment. The shallow risk values ​​and deep risk nodes generated by the multi-source data acquisition module are matched against the triggering conditions of the disease management plans.

[0055] If only the routine logs and trend signals are normal, then the standard nursing care plan should be matched.

[0056] If the log contains abnormal fluctuations or normal trend signals, then an enhanced monitoring plan will be implemented.

[0057] If there are abnormal indicators or signs of worsening trends, then a matching intervention and treatment plan will be implemented.

[0058] Integrating matching results with patient basic information generates a predictive diagnostic report containing risk levels and recommended measures. It should be noted that: Risk and countermeasures are precisely linked through plan matching. Risk levels are defined as three levels based on comprehensive assessment results, corresponding to clinical intervention priorities: Level 1: Stable condition, no emergency intervention required; Level 2: Fluctuations in indicators, requiring enhanced monitoring; Level 3: High risk of disease progression, requiring immediate clinical intervention. The comprehensive assessment results are weighted and scored based on shallow risk values ​​and deep risk nodes using the comprehensive assessment analysis module. Shallow risk values, such as a deviation rate <10%, are low risk; 10%-30% are medium risk; and >30% are high risk. Deep risk nodes, such as no inflection point, are low risk; a single inflection point is medium risk; and multiple inflection points are high risk. A score <30 indicates Level 1 risk, 30%-60 indicates Level 2 risk, and >60 indicates Level 3 risk. Risk levels are linked to plan matching results: Level 1 corresponds to a routine nursing plan, Level 2 to an enhanced monitoring plan, and Level 3 to an intervention treatment plan.

[0059] Example 3:

[0060] The resource management platform acquires new and old data from the resource integration scope and builds temporary cache databases and historical databases.

[0061] Generate multi-level alert signals based on the risk level predicted in the diagnostic report:

[0062] Level 1 alerts correspond to routine nursing care plans, with weekly data summaries pushed to healthcare terminals; for cases triggering three consecutive Level 1 alerts, the data archiving cycle is automatically extended to once a month.

[0063] The Level 2 alert corresponds to an enhanced monitoring plan, which increases the data collection frequency to once a day and provides a pop-up reminder.

[0064] The three-level alarm corresponds to the intervention and treatment plan, and immediately sends an audible and visual alarm to the responsible physician's terminal. For cases with a three-level alarm, an examination list is generated simultaneously, including liver function re-examination and imaging consultation. It should be noted that differentiated clinical responses are achieved through multi-level alarms.

[0065] Every day at midnight, the data in the temporary cache database is quality checked to remove duplicate entries and data with logical errors; the data that passes the check is archived to the historical database by disease type and indicator type, and the number of cases and time range in the database are updated.

[0066] The thresholds for each indicator are updated using the moving average method based on the newly added data. The calculation formula is as follows:

[0067] New threshold = α × old threshold + (1-α) × mean of new data

[0068] Where α=0.7; the old threshold represents multiple thresholds retrieved from the historical database; the mean of the newly added data represents the latest collected dynamic case parameter set data, and the baseline curve of the trend analysis model is updated synchronously to ensure the timeliness of historical data. It should be noted that the acquisition process of the new threshold data adopts a dynamic threshold update strategy. α is the weight coefficient of the old threshold, referring to the classic parameter setting of the moving average method in time series analysis. Because the historical threshold needs to retain high stability, it is assigned a weight of 0.7. The mean of the newly added data has a weight of 0.3, balancing historical stability and the timeliness of new data, ensuring that the threshold update both refers to historical experience and incorporates the latest clinical data, and ensuring the accuracy of the threshold and model through dynamic updates.

[0069] The accuracy rate of predicted diagnostic reports based on clinical feedback is collected monthly. If the accuracy rate is <90%, optimization is required. Data acquisition process: Based on the basic requirements of clinical reliability for predicted diagnostic reports, when the accuracy rate is ≥90%, the clinical decision-making compliance rate reaches 88%. If it is below 90%, the risk of misdiagnosis may increase. Therefore, it is set as the system optimization trigger threshold. The data preprocessing stage is reviewed, and the desensitization and standardization parameters are adjusted. For abnormal signals with frequent false alarms, the rationality of the threshold is analyzed using a confusion matrix. Target indicators for abnormal signals with frequent false alarms are selected, such as the ALT warning threshold for hepatitis B cirrhosis. Case data corresponding to this indicator in the past 3 months are extracted from the historical database. Valid samples containing actual clinical diagnostic results are screened out. For example, liver injury is diagnosed as "actual positive", and no liver injury is "actual negative". The sample size needs to be ≥300 cases to ensure statistical reliability.

[0070] Based on the current warning threshold, such as ALT > 80 U / L, predict and classify the samples:

[0071] Cases that are actually positive and predicted to be abnormal ≥ the threshold are marked as true positive TP;

[0072] Cases that are actually negative but are predicted to be abnormal (≥) the threshold are marked as false positives (FP), i.e., misreported cases.

[0073] Cases that are actually negative but predicted to be normal (below the threshold) are marked as true negative TN.

[0074] Cases that are actually positive but predicted to be normal (below the threshold) are marked as false negatives (FN), i.e., underreported cases; a 2×2 confusion matrix is ​​constructed based on this.

[0075] False positive rate / false alarm rate = FP / (FP+TN), which reflects the proportion of normal cases that are misclassified as abnormal. If this value is >20%, the threshold may be too low, considering the clinically acceptable false alarm criteria.

[0076] The false negative rate / false negative rate = FN / (TP+FN) reflects the proportion of abnormal cases that are misclassified as normal. If this value is >10%, the threshold may be too high.

[0077] Precision = TP / (TP+FP), which reflects the proportion of actual abnormalities among predicted abnormal cases. If this value is < 70%, the threshold needs to be optimized and the warning threshold needs to be corrected. The corrected threshold, such as ALT>90U / L, is applied to 200 newly collected samples, and the confusion matrix is ​​reconstructed to verify the improvement of the indicator. After confirming that the precision is ≥ 80%, the new threshold is stored in the system configuration library to complete the threshold update.

[0078] Based on the newly added disease case data, the disease classification and contingency plan database are expanded; the optimized parameters and new contingency plans are stored in the system configuration database to achieve self-iterative optimization. It should be noted that the system performance is continuously improved through feedback loop.

[0079] Combining Examples 1, 2, and 3, a closed-loop system of data integration, risk analysis, and clinical intervention is constructed. The standardized integration of preceding data ensures the reliability of the analysis, while precise risk analysis provides direction for intervention. The two work together to achieve multi-level alarms and plan matching, automatically generate predictive diagnostic reports and intervention suggestions, reduce ineffective operations, and improve the efficiency of clinical decision-making. At the same time, through dynamic threshold updates and system self-iterative optimization, the adaptability and practicality of the system are continuously enhanced.

[0080] The above description is merely an example and illustration of the structure of the present invention. Those skilled in the art can make various modifications or additions to the specific examples described, or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined in the claims, and all such modifications and additions should fall within the protection scope of the present invention.

[0081] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0082] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A case resource integration data system based on big data analysis, characterized in that, To establish a resource integration scope, we need to acquire hospital medical records and corresponding disease type data. Includes the following steps; The resource management platform acquires new and old data from the resource integration scope and builds temporary cache databases and historical databases. Multi-source data acquisition module: Acquires the personal medical information set provided by patients during medical consultations, extracts sensitive data from the personal medical information set, obtains a dynamic medical parameter set, and sends it to the data risk analysis module; Data Risk Analysis Module: Constructs shallow analysis unit and deep analysis unit. The shallow analysis unit obtains dynamic case parameter set, decomposes the indicators to obtain related dataset, and performs shallow analysis with threshold and historical database to generate normal log, fluctuation abnormal log, indicator abnormal signal and shallow risk value and send them to the comprehensive evaluation and analysis module. The deep analysis unit acquires the associated dataset and the rate of change of indicators △Zt, selects the time series data of key indicators in the historical database to construct a trend curve, generates a normal trend signal, a deteriorating trend signal and deep risk nodes, and sends them to the comprehensive evaluation and analysis module. Comprehensive assessment and analysis module: acquires historical data and several pre-stored symptom data, constructs several corresponding symptom treatment plans, obtains comparison results and matches them with the treatment plans, generates diagnostic reports and feeds them back to the resource management platform.

2. The case resource integration data system based on big data analysis according to claim 1, characterized in that, The multi-source data acquisition module processes the individual medical record information set as follows: The system acquires a personal medical record set, which includes hospital electronic medical records, laboratory reports, imaging diagnoses, pathology results, and basic information about the symptoms provided by patients through consultations. Sensitive data is anonymized to remove miscellaneous items caused by entry / registration errors. Data of different formats is standardized and converted to unify the units and timestamps of laboratory indicators. The processed data is spatiotemporally labeled according to patient name, disease type, and collection time to construct a dynamic case parameter set, which is then stored in batches to a temporary cache database. The historical database is synchronously updated with data of similar cases from previous periods.

3. The case resource integration data system based on big data analysis according to claim 1, characterized in that, The data risk analysis module processes the dynamic case parameter set as follows: The dynamic case parameter set is classified into primary categories according to disease type, and then further broken down into clinical indicator types, namely laboratory indicators, imaging indicators, and pathological indicators. Numerical and textual data of each indicator are extracted, and the textual indicators are structured. A three-level indicator matrix of disease type, indicator type, and specific indicator is constructed, and patient basic information is associated to form an associated dataset.

4. The case resource integration data system based on big data analysis according to claim 3, characterized in that, The analysis process of the shallow analysis unit is as follows: Retrieve normal and warning thresholds for indicators of the same disease and stage from the historical database, and compare the associated dataset of the current case with the normal thresholds: if the associated dataset is within the normal threshold and the fluctuation range is less than the warning threshold, generate a normal log. If the associated dataset is within the normal threshold but the fluctuation range is greater than or equal to the warning threshold, an abnormal fluctuation log is generated; if the associated dataset exceeds the normal threshold and the fluctuation range is greater than or equal to the warning threshold, an abnormal indicator signal is generated; at the same time, the deviation value between the indicator compared in the associated dataset and the normal threshold value of the indicator in the historical data is obtained and marked as a shallow risk value.

5. A case resource integration data system based on big data analysis according to claim 4, characterized in that, The analysis process of the deep analysis unit is as follows: Obtain the associated dataset, extract the recording period of the disease in the associated dataset, construct several period nodes according to the recording period from the onset of the disease to the current time, obtain the numerical change of the same indicator under different period nodes, and mark it as the current indicator reference value. Retrieve the past records of the same type of data in the historical database and mark them as past indicator reference values. Obtain the indicator value difference between the current indicator reference value and the past indicator reference value. Finally, with the time interval of the period nodes as the denominator and the numerical difference as the numerator, obtain the indicator change rate ΔZt per unit time.

6. The case resource integration data system based on big data analysis according to claim 5, characterized in that, Based on the historical database, key indicator time series data of the past 3 months were selected, and a trend curve was constructed using a time series analysis model. The collection time stamp was used as the X-axis and the indicator value was used as the Y-axis. The rate of change of the indicator ΔZt of two consecutive periods was compared, and the recovery trend curve of similar cases in the historical database was retrieved as a reference baseline. If the current curve fits the baseline by ≥85%, a normal trend signal is generated. If the fit is less than 85% and the rate of change of the index ΔZt is greater than 0.2 / week, a trend deterioration signal is generated; the time point and index value corresponding to the inflection point of the curve are extracted and marked as deep risk nodes.

7. The case resource integration data system based on big data analysis according to claim 1, characterized in that, The processing procedure of the comprehensive evaluation and analysis module is as follows: Based on the historical database, the pre-stored disease management plan set is retrieved. The disease management plan set is divided into three levels of triggering conditions: routine care, enhanced monitoring and intervention treatment. The shallow risk value and deep risk node generated by the multi-source data acquisition module are matched with the triggering conditions of the disease management plan set: if only normal logs and trend normal signals are available, then the routine care plan is matched. If the log contains abnormal fluctuations or normal trend signals, an enhanced monitoring plan is matched; if the log contains abnormal indicators or worsening trend signals, an intervention treatment plan is matched; the matching results are integrated with the patient's basic information to generate a predictive diagnostic report that includes risk level and recommended measures.

8. A case resource integration data system based on big data analysis according to claim 6, characterized in that, Based on the risk level of the predicted diagnostic report, a multi-level alarm signal is generated. The first-level alarm corresponds to the routine nursing plan, and weekly data summaries are pushed to the medical staff terminal. The second-level alarm corresponds to the enhanced monitoring plan, which increases the data collection frequency to once a day and sends a pop-up reminder. The third-level alarm corresponds to the intervention treatment plan, and an audible and visual alarm is sent to the terminal of the responsible physician in an instant. For cases that trigger Level 1 alerts three times consecutively, the data archiving cycle will be automatically extended to once a month. For cases triggering a Level 3 alert, a checklist is generated simultaneously.

9. A case resource integration data system based on big data analysis according to claim 1, characterized in that, Every day at midnight, the data in the temporary cache database is quality checked to remove duplicate entries and data with logical errors; the data that passes the check is archived to the historical database by disease type and indicator type, and the number of cases and time range in the database are updated; based on the new data, the threshold of each indicator is updated using the moving average method, and the baseline curve of the trend analysis model is updated synchronously to ensure the timeliness of historical data.

10. A case resource integration data system based on big data analysis according to claim 1, characterized in that, The accuracy rate of predictive diagnostic reports based on clinical feedback is collected monthly. If the accuracy rate is <90%, the data preprocessing stage is reviewed, and the desensitization and standardization parameters are adjusted. For abnormal signals with frequent false alarms, the rationality of the threshold is analyzed using a confusion matrix, and the warning threshold is revised. Based on the newly added disease case data, the disease classification and contingency plan database are expanded; the optimized parameters and new contingency plans are stored in the system configuration database to achieve self-iterative optimization.

Citation Information

Cited By

  • Multi-source data-oriented medical hidden danger analysis system

    CN121460046A

  • A multi-source data-oriented medical hidden danger analysis system

    CN121460046B