Biomedical data monitoring analysis method and system
By optimizing the storage resources of the medical data platform through dynamic scoring models and hash summary technology, the query response time delay problem caused by frequent data migration is solved, efficient data storage and query response are achieved, and storage costs and system risks are reduced.
Patent Information
- Application Number
- CN202510777632.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
AI Technical Summary
The storage I/O fluctuations caused by frequent data migration in existing medical data platforms lead to delayed data query response times, affecting clinical decision-making and systemic risks.
The storage value of medical data packets is evaluated through a dynamic scoring model to determine their pre-storage tier. Hash summary and physical reorganization technology are then used to optimize storage resources, eliminate data redundancy and inconsistency, and optimize storage tiers to reduce query response latency.
Effectively identify and eliminate unnecessary waste of storage resources, reduce platform costs, avoid data query response time delays, and reduce clinical decision delays and systemic risks.
Smart Images

Figure CN120673957A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical data analysis, and in particular to a monitoring and analysis method and system for biomedical data. Background Art
[0002] A medical data platform is a technical infrastructure that integrates multiple medical and health data resources and provides data storage, processing, analysis, and sharing services. This type of platform plays a central role in the development of healthcare informatization, integrating heterogeneous data from multiple sources, including electronic health records (EHRs), electronic medical records (EMRs), picture archiving and communication systems (PACS), laboratory information systems (LISs), and wearable devices. Biomedical data serves as the core foundation of the platform. The accuracy of medical data and query response time directly impact diagnosis and treatment, making data quality crucial for monitoring and ensuring its integrity and consistency.
[0003] With existing technologies, medical data platforms are already able to achieve real-time dynamic optimization of storage resources (such as automatic tiered storage, migration of hot and cold data storage levels, and compression strategy adjustment). However, frequent data migration can cause storage I / O fluctuations, leading to delayed data query response times, which in turn can trigger a series of chain reactions, ranging from delayed clinical decision-making to systemic risks. Summary of the Invention
[0004] The object of the present invention is to provide a method and system for monitoring and analyzing biomedical data to solve the problems raised in the above-mentioned background technology.
[0005] In order to solve the above technical problems, the present invention provides the following technical solution: a method for monitoring and analyzing biomedical data, the method comprising the following steps:
[0006] Obtaining a medical data package to be stored, and scoring the storage value of the medical data package based on the access frequency, clinical criticality score, and package volume of the medical data package in combination with a dynamic scoring model;
[0007] Determining a pre-storage level for the medical data packet based on the storage value scoring result, wherein the pre-storage level is a target storage location for the medical data packet;
[0008] determining an actual storage level of the medical data packet based on real-time storage space, and analyzing an actual query response delay time of the medical data packet based on a level difference between the pre-storage level and the actual storage level;
[0009] Determine whether the query response time is compliant based on the actual query response delay time of the medical data packet;
[0010] When it is determined that the query response time of the medical data package is not in compliance, the storage resources in the platform are optimized, including the upgrade processing of the data package storage layer and the replacement candidates of the medical data package pre-storage layer data.
[0011] The specific method of scoring the storage value of the medical data package in combination with the dynamic scoring model and determining the pre-storage level of the medical data package is as follows:
[0012] According to the dynamic scoring model: The storage value of the medical data package is scored, where the score is affected by the access frequency, clinical criticality, and data volume of the medical data package. It is positively correlated with the access frequency and clinical criticality score, and negatively correlated with the data volume. f represents the dynamic score of the medical data package, p represents the access frequency of the medical data package, and p represents the data volume. max α represents the highest access frequency of all data in the platform, m represents the clinical criticality score of the medical data package, which is the data importance assessed by clinical experts; M represents the total clinical criticality score, and g represents the data volume of the medical data package; α represents the access frequency weight coefficient, which is used to adjust the impact of access frequency on the total score; β represents the clinical criticality weight coefficient, which is used to adjust the impact of clinical importance on the total score; γ represents the data volume weight coefficient, which is used to adjust the negative impact of data size on the total score;
[0013] The pre-storage level of the medical data packet is determined according to the dynamic score, wherein the dynamic score and the pre-storage level are in a positive hierarchical mapping relationship, that is, the higher the score, the higher the storage level; the medical data storage platform includes multiple storage levels, each storage level is a logical subset of the storage level above it, and the storage level is positively correlated with the impact speed of data query.
[0014] Furthermore, the specific method for optimizing the storage resources in the platform is as follows:
[0015] Since it is impossible to determine whether the data in the pre-storage layer of the medical data package is fragmentary and duplicated, the core metadata fields of the data package in the pre-storage layer of the medical data package are extracted as the benchmark anchor for deduplication and integration. The core data includes unique identifiers, timestamp sequences and clinical event codes. Fragmentary duplicate entry refers to the phenomenon that the same or highly similar data fragments are entered into the system multiple times during the data entry process, resulting in data redundancy, inconsistency or errors. A field hash summary is established as a comparison basis to generate a unique summary value for the core field. Data fields with the same summary value are retrieved and located in the dispersed storage associated field blocks, and physical reorganization technology (column storage reorganization) is used to integrate and compress the data packages in the pre-storage layer of the medical data package.
[0016] Compare the remaining space in the pre-storage layer of the compressed medical data package with the volume of the medical data package. If the remaining space in the pre-storage layer is less than the volume of the medical data package, calculate the storage value of the data package in the pre-storage layer using a dynamic scoring model, and upgrade the scoring model based on the storage layer:
[0017]
[0018] The storage level of the medical data package and the data package in the pre-storage level is evaluated for upgrade, where F represents the medical data package storage level upgrade score, ε, μ, and ρ are coefficients, and ρ is a negative number; the storage level of the data package whose storage level upgrade score exceeds the preset upgrade score threshold is upgraded until the remaining space in the pre-storage level of the medical data package exceeds the size of the medical data package;
[0019] If the highest storage level upgrade score in the data package is lower than the preset upgrade score threshold, the data packages in the pre-storage level will be replaced with data packages with a storage value lower than the medical data package storage value according to the storage value from low to high, and the storage level downgrade process will be performed first until the remaining pre-storage level exceeds the size of the medical data package.
[0020] Furthermore, the access frequency of the medical data packet is predicted by the following method:
[0021] Compare the preset medical data types with the fields in the medical data package for similarity, determine the data types corresponding to the fields in the medical data package, and calculate the ratio of the number of fields of different medical data types to the total number of fields in the medical data package; different medical data fields correspond to different medical data types;
[0022] The access frequency of historical medical data types is obtained, and for the medical data packet, the access frequency of the medical data packet is obtained by weighted accumulation according to the proportion of the medical data field types possessed and the corresponding access frequency.
[0023] Furthermore, the specific method for determining whether the query response time is compliant based on the actual query response delay time of the medical data packet is as follows:
[0024] The actual query response delay time of the medical data packet is calculated by performing a weighted product calculation based on the level difference between the actual storage level and the pre-storage level of the medical data packet, the volume of the medical data packet, and the ratio of the priority response ranking of the medical data packet in the actual query volume; the query scenario and response compliance time of the data type corresponding to the field in the medical data packet are determined by using the query scenario and response compliance time of the historical medical data type in the platform;
[0025] Obtain the shortest query response compliance time for the data type corresponding to the field in the medical data packet. Utilizing the positive relationship between storage tiers and data query response speeds and the actual query response delay time of the medical data packet, the actual query response time of the medical data packet is obtained. When the actual query response time of the medical data packet exceeds the shortest query response compliance time, the query response time of the medical data is determined to be unqualified.
[0026] The ratio of the priority response ranking of the medical data packet in the actual query volume includes:
[0027] Based on the access frequency of the medical data packet, the platform intercepts the corresponding historical time period, selects the actual query volume closest to the average query volume in the corresponding historical time period, and uses it as the benchmark reference for the medical data packet query volume on the day of the medical data packet query; according to the formula: Determine the priority response value of the medical data packet and the query benchmark with reference to the medical data, and sort the data packets according to the priority response value, where Y is the priority response value of the medical data packet.
[0028] Furthermore, the access frequency prediction of the medical data package is based on the premise that the data quality of the medical data package is qualified. The specific methods for analyzing the data quality of the medical data package include:
[0029] Perform anomaly detection on the fields in the medical data packet to be entered, including:
[0030] Obtaining a related field of each field in the medical data packet by using a medical data model for representing the relationship between medical data;
[0031] The indicator function "is null" is used to mark missing fields and abnormal fields in the medical data packet, and the abnormality rate of the medical data packet is calculated by calculating the proportion of the marked fields to the total number of fields in the medical data packet. The missing fields are the missing required fields and associated fields in the medical data packet, and the abnormal fields are the blank fields in the medical data packet whose field values are not zero.
[0032] When the abnormality rate of the medical data packet exceeds the abnormality threshold, the platform determines that the data quality of the medical data packet collected by the self-collection device is unqualified.
[0033] A monitoring and analysis system for biomedical data, comprising a data quality detection module, an access frequency analysis module, a query response compliance analysis module, and a storage resource optimization module;
[0034] The data quality detection module is configured to detect anomalies in fields in a medical data packet, including: obtaining associated fields for each field in the medical data packet using a medical data model for characterizing associations between medical data; marking missing fields and abnormal fields in the medical data packet using an indicator function, and calculating an anomaly rate of the medical data packet based on the proportion of the marked fields to the total number of fields in the medical data packet; and determining that the data quality of the medical data packet obtained by the platform from the collection device is unqualified when the anomaly rate of the medical data packet exceeds an anomaly threshold;
[0035] The access frequency analysis module is used to compare the similarity between the preset medical data types and the fields in the medical data packet based on the medical data types in the platform, determine the data types corresponding to the fields in the medical data packet, and calculate the ratio of the number of fields of different medical data types to the total number of fields in the medical data packet; obtain the access frequency of the historical medical data types, and for the medical data packet, perform weighted accumulation according to the ratio of the medical data field types and the corresponding access frequencies to obtain the access frequency of the medical data packet;
[0036] The query response compliance analysis module is configured to score the storage value of the medical data package using a dynamic scoring model and determine the pre-storage level of the medical data package according to the scoring result; calculate the actual query response delay time of the medical data package by weighted product based on the level difference between the actual storage level of the medical data package and the pre-storage level and the volume of the medical data package; and obtain the actual query response time of the medical data package by utilizing the positive relationship between the storage level and the response speed of the data query and the actual query response delay time of the medical data package;
[0037] Using the query scenarios and response compliance times of historical medical data types in the platform, determine the query scenarios and query response compliance times of the data types corresponding to the fields in the medical data package; compare the actual query response time of the medical data package with the shortest query response compliance time in the medical data package. When the actual query response time of the medical data package exceeds the shortest query response compliance time, the query response time of the medical data is determined to be unqualified;
[0038] The storage resource optimization module is configured to use the core data fields of data objects in the pre-storage layer of the medical data packet as a benchmark anchor for deduplication and integration, establish a field hash summary as a comparison basis, locate the dispersed storage associated field blocks of data fields with the same unique summary value through an inverted index, integrate and compress the data fields in the pre-storage layer of the medical data packet using a physical reorganization technology, and use a storage level upgrade scoring model to evaluate the storage level of the medical data packet and the data packets in the pre-storage layer. The storage level of the data corresponding to the data packet storage level upgrade score exceeding a preset upgrade score threshold is upgraded until the remaining space in the pre-storage layer of the medical data packet exceeds the size of the medical data packet.
[0039] If the highest storage level upgrade score in the data package is lower than the preset upgrade score threshold, the data packages in the pre-storage level will be replaced with data packages with a storage value lower than the medical data package storage value according to the storage value from low to high, and the storage level downgrade process will be performed first until the remaining pre-storage level exceeds the size of the medical data package.
[0040] Furthermore, the query response compliance analysis module includes a storage value scoring unit, a pre-storage level determination unit, and a response time compliance determination unit;
[0041] The storage value scoring unit is configured to score the storage value of the medical data packet using a dynamic scoring model; the pre-storage level determination unit is configured to determine the pre-storage level of the medical data packet by dividing the scoring area according to different scoring thresholds, wherein the storage level of the medical data packet and its dynamic score present a strict positive hierarchical mapping relationship; the response time compliance determination unit is configured to calculate the actual query response delay time of the medical data packet by weighted product calculation based on the level difference between the actual storage level of the medical data packet and the pre-storage level and the volume of the medical data packet; and to obtain the actual query response time of the medical data packet based on the positive relationship between the storage level and the response speed of the data query;
[0042] Utilize the query scenarios and response compliance times of historical medical data types in the platform to determine the query scenarios and query response compliance times of the data types corresponding to the fields in the medical data package; compare the actual query response time of the medical data package with the shortest query response compliance time in the medical data package to determine whether the query response time of the medical data package is compliant.
[0043] Furthermore, the storage resource optimization module includes a data field integration unit, a storage level adjustment unit, and a data replacement unit;
[0044] The data field integration unit is configured to extract core metadata fields of data objects in the pre-storage layer of the medical data packet as reference anchors for deduplication and integration, establish a field hash digest, locate associated field blocks in dispersed storage using an inverted index for data fields with the same unique digest value, and integrate and compress the data fields in the pre-storage layer of the medical data packet using a physical reorganization technique;
[0045] The storage level adjustment unit is configured to evaluate the storage level of the medical data package and the data package in the pre-storage level using the storage level upgrade scoring model, and upgrade the storage level of the data package whose storage level upgrade score exceeds a preset upgrade score threshold until the remaining space in the pre-storage level of the medical data package exceeds the size of the medical data package;
[0046] The data replacement unit is used to replace the data packets with storage values lower than the storage value of the medical data packets according to the storage value of the data packets in the pre-storage layer when the highest storage level upgrade score in the data packet is lower than the preset upgrade score threshold, and prioritize the storage level downgrade processing according to the storage value from low to high until the remaining pre-storage layer exceeds the size of the medical data packet.
[0047] Compared with the existing technology, the beneficial effects achieved by the present invention are: the present invention optimizes storage resources by performing data quality detection and query response time analysis. The optimization can identify and eliminate unnecessary waste of storage resources, thereby reducing the storage cost of the platform; and when the data query response time is not compliant when the medical data packet is entered into the platform, the storage resources in the platform are optimized in a targeted manner to avoid frequent data migration leading to data query response time delays, which will trigger a series of chain reactions, from clinical decision delays to systemic risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0049] Figure 1 This is a flowchart of the steps of a method for monitoring and analyzing biomedical data of the present invention. DETAILED DESCRIPTION
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0051] See also Figure 1 The present invention provides a technical solution: a method for monitoring and analyzing biomedical data, which specifically includes the following steps:
[0052] Obtaining a medical data package to be stored, and scoring the storage value of the medical data package based on the access frequency, clinical criticality score, and package volume of the medical data package in combination with a dynamic scoring model;
[0053] Determining a pre-storage level for the medical data packet based on the storage value scoring result, wherein the pre-storage level is a target storage location for the medical data packet;
[0054] determining an actual storage level of the medical data packet based on real-time storage space, and analyzing an actual query response delay time of the medical data packet based on a level difference between the pre-storage level and the actual storage level;
[0055] Determine whether the query response time is compliant based on the actual query response delay time of the medical data packet;
[0056] When it is determined that the query response time of the medical data package is not in compliance, the storage resources in the platform are optimized, including the upgrade processing of the data package storage layer and the replacement candidates of the medical data package pre-storage layer data.
[0057] The specific method of scoring the storage value of the medical data package in combination with the dynamic scoring model and determining the pre-storage level of the medical data package is as follows:
[0058] According to the dynamic scoring model: Score the storage value of the medical data package, where f represents the dynamic score of the medical data package, p represents the access frequency of the medical data package, and p max α represents the highest access frequency of all data in the platform, m represents the clinical criticality score of the medical data package, which is the data importance assessed by clinical experts; M represents the total clinical criticality score, and g represents the data volume of the medical data package; α represents the access frequency weight coefficient, which is used to adjust the impact of access frequency on the total score. For example, if a field is accessed 100 times a day and the system has a maximum access of 1000 times, the access frequency weight coefficient is 0.1; β represents the clinical criticality weight coefficient, which is used to adjust the impact of clinical importance on the total score. For example, if the criticality of "intraoperative real-time blood pressure" is 9 / 10, as scored by clinical experts, the clinical criticality weight coefficient is 0.9; γ represents the data volume weight coefficient, which is used to adjust the negative impact of data size on the total score.
[0059] The pre-storage level of the medical data packet is determined according to the dynamic score, wherein the dynamic score and the pre-storage level are in a positive hierarchical mapping relationship, that is, the higher the score, the higher the storage level; the medical data storage platform includes multiple storage levels, each storage level is a logical subset of the storage level above it, and the storage level is positively correlated with the impact speed of data query.
[0060] Furthermore, the specific method for optimizing the storage resources in the platform is as follows:
[0061] Since it is impossible to determine whether the data in the pre-storage layer of the medical data package is fragmentary and duplicated, the core metadata fields of the data package in the pre-storage layer of the medical data package are extracted as the benchmark anchor for deduplication and integration. The core data includes unique identifiers, timestamp sequences and clinical event codes. Fragmentary duplicate entry refers to the phenomenon that the same or highly similar data fragments are entered into the system multiple times during the data entry process, resulting in data redundancy, inconsistency or errors. A field hash summary is established as a comparison basis to generate a unique summary value for the core field. Data fields with the same summary value are retrieved and located in the dispersed storage associated field blocks, and physical reorganization technology (column storage reorganization) is used to integrate and compress the data packages in the pre-storage layer of the medical data package.
[0062] Compare the remaining space in the pre-storage layer of the compressed medical data package with the volume of the medical data package. If the remaining space in the pre-storage layer is less than the volume of the medical data package, calculate the storage value of the data package in the pre-storage layer using a dynamic scoring model, and upgrade the scoring model based on the storage layer:
[0063]
[0064] The storage level of the medical data package and the data package in the pre-storage level is evaluated for upgrade, where F represents the medical data package storage level upgrade score, ε, μ, and ρ are coefficients, and ρ is a negative number; the storage level of the data package whose storage level upgrade score exceeds the preset upgrade score threshold is upgraded until the remaining space in the pre-storage level of the medical data package exceeds the size of the medical data package;
[0065] If the highest storage level upgrade score in the data package is lower than the preset upgrade score threshold, the data packages in the pre-storage level will be replaced with data packages with a storage value lower than the medical data package storage value according to the storage value from low to high, and the storage level downgrade process will be performed first until the remaining pre-storage level exceeds the size of the medical data package.
[0066] Furthermore, the access frequency of the medical data packet is predicted by the following method:
[0067] Compare the preset medical data types with the fields in the medical data package for similarity, determine the data types corresponding to the fields in the medical data package, and calculate the ratio of the number of fields of different medical data types to the total number of fields in the medical data package; different medical data fields correspond to different medical data types;
[0068] The access frequency of historical medical data types is obtained, and for the medical data packet, the access frequency of the medical data packet is obtained by weighted accumulation according to the proportion of the medical data field types possessed and the corresponding access frequency.
[0069] Furthermore, the specific method for determining whether the query response time is compliant based on the actual query response delay time of the medical data packet is as follows:
[0070] The actual query response delay time of the medical data packet is calculated by performing a weighted product calculation based on the level difference between the actual storage level and the pre-storage level of the medical data packet, the volume of the medical data packet, and the ratio of the priority response ranking of the medical data packet in the actual query volume; the query scenario and response compliance time of the data type corresponding to the field in the medical data packet are determined by using the query scenario and response compliance time of the historical medical data type in the platform;
[0071] Obtain the shortest query response compliance time for the data type corresponding to the field in the medical data packet. Utilizing the positive relationship between storage tiers and data query response speeds and the actual query response delay time of the medical data packet, the actual query response time of the medical data packet is obtained. When the actual query response time of the medical data packet exceeds the shortest query response compliance time, the query response time of the medical data is determined to be unqualified.
[0072] The ratio of the priority response ranking of the medical data packet in the actual query volume includes:
[0073] Based on the access frequency of the medical data packet, the platform intercepts the corresponding historical time period, selects the actual query volume closest to the average query volume in the corresponding historical time period, and uses it as the benchmark reference for the medical data packet query volume on the day of the medical data packet query; according to the formula: Determine the priority response value of the medical data packet and the query benchmark with reference to the medical data, and sort the data packets according to the priority response value, where Y is the priority response value of the medical data packet.
[0074] Furthermore, the access frequency prediction of the medical data package is based on the premise that the data quality of the medical data package is qualified. The specific methods for analyzing the data quality of the medical data package include:
[0075] Perform anomaly detection on the fields in the medical data packet to be entered, including:
[0076] Obtaining a related field of each field in the medical data packet by using a medical data model for representing the relationship between medical data;
[0077] The indicator function "is null" is used to mark missing fields and abnormal fields in the medical data packet, and the abnormality rate of the medical data packet is calculated by calculating the proportion of the marked fields to the total number of fields in the medical data packet. The missing fields are the missing required fields and associated fields in the medical data packet, and the abnormal fields are the blank fields in the medical data packet whose field values are not zero.
[0078] When the abnormality rate of the medical data packet exceeds the abnormality threshold, the platform determines that the data quality of the medical data packet collected by the self-collection device is unqualified.
[0079] A monitoring and analysis system for biomedical data, comprising a data quality detection module, an access frequency analysis module, a query response compliance analysis module, and a storage resource optimization module;
[0080] The data quality detection module is configured to detect anomalies in fields in a medical data packet, including: obtaining associated fields for each field in the medical data packet using a medical data model for characterizing associations between medical data; marking missing fields and abnormal fields in the medical data packet using an indicator function, and calculating an anomaly rate of the medical data packet based on the proportion of the marked fields to the total number of fields in the medical data packet; and determining that the data quality of the medical data packet obtained by the platform from the collection device is unqualified when the anomaly rate of the medical data packet exceeds an anomaly threshold;
[0081] The access frequency analysis module is used to compare the similarity between the preset medical data types and the fields in the medical data packet based on the medical data types in the platform, determine the data types corresponding to the fields in the medical data packet, and calculate the ratio of the number of fields of different medical data types to the total number of fields in the medical data packet; obtain the access frequency of the historical medical data types, and for the medical data packet, perform weighted accumulation according to the ratio of the medical data field types and the corresponding access frequencies to obtain the access frequency of the medical data packet;
[0082] The query response compliance analysis module is configured to score the storage value of the medical data package using a dynamic scoring model and determine the pre-storage level of the medical data package according to the scoring result; calculate the actual query response delay time of the medical data package by weighted product based on the level difference between the actual storage level of the medical data package and the pre-storage level and the volume of the medical data package; and obtain the actual query response time of the medical data package by utilizing the positive relationship between the storage level and the response speed of the data query and the actual query response delay time of the medical data package;
[0083] Using the query scenarios and response compliance times of historical medical data types in the platform, determine the query scenarios and query response compliance times of the data types corresponding to the fields in the medical data package; compare the actual query response time of the medical data package with the shortest query response compliance time in the medical data package. When the actual query response time of the medical data package exceeds the shortest query response compliance time, the query response time of the medical data is determined to be unqualified;
[0084] The storage resource optimization module is configured to use the core data fields of data objects in the pre-storage layer of the medical data packet as a benchmark anchor for deduplication and integration, establish a field hash summary as a comparison basis, locate the dispersed storage associated field blocks of data fields with the same unique summary value through an inverted index, integrate and compress the data fields in the pre-storage layer of the medical data packet using a physical reorganization technology, and use a storage level upgrade scoring model to evaluate the storage level of the medical data packet and the data packets in the pre-storage layer. The storage level of the data corresponding to the data packet storage level upgrade score exceeding a preset upgrade score threshold is upgraded until the remaining space in the pre-storage layer of the medical data packet exceeds the size of the medical data packet.
[0085] If the highest storage level upgrade score in the data package is lower than the preset upgrade score threshold, the data packages in the pre-storage level will be replaced with data packages with a storage value lower than the medical data package storage value according to the storage value from low to high, and the storage level downgrade process will be performed first until the remaining pre-storage level exceeds the size of the medical data package.
[0086] Furthermore, the query response compliance analysis module includes a storage value scoring unit, a pre-storage level determination unit, and a response time compliance determination unit;
[0087] The storage value scoring unit is configured to score the storage value of the medical data packet using a dynamic scoring model; the pre-storage level determination unit is configured to determine the pre-storage level of the medical data packet by dividing the scoring area according to different scoring thresholds, wherein the storage level of the medical data packet and its dynamic score present a strict positive hierarchical mapping relationship; the response time compliance determination unit is configured to calculate the actual query response delay time of the medical data packet by weighted product calculation based on the level difference between the actual storage level of the medical data packet and the pre-storage level and the volume of the medical data packet; and to obtain the actual query response time of the medical data packet based on the positive relationship between the storage level and the response speed of the data query;
[0088] Utilize the query scenarios and response compliance times of historical medical data types in the platform to determine the query scenarios and query response compliance times of the data types corresponding to the fields in the medical data package; compare the actual query response time of the medical data package with the shortest query response compliance time in the medical data package to determine whether the query response time of the medical data package is compliant.
[0089] Furthermore, the storage resource optimization module includes a data field integration unit, a storage level adjustment unit, and a data replacement unit;
[0090] The data field integration unit is configured to extract core metadata fields of data objects in the pre-storage layer of the medical data packet as reference anchors for deduplication and integration, establish a field hash digest, locate associated field blocks in dispersed storage using an inverted index for data fields with the same unique digest value, and integrate and compress the data fields in the pre-storage layer of the medical data packet using a physical reorganization technique;
[0091] The storage level adjustment unit is configured to evaluate the storage level of the medical data package and the data package in the pre-storage level using the storage level upgrade scoring model, and upgrade the storage level of the data package whose storage level upgrade score exceeds a preset upgrade score threshold until the remaining space in the pre-storage level of the medical data package exceeds the size of the medical data package;
[0092] The data replacement unit is used to replace the data packets with storage values lower than the storage value of the medical data packets according to the storage value of the data packets in the pre-storage layer when the highest storage level upgrade score in the data packet is lower than the preset upgrade score threshold, and prioritize the storage level downgrade processing according to the storage value from low to high until the remaining pre-storage layer exceeds the size of the medical data packet.
[0093] In this embodiment:
[0094] Through the collaboration of multiple technologies, the platform achieves comprehensive coverage of medical data. It obtains medical data packets collected by self-collection devices and performs rule verification on the data in the medical data packets, including:
[0095] Obtaining a related field of each field in the medical data packet by using a medical data model for representing the relationship between medical data;
[0096] Use the indicator function is null to mark missing fields and abnormal fields in the medical data packet, where missing fields refer to the missing required fields and associated fields in the medical data packet, and abnormal fields refer to blank fields in the medical data packet whose field value is not zero;
[0097] By calculating the proportion of the marked fields in the total number of fields in the medical data packet, the abnormality rate of the medical data packet is obtained as q = (a / A) = 0.35, where a and A represent the number of marked fields and the total number of fields in the medical data packet, respectively; q is less than the abnormality threshold Q = 0.4, and it is determined that the data quality of the medical data packet obtained by the platform from the collection device is qualified.
[0098] The data quality inspection of the medical data packet is qualified. The similarity comparison is performed between the medical data types in the platform and the required fields and optional fields in the medical data packet. The data types corresponding to the required fields and optional fields in the medical data packet are determined based on the maximum similarity value. Based on the different access frequencies corresponding to the medical data types in the platform, the access frequency of different fields in the medical data packet in the platform is determined. The access frequency of the medical data packet is quantified using the weighted accumulation method based on the proportion of different fields to the total fields in the medical data packet.
[0099] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0100] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for monitoring and analyzing biomedical data, applied to a medical data storage platform, characterized by: Obtaining a medical data package to be stored, and scoring the storage value of the medical data package based on the access frequency, clinical criticality score, and package volume of the medical data package in combination with a dynamic scoring model; Determining a pre-storage level for the medical data packet based on the storage value scoring result, wherein the pre-storage level is a target storage location for the medical data packet; determining an actual storage level of the medical data packet based on real-time storage space, and analyzing an actual query response delay time of the medical data packet based on a level difference between the pre-storage level and the actual storage level; Determine whether the query response time is compliant based on the actual query response delay time of the medical data packet; When it is determined that the query response time of the medical data package is not in compliance, the storage resources in the platform are optimized, including the upgrade processing of the data package storage layer and the replacement candidates of the medical data package pre-storage layer data.
2. The method for monitoring and analyzing biomedical data according to claim 1, wherein: The specific method of scoring the storage value of the medical data package in combination with the dynamic scoring model and determining the pre-storage level of the medical data package is as follows: According to the dynamic scoring model: Score the storage value of the medical data package, where f represents the dynamic score of the medical data package, p represents the access frequency of the medical data package, and p max It is represented by the highest access frequency of all data in the platform, m is the score of clinical criticality of the medical data package, M is the total score of clinical criticality, and g is the data volume of the medical data package; α is the access frequency weight coefficient, β is the clinical criticality weight coefficient, and γ is the data volume weight coefficient, where α and β>0>γ; The pre-storage level of the medical data packet is determined according to the dynamic score, wherein the dynamic score and the pre-storage level are in a positive hierarchical mapping relationship; the medical data storage platform includes multiple storage levels, each storage level is a logical subset of the storage level above it, and the storage level is positively correlated with the impact speed of data query.
3. The method for monitoring and analyzing biomedical data according to claim 2, wherein: The specific method for optimizing storage resources in the platform is as follows: By extracting the core metadata fields of the medical data package in the pre-storage layer as the benchmark anchor for deduplication and integration, the core data includes unique identifiers, timestamp sequences, and clinical event codes; Establish a field hash summary as a comparison basis to generate a unique summary value for the core field; locate the associated field blocks stored in a dispersed manner by searching for data fields with the same summary value, and use physical reorganization technology to integrate and compress the data packets at the pre-storage level of the medical data packet; Compare the remaining space in the pre-storage layer of the compressed medical data package with the volume of the medical data package. If the remaining space in the pre-storage layer is less than the volume of the medical data package, calculate the storage value of the data package in the pre-storage layer using a dynamic scoring model, and upgrade the scoring model based on the storage layer: The storage level of medical data packets and data packets in the pre-storage level is upgraded and evaluated. F represents the medical data packet storage level upgrade score, ε, μ, and ρ are coefficients, and ε, μ>0>ρ; Upgrading the storage level of the data whose data package storage level upgrade score exceeds the preset upgrade score threshold until the remaining space of the medical data package pre-storage level exceeds the size of the medical data package; If the highest storage level upgrade score in the data package is lower than the preset upgrade score threshold, the data packages in the pre-storage level will be replaced with data packages with a storage value lower than the medical data package storage value according to the storage value from low to high, and the storage level downgrade process will be performed first until the remaining pre-storage level exceeds the size of the medical data package.
4. The method for monitoring and analyzing biomedical data according to claim 1, wherein: The access frequency of the medical data packet is predicted by the following method: Compare the preset medical data types with the fields in the medical data package for similarity, determine the data types corresponding to the fields in the medical data package, and calculate the ratio of the number of fields of different medical data types to the total number of fields in the medical data package; different medical data fields correspond to different medical data types; The access frequency of historical medical data types is obtained, and for the medical data packet, the access frequency of the medical data packet is obtained by weighted accumulation according to the proportion of the medical data field types possessed and the corresponding access frequency.
5. The method for monitoring and analyzing biomedical data according to claim 1, wherein: The specific method for determining whether the query response time is compliant based on the actual query response delay time of the medical data packet is as follows: The actual query response delay time of the medical data packet is calculated by weighted product calculation based on the level difference between the actual storage level and the pre-storage level of the medical data packet and the volume of the medical data packet; the query scenario and response compliance time of the data type corresponding to the field in the medical data packet are determined by using the query scenario and response compliance time of the historical medical data type in the platform; Obtain the shortest query response compliance time for the data type corresponding to the field in the medical data packet. By utilizing the positive relationship between the storage level and the response speed of data query and the actual query response delay time of the medical data packet, the actual query response time of the medical data packet is obtained. When the actual query response time of the medical data packet exceeds the shortest query response compliance time, the query response time of the medical data is determined to be unqualified.
6. The method for monitoring and analyzing biomedical data according to claim 4, wherein: The premise for predicting the access frequency of medical data packages is that the data quality of the medical data packages is qualified. The specific methods for analyzing the data quality of medical data packages include: Perform anomaly detection on the fields in the medical data packet to be entered, including: Obtaining a related field of each field in the medical data packet by using a medical data model for representing the relationship between medical data; Using an indicator function to mark missing fields and abnormal fields in a medical data packet, the abnormality rate of the medical data packet is obtained by calculating the proportion of the marked fields to the total number of fields in the medical data packet; the missing fields are the missing required fields and associated fields in the medical data packet, and the abnormal fields are the blank fields in the medical data packet whose field values are not zero; When the abnormality rate of the medical data packet exceeds the abnormality threshold, the platform determines that the data quality of the medical data packet collected by the self-collection device is unqualified.
7. A biomedical data monitoring and analysis system, characterized by: The monitoring and analysis system includes a data quality detection module, an access frequency analysis module, a query response compliance analysis module, and a storage resource optimization module; The data quality detection module is configured to detect anomalies in fields in a medical data packet, including: obtaining associated fields for each field in the medical data packet using a medical data model for characterizing associations between medical data; marking missing fields and abnormal fields in the medical data packet using an indicator function, and calculating an anomaly rate of the medical data packet based on the proportion of the marked fields to the total number of fields in the medical data packet; and determining that the data quality of the medical data packet obtained by the platform from the collection device is unqualified when the anomaly rate of the medical data packet exceeds an anomaly threshold; The access frequency analysis module is used to compare the similarity between the preset medical data types and the fields in the medical data packet based on the medical data types in the platform, determine the data types corresponding to the fields in the medical data packet, and calculate the ratio of the number of fields of different medical data types to the total number of fields in the medical data packet; obtain the access frequency of the historical medical data types, and for the medical data packet, perform weighted accumulation according to the ratio of the medical data field types and the corresponding access frequencies to obtain the access frequency of the medical data packet; The query response compliance analysis module is configured to score the storage value of the medical data package using a dynamic scoring model and determine the pre-storage level of the medical data package according to the scoring result; calculate the actual query response delay time of the medical data package by weighted product based on the level difference between the actual storage level of the medical data package and the pre-storage level and the volume of the medical data package; and obtain the actual query response time of the medical data package by utilizing the positive relationship between the storage level and the response speed of the data query and the actual query response delay time of the medical data package; Using the query scenarios and response compliance times of historical medical data types in the platform, determine the query scenarios and query response compliance times of the data types corresponding to the fields in the medical data package; compare the actual query response time of the medical data package with the shortest query response compliance time in the medical data package. When the actual query response time of the medical data package exceeds the shortest query response compliance time, the query response time of the medical data is determined to be unqualified; The storage resource optimization module is configured to use the core data fields of data objects in the pre-storage layer of the medical data packet as a benchmark anchor for deduplication and integration, establish a field hash summary as a comparison basis, locate the dispersed storage associated field blocks of data fields with the same unique summary value through an inverted index, integrate and compress the data fields in the pre-storage layer of the medical data packet using a physical reorganization technology, and use a storage level upgrade scoring model to evaluate the storage level of the medical data packet and the data packets in the pre-storage layer. The storage level of the data corresponding to the data packet storage level upgrade score exceeding a preset upgrade score threshold is upgraded until the remaining space in the pre-storage layer of the medical data packet exceeds the size of the medical data packet. If the highest storage level upgrade score in the data package is lower than the preset upgrade score threshold, the data packages in the pre-storage level will be replaced with data packages with a storage value lower than the medical data package storage value according to the storage value from low to high, and the storage level downgrade process will be performed first until the remaining pre-storage level exceeds the size of the medical data package.
8. The biomedical data monitoring and analysis system according to claim 7, characterized in that: The query response compliance analysis module includes a storage value scoring unit, a pre-storage level determination unit, and a response time compliance determination unit; The storage value scoring unit is used to score the storage value of the medical data package using a dynamic scoring model; the pre-storage level determination unit is used to determine the pre-storage level of the medical data package by dividing the scoring area according to different scoring thresholds, wherein the storage level of the medical data package and its dynamic score present a strict positive hierarchical mapping relationship; The response time compliance determination unit is configured to calculate the actual query response delay time of the medical data packet by weighted product calculation based on the level difference between the actual storage level and the pre-storage level of the medical data packet and the volume of the medical data packet; The actual query response time of medical data packets is obtained through the positive relationship between storage level and data query response speed; Using the query scenarios and response compliance times of historical medical data types in the platform, determine the query scenarios and query response compliance times for the data types corresponding to the fields in the medical data package; The actual query response time of the medical data package is compared with the shortest query response compliance time in the medical data package to determine whether the query response time of the medical data package is compliant.
9. The biomedical data monitoring and analysis system according to claim 8, characterized in that: The storage resource optimization module includes a data field integration unit, a storage level adjustment unit and a data replacement unit; The data field integration unit is configured to extract core metadata fields of data objects in the pre-storage layer of the medical data packet as reference anchors for deduplication and integration, establish a field hash digest, locate associated field blocks in dispersed storage using an inverted index for data fields with the same unique digest value, and integrate and compress the data fields in the pre-storage layer of the medical data packet using a physical reorganization technique; The storage level adjustment unit is configured to evaluate the storage level of the medical data package and the data package in the pre-storage level using the storage level upgrade scoring model, and upgrade the storage level of the data package whose storage level upgrade score exceeds a preset upgrade score threshold until the remaining space in the pre-storage level of the medical data package exceeds the size of the medical data package; The data replacement unit is used to replace the data packets with storage values lower than the storage value of the medical data packets according to the storage value of the data packets in the pre-storage layer when the highest storage level upgrade score in the data packet is lower than the preset upgrade score threshold, and prioritize the storage level downgrade processing according to the storage value from low to high until the remaining pre-storage layer exceeds the size of the medical data packet.