Data management system and method for medical affairs
By constructing a joint feature matrix and a multivariate fitting algorithm, the problem of unified monitoring and management of data collection tasks in medical consortia was solved, enabling accurate quantitative assessment and trend early warning of data quality, improving the integrity and timeliness of data collection, and reducing the need for manual intervention.
Patent Information
- Application Number
- CN202511455600.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-13
AI Technical Summary
In medical consortia, existing technologies are insufficient for unified monitoring and management of data collection tasks, leading to problems such as task execution failure, node disconnection, and delayed data transmission, which affect the integrity and timeliness of data collection, and lack automated and visualized management mechanisms.
By acquiring data collection task logs and server node operation status logs from each target medical institution, statistical data quality and operation status indicators are collected, a joint feature matrix is constructed, multivariate fitting algorithms are used to capture trend changes, trend early warning information is generated, and task management functions, including start/stop control and priority adjustment, are provided.
It enables precise quantitative assessment of data quality within the medical consortium, improves the real-time performance and accuracy of abnormal trend detection, reduces the latency and subjective error of manual analysis, and enhances the sensitivity and rapid response capability to data quality anomalies.
Smart Images

Figure CN120932792A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data management technology, and more specifically, to a data management system and method for medical consortia. Background Technology
[0002] The demand for data integration and sharing among multiple medical institutions within the region is constantly growing. To meet the requirements of unified management and data interoperability, data acquisition server nodes are typically deployed locally at each member unit. Structured data from the business systems is periodically transmitted to the medical consortium's data center via ETL (Extract, Transform, Load) tasks. However, due to the large number of acquisition nodes, the complexity of task types, and the significant differences in operating environments, traditional manual monitoring methods struggle to promptly grasp the operational status of each node and its tasks. This often leads to problems such as task execution failures, node disconnections, and delayed data transmission, severely impacting the integrity and timeliness of data acquisition.
[0003] Current technologies typically focus only on task completion, lacking systematic monitoring and coordinated management of server operating status and task execution processes. Especially in actual medical consortium deployments, a single server often runs multiple data collection tasks, and task status anomalies are often highly correlated with node resource anomalies, such as CPU overload, memory overflow, and network interruptions. Without coordinated monitoring and management of nodes and tasks, troubleshooting efficiency will be low, task failure rates will increase, and data interruption time will be prolonged. Furthermore, some systems still rely on manual methods for starting and stopping tasks and handling anomalies, lacking a unified task scheduling and management mechanism, making it difficult to meet the visualization, automation, and stabilization requirements of the data collection process in large-scale medical consortium scenarios.
[0004] Therefore, there is an urgent need for a data management method for medical consortia that can achieve full-process monitoring of data collection and exchange under a unified architecture. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this application provides a data management system and method for medical consortia.
[0006] Firstly, this application provides a data management method for medical consortia, including:
[0007] According to the first preset cycle, the data collection task logs of each target medical institution and the corresponding server node operation status logs are obtained; the server node operation status logs include: node identifier, node heartbeat information, node resource occupancy rate and node availability status;
[0008] Based on the data acquisition task log and the server node's operation status log, the data quality indicators and server node operation status indicators of each target medical institution are statistically analyzed, and time-series data of each data quality indicator and operation status indicator are generated according to a second preset period; the operation status indicators include: node failure rate and abnormal resource occupancy rate.
[0009] For the same data quality indicator, the time series data corresponding to each target medical institution are stacked according to the institution dimension to construct a joint feature matrix, and a multivariate fitting algorithm is executed based on the joint feature matrix to obtain the trend change parameters of each target medical institution.
[0010] Based on the trend change parameters and the operating status indicators of the server nodes, a degradation trend detection result is generated according to preset judgment conditions; based on the degradation trend detection result, a trend warning information is generated and the trend warning information is output to the management terminal.
[0011] As an optional implementation, the management terminal is further used for:
[0012] It displays the real-time running status information of the corresponding server nodes of each target medical institution and the execution status information of the data collection tasks;
[0013] The operational status information includes: node availability status, node resource utilization rate, and node fault indication; the execution status information includes the current task status, task execution progress, and task execution result.
[0014] The management terminal is also used to: provide task management functions based on the execution status information of the data acquisition task and the running status information of the server node; the task management functions include task start / stop control, task priority adjustment and task retry strategy management.
[0015] As an optional implementation, the data quality indicators include: data missing rate, format error rate, and data acquisition delay; the data acquisition task log includes: acquisition task identifier, acquisition time, number of acquisition items, acquisition results, and acquisition status.
[0016] Based on the aforementioned trend change parameters and preset judgment conditions, a degradation trend detection result is generated, including:
[0017] Based on the trend change parameters of each target medical institution, generate corresponding single-institution judgment result information;
[0018] The single-institution judgment result information is used as multi-source evidence information input, and evidence theory is used for synthesis processing to obtain a comprehensive confidence result for the overall data collection situation;
[0019] The degree of conflict of the comprehensive confidence result is calculated, and when the degree of conflict exceeds a preset range, the weight of the conflicting evidence information is adjusted to obtain the corrected comprehensive confidence result.
[0020] By comparing the corrected overall confidence result with the preset confidence threshold, the abnormal confidence level at the medical consortium level is determined;
[0021] The degradation trend detection results are generated based on the abnormal confidence level and the single-institution judgment results of each target medical institution.
[0022] As an optional implementation, obtaining a comprehensive confidence result for the overall data collection situation includes:
[0023] Based on the individual assessment results of each target medical institution, an initial basic confidence allocation is constructed;
[0024] Calculate the degree of conflict between the initial basic confidence assignments, and when the degree of conflict exceeds a preset range, adjust the weight of conflict evidence information to update the basic confidence assignments;
[0025] An iterative synthesis process is performed on the updated basic confidence assignment to obtain a first synthesized confidence level and a second synthesized confidence level; wherein, the first synthesized confidence level is used to characterize that the data acquisition quality is in a deteriorating trend, and the second synthesized confidence level is used to characterize that the data acquisition quality is in a normal trend.
[0026] Based on the first composite confidence level and the second composite confidence level, the uncertainty entropy value is determined;
[0027] Based on the first composite confidence level, the second composite confidence level, and the uncertainty entropy value, a comprehensive confidence result is generated;
[0028] The uncertainty entropy value is used to characterize the degree of consistency of the overall evidence of the medical consortium.
[0029] As an optional implementation, it also includes:
[0030] A geographic proximity matrix is generated based on the geographic location information between the target medical institutions.
[0031] Based on the geographic proximity matrix, the single-institution determination results of each target medical institution are subjected to proximity-based weighted processing.
[0032] The weighted single-institution determination results are used to generate the initial basic confidence assignment.
[0033] The generation of the geographic proximity matrix based on the geographic location information between each target medical institution includes:
[0034] Geographic parsing is performed on the address information of each target medical institution to obtain the corresponding institution location information parameters;
[0035] When the address of any target medical institution cannot be successfully resolved or the confidence level of the corresponding institution location information parameter is lower than a preset threshold, the institution location information parameter is supplemented based on a predefined auxiliary location information generation rule.
[0036] Based on the location information parameters of the institutions, the geographical proximity between any two target medical institutions is calculated, and the geographical proximity matrix is constructed.
[0037] As an optional implementation, after calculating the geographical proximity between any two target medical institutions based on the institution location information parameters, the method further includes:
[0038] Within the third preset period, business interaction logs of each target medical institution are collected. The business interaction logs record at least the interaction timestamp, the identifier of the other party, the interaction type identifier, and the amount of interaction data.
[0039] For any two target medical institutions, count the number of interactions and the amount of interaction data within the third preset period;
[0040] Based on the number of interactions, the amount of interaction data, and the geographical proximity, the corresponding proximity value is weighted and adjusted to obtain the updated proximity value.
[0041] Based on the updated proximity values, the geographic proximity matrix is constructed.
[0042] As an optional implementation, the step of stacking the time-series data corresponding to each target medical institution according to the institution dimension to construct a joint feature matrix for the same data quality indicator includes:
[0043] Based on the geographic proximity matrix, target medical institutions with geographic proximity higher than a preset grouping threshold are divided into the same neighborhood cluster according to the proximity aggregation rule;
[0044] For each neighboring cluster, the anomaly count of the target medical institution within the cluster is counted at each time-series index, and a regional synchronization anomaly marker is generated when the anomaly count reaches a preset cluster threshold.
[0045] When the regional synchronization anomaly flag exists, the corresponding time-series data is merged at the cluster level according to the preset merging rules to obtain a cluster-level representative data column.
[0046] The cluster-level representative data column and the time-series data column of the target medical institution outside the cluster are stacked according to the institution dimension and filled into the joint feature matrix.
[0047] As an optional implementation, stacking the cluster-level representative data column and the time-series data column of the target medical institution outside the cluster along the institution dimension to fill the joint feature matrix includes:
[0048] For each neighboring cluster, the cluster-level representative data column is used to perform differential processing on the time-series data column of the member institutions within the cluster to generate the corresponding residual data column;
[0049] The cluster-level representative data column and each of the residual data columns are stacked according to the institution dimension, and merged with the time-series data column of the target medical institution outside the cluster to fill the joint feature matrix.
[0050] As an optional implementation, a multivariate fitting algorithm is performed based on the joint feature matrix to obtain trend change parameters for each target medical institution, including:
[0051] The joint feature matrix is partitioned according to a preset sliding time window to obtain multiple window feature sub-matrices;
[0052] For each of the aforementioned window feature submatrices, a weighted regression fitting is performed to obtain the corresponding candidate trend parameters and residual confidence values.
[0053] The candidate trend parameters and residual confidence values of each window are used as multi-source evidence information to construct a window-level basic confidence assignment.
[0054] The window-level basic confidence assignment is processed by evidence theory synthesis, and the weight of conflicting evidence information is dynamically adjusted when the degree of conflict exceeds a preset range, generating comprehensive trend change parameters for each target medical institution.
[0055] Secondly, this application provides a data management system for medical consortia, including:
[0056] The data acquisition module is used to acquire data acquisition task logs and corresponding server node operation status logs of each target medical institution according to a first preset period; the server node operation status logs include: node identifier, node heartbeat information, node resource occupancy rate, and node availability status;
[0057] The first processing module is used to calculate the data quality indicators and server node operation status indicators of each target medical institution based on the data acquisition task log and the server node operation status log, and generate time-series data of each data quality indicator and operation status indicator according to a second preset period; the operation status indicators include: node failure rate and abnormal resource occupancy rate.
[0058] The second processing module is used to stack the time series data corresponding to each target medical institution according to the institution dimension for the same data quality index, construct a joint feature matrix, and perform a multivariate fitting algorithm based on the joint feature matrix to obtain the trend change parameters of each target medical institution.
[0059] The detection module is used to generate a degradation trend detection result based on the trend change parameters and the operating status indicators of the server node, according to preset judgment conditions; based on the degradation trend detection result, generate trend warning information, and output the trend warning information to the management terminal.
[0060] Compared to existing technologies, this application automates the acquisition of data collection task logs from various medical institutions and further analyzes multiple indicators such as data missing rate, format error rate, and data collection delay, enabling precise quantitative assessment of data quality within the medical consortium. Simultaneously, by employing joint feature matrix construction and multivariate fitting algorithms, the method accurately and efficiently captures and reflects the trend changes in data quality over time, significantly improving the real-time performance and accuracy of abnormal trend detection. Furthermore, this application's solution can automatically generate trend-based early warning information based on trend change parameters and push it to the management terminal in real time, helping managers take timely countermeasures and effectively avoid the risk of medical decision-making bias and service interruption caused by continuous data quality deterioration. Compared to existing technologies, this application's method significantly reduces the latency and subjective errors of manual analysis and enhances the sensitivity and rapid response capability to data quality anomalies within the region, making it more suitable for large-scale medical consortium collaborative management scenarios with massive data volumes. Attached Figure Description
[0061] Figure 1 A flowchart of the data management method for medical consortia provided in this application;
[0062] Figure 2 A flowchart of the method for generating degradation trend detection results provided in this application;
[0063] Figure 3 A flowchart illustrating the method for obtaining a comprehensive confidence result for the overall data collection situation provided in this application;
[0064] Figure 4 A schematic diagram of the data management system for medical consortia provided in this application. Detailed Implementation
[0065] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0066] Reference Figure 1 The diagram shown is a flowchart of the data management method for medical consortia provided in this application. The method includes steps S101 to S104, wherein:
[0067] S101: Acquire the data collection task logs of each target medical institution and the corresponding server node's operation status logs according to the first preset cycle; the server node's operation status logs include: node identifier, node heartbeat information, node resource occupancy rate, and node availability status;
[0068] S102: Based on the data acquisition task log and the server node's operation status log, respectively, the data quality indicators and server node operation status indicators of each target medical institution are statistically analyzed, and time-series data of each data quality indicator and operation status indicator are generated according to the second preset period; the operation status indicators include: node failure rate and abnormal resource occupancy rate;
[0069] S103: For the same data quality indicator, stack the time series data corresponding to each target medical institution according to the institution dimension to construct a joint feature matrix, and perform a multivariate fitting algorithm based on the joint feature matrix to obtain the trend change parameters of each target medical institution.
[0070] S104: Based on the trend change parameters and the operating status indicators of the server nodes, generate a degradation trend detection result according to preset judgment conditions; generate trend warning information based on the degradation trend detection result, and output the trend warning information to the management terminal.
[0071] As an optional implementation, the management terminal is further used for:
[0072] It displays the real-time running status information of the corresponding server nodes of each target medical institution and the execution status information of the data collection tasks;
[0073] The operational status information includes: node availability status, node resource utilization rate, and node fault indication; the execution status information includes the current task status, task execution progress, and task execution result.
[0074] The management terminal is also used to: provide task management functions based on the execution status information of the data acquisition task and the running status information of the server node; the task management functions include task start / stop control, task priority adjustment and task retry strategy management.
[0075] As an optional implementation, the data quality indicators include: data missing rate, format error rate, and data acquisition delay; the data acquisition task log includes: acquisition task identifier, acquisition time, number of acquisition items, acquisition results, and acquisition status.
[0076] The data acquisition task log refers to a set of structured records automatically generated by the medical information system. The meanings of each field are as follows: Acquisition Task Identifier uniquely identifies a data extraction operation; Acquisition Time indicates the trigger time; Number of Acquisition Items describes the total number of data tables or interfaces extracted; Acquisition Result provides an enumeration of successful, failed, or partially successful extraction statuses; Acquisition Status records the real-time status of the current operation, such as queuing, executing, or completed. The first preset period can be a five-minute rolling time window, triggered by a unified scheduling strategy.
[0077] This application continuously reads the task logs uploaded to the message queue by each target medical institution, writes the log entries to a time-series database, and maintains them in a chained manner. Then, it executes the indicator calculation process: it iterates through the log slices of each medical institution within the second preset period, and calculates three indicators: missing rate, format error rate, and collection latency. The missing rate is obtained by comparing the number of extracted items with the actual number of successful items; the format error rate is determined based on the count of failed parsing entries; and the collection latency is mapped from the difference between the log reporting time and the collection time to a latency interval label. The calculation results are written back to the original database as three columns of indicator values and supplemented in chronological order to form a time-series segment of indicators corresponding to each institution.
[0078] The system performs a stacking operation on the same indicator across institutions. It first reads the indicator columns for which gaps have been filled in each institution, then performs horizontal concatenation according to the institution's unique code to generate a joint feature matrix. The row dimension of the matrix represents the time index, the column dimension is sorted by institution, and the matrix cell contains the discrete values of the same indicator in the same time slice. This application calls a multivariate fitting algorithm, using a sliding time window as the basic unit, to fit the matrix column vectors. The fitting algorithm is implemented without formulas using a column-wise residual recursive approach: first, all columns are normalized to zero mean - unit variance; then, iterative least squares is applied to estimate the column trend slope; then, using the slope of earlier samples as the initial value, the latest trend change parameters are continuously smoothed and updated through a rolling window. The trend change parameters are categorized into three ranges: positive, negative, and near-zero. Positive indicates continuous deterioration of the indicator, negative indicates improvement, and near-zero indicates stability.
[0079] The system compares trend change parameters with pre-set judgment thresholds, which are determined by a percentile voting strategy based on the historical indicator variance. When the trend change parameter of any institution exceeds the corresponding threshold, the system immediately marks it as a single institution deteriorating. When the proportion of deteriorated institutions exceeds 15 percentage points within the same time window, this application aggregates the single institution marking information to generate deterioration trend detection results and records the detection items according to institution code and indicator type.
[0080] The detected items enter the early warning generation process, where the system assembles trend-based early warning information according to the early warning strategy. Early warning information fields include: early warning identifier, list of triggering organizations, impact indicators, trigger time, trend parameter snapshot, and suggested action. This application pushes the generated JSON structure to the management terminal via a REST interface, and simultaneously displays it in real-time on the alarm dashboard using red, orange, and yellow color blocks.
[0081] This implementation method ensures that data quality trends across all institutions within the medical consortium are quickly visible and automatically alerted through unified zipper-style task log management, three-indicator time-series calculation, joint feature matrix construction, multivariate slope recursion, and threshold determination closed-loop. It can detect indicator degradation within a five-minute timeframe, identifying anomalies hours earlier than manual weekly or daily reports, significantly improving data governance timeliness. The joint feature matrix retains institutional granularity information and avoids analytical gaps caused by column incompleteness, laying a stable input foundation for subsequent evidence synthesis.
[0082] Regarding the above S101:
[0083] In this embodiment, the application deploys a log push service within each target medical institution, with logs written to a message queue topic using a unified JSON structure. The scheduling module triggers a retrieval operation based on a first preset period. The first preset period can be set to five minutes; an exemplary range is one to fifteen minutes, allowing maintenance personnel to select based on data volume and monitoring real-time requirements. The number of message entries retrieved by the scheduling module at one time is limited by the batch size, which can be set to 500 entries, with an exemplary range of 100 to 1000 entries, to balance network bandwidth usage and processing latency.
[0084] Upon receiving the logs, the system sequentially parses five fields: task identifier, collection time, number of collection items, collection result, and collection status. The task identifier must remain unique within the current period; the collection time is converted to UTC; the number of collection items and the collection result must pass range and enumeration checks; and the collection status is mapped to generate an internal status code. An example collection status can be set to one of three enumerated values: running, completed, or failed.
[0085] After successful parsing and verification, the system writes the organization code and task identifier combination as the primary key to the time-series database, along with fields such as timestamp, number of items, result code, status code, and organization identifier. The write operation uses a batch submission mode, with the batch size matching the fetch batch size. If a write failure occurs, this application automatically executes a retry mechanism. The retry threshold can be set to two times, with an exemplary range of one to five attempts. If writing still fails after reaching the threshold, the system transfers the corresponding batch log to an error channel and generates an alarm event to alert maintenance personnel.
[0086] For medical institutions with smaller data volumes or those using traditional databases, this application can also enable an active pull method based on direct database connections. This involves querying the log table by time interval to perform the same field extraction and validation, and then writing the results to a time-series database. Time zone conversion, uniqueness verification, enumeration mapping, and retry strategies remain consistent with the message queue path; only the triggering mechanism changes from passive consumption to active polling.
[0087] For example, in one optional implementation, the data acquisition task can be executed by an ETL engine, and its running status log is the ETL task execution log.
[0088] Through the above parameterized configuration and pipelined processing, this application can collect data collection task logs from each target medical institution in the medical consortium completely and reliably within each first preset cycle, providing a timely and consistent original data foundation for subsequent data quality indicator statistics and trend analysis.
[0089] Regarding S102 above:
[0090] After the logs are written to the database, this application initiates an indicator generation task. The task scheduler executes the task cyclically with a second preset period as the trigger granularity. The second preset period can be set to thirty minutes; an exemplary range is ten minutes to one hour, which can be configured according to the need to balance trend sensitivity and computing resource consumption. The scheduler iterates through the log records newly added in the time-series database within the current period window by institution code, and calculates three data quality indicators for each target medical institution: data missing rate, format error rate, and data acquisition delay.
[0091] To obtain the data missing rate, this application first reads the total number of data collection tasks and the number of successful records for the organization within the window. Missing tasks are defined as those marked as failed or partially successful in the data collection result field. If the total number of records is zero, the system marks the missing rate for this period as invalid and adds a coverage indicator; otherwise, it calculates the percentage of missing tasks, maps the result to a percentage format, and retains two decimal places.
[0092] For example, the range can be judged as 0-5% as excellent, 5-15% as acceptable, and above 15% as unacceptable. The range thresholds are set by the medical consortium data governance specifications during the deployment phase and can be dynamically adjusted through the management terminal afterwards.
[0093] The format error rate is calculated based on the number of format validation failures recorded during log parsing. This application calculates the format error rate by comparing the number of failed parsing entries for the organization within the statistical window with the total number of task records. To prevent misjudgment when the number of collected items is too low, the system does not calculate the format error rate in windows with fewer than ten records, but instead marks the indicator as insufficient coverage.
[0094] For example, the thresholds for differentiation are 0.5%, 1%, and 3%, which correspond to three levels: excellent, warning, and risk, respectively.
[0095] Data acquisition latency measures the degree of delay between log reporting time and acquisition time. This application calculates this difference for each log entry and maps it to a latency level based on the latency duration. For example, the levels can be set to four levels: 0–5 minutes, 5–15 minutes, 15–30 minutes, and over 30 minutes. This application takes the median of the latency levels for all logs from the organization within the same window as the window latency index. If more than one-third of the logs within the window are in the "over 30 minutes" latency level, the system directly marks the latency index as the highest risk level to ensure that late batch reporting is not masked by the average.
[0096] After generating the three indicators, this application writes the missing rate, format error rate, and latency level into the institutional indicator time series table, with corresponding fields including window end timestamp, institutional code, indicator type, and indicator value. During writing, if an old value already exists for the timestamp and institutional combination, the system overwrites it using time priority and records the overwrite log to prevent duplicate statistics. For cases where there are no valid logs within the window, the system writes a placeholder record and marks the indicator value field with a null value, which can be detected and handled as a gap in the subsequent trend fitting stage.
[0097] After the time series table of indicators is written, the data completion module is triggered to fill the gaps by forward filling or linear interpolation. This application uses the forward filling strategy by default; exemplary optional strategies include backward filling or external mean filling, which can be switched through the configuration file. After completion, the system will update the identifier and write it to the task progress table, providing continuous, gap-free, and standardized time series data for the subsequent construction of the joint feature matrix in S103.
[0098] In this cycle, this application can generate traceable and comparable time series of data quality indicators for each target medical institution in the medical consortium within each second preset cycle, laying an accurate and stable input foundation for subsequent trend analysis.
[0099] Regarding the above S103:
[0100] After completing the time series completion of the indicators, this application enters the joint feature matrix construction stage. The system first reads the current indicator generation completion flag from the task schedule table and processes them one by one according to indicator type. Taking "data missing rate" as an example, the system scans the indicator time series of all institutions according to a unified time index, aligns the time index to the second-level Unix timestamp, and completes time zone unification. For any time point that still has null value markers, the system applies the most recent valid value to fill forward; if the length of consecutive null values exceeds thirty index intervals, an additional gap marker column is inserted to facilitate the subsequent model to identify coverage defects. After the filling is completed, the system horizontally concatenates the columns according to the ascending order of the institution code to generate a joint feature matrix with the row dimension being the time index and the column dimension being the institution code.
[0101] To reduce the impact of inter-column dimensional differences on the fitting, this application performs column-level standardization on the matrix. The standardization method adopts a zero-mean plus unit variance scheme. The standardization parameters are calculated and persisted during the first run, and are directly reused in subsequent cycles to ensure that the dimensions remain consistent across different rolling windows. If the variance of an institution in the historical samples is lower than a one percent threshold, the system automatically marks the column as a low-variability column and skips it during fitting to prevent numerical instability caused by near-zero variance.
[0102] After the matrix is prepared, this application performs multivariate fitting on each window. The system scrolls with a third preset window size; the exemplary window length can be set to 120 time indices, and the specific range can be adjusted from 30 to 240. Within each window, this application uses an incremental least squares recursive algorithm to calculate the linear slope of the matrix column vectors, using the slope of the previous window as the initial value, and updates in five microsteps to converge to the latest estimate. To mitigate the instantaneous impact of abnormal peaks on the slope, the system performs ternary truncation on each row before recursion, replacing outliers exceeding one quartile distance from the upper or lower limit with neighboring non-outliers.
[0103] After obtaining the window slope, the system generates a trend change parameter time series for each column. To facilitate threshold determination, this application maps the slope values to five ranges: "accelerated degradation," "slight deterioration," "stable," "slight improvement," and "accelerated improvement." Exemplary boundaries can be set at 25%, 50%, 0%, -50%, and -25% of the standard deviation, and can be automatically learned based on historical fluctuation ranges during deployment. The system adds a trend range column below the same column in the matrix, synchronously storing the original slope and range labels for use in the next stage of degradation detection.
[0104] To reduce the persistence of the full matrix, this application releases the row and vector buffer of the previous window after each window fitting, retaining only four types of information: archived sample index, standardized parameters, slope vector, and trend level labels. For monthly or quarterly trends requiring long-term comparison, the system resamples and aggregates the slope sequences in offline batch tasks. Through a window-based iterative construction-fitting-compression pipeline, this application maintains the rolling real-time performance of trend calculation while controlling storage and memory overhead.
[0105] Regarding S104 above:
[0106] After acquiring the window slope and mapping it to trend levels, the system immediately enters the degradation detection phase. A dynamic status register is maintained for each medical institution, with fields including: the trend level sequence of the last five windows, the cumulative degradation count, the start time of continuous degradation, and the current degradation level. This application performs the following actions during each window update.
[0107] First, the system reads the latest trend level. If the level is "accelerated degradation" or "slight deterioration," the cumulative degradation count in the register is incremented by one; otherwise, it is reset to zero. The cumulative count threshold can be set to three (for example), with an optional range of two to five, to control the trigger sensitivity of "continuous window degradation." When the cumulative count reaches the threshold, the system marks the institution as "single institution degradation" and records the start time of continuous degradation. If the trend level does not fall back to "remaining stable" or better within the subsequent two windows, the system upgrades the degradation level to "aggravated." The level enumeration includes four states: initial, degradation, aggravated, and mitigation. This application automatically pushes a "technical operation and maintenance verification" task to institutions in the aggravated state that have not been mitigated for three consecutive windows.
[0108] Subsequently, the system performs a horizontal aggregation of all institutions at the window granularity. This application counts the number of institutions in the current window that are in a "deteriorated" or "aggravated" state, and divides this number by the total number of institutions to obtain the deterioration percentage. The deterioration percentage threshold can be configured to 15%, for example, with an adjustable range of 10% to 30%. If the percentage exceeds the threshold and covers at least two data quality indicators, the system generates a "regional deterioration event." The event fields include: event identifier (composed of indicator type + window end time), list of triggering institutions, deterioration indicator, trigger threshold, percentage value, and deterioration level distribution.
[0109] After an event is generated, it enters the early warning assembly process. This application matches the alarm strategy according to the event level: single-organization degradation is mapped to a yellow warning, regional degradation events are mapped to an orange warning, and regional escalation events are mapped to a red warning. In addition to basic information, the assembly fields also include suggested operation text. Suggested operation examples can be "Check VPN tunnel connectivity," "Verify interface field mapping rules," "Reset ETL extraction nodes," etc., and the content is automatically matched by the operation and maintenance knowledge base according to the indicator type and organization level.
[0110] After the warning information is generated, it is written to the alarm table and pushed to the management terminal via asynchronous message. This application uses the REST interface by default; for example, optional channels also include WebSocket push streaming and email notification. After receiving the information, the terminal displays it on the large screen simultaneously as a color block and a scrolling list. If the same organization triggers the same level of warning repeatedly within ten minutes, the system merges them into one warning and updates the trigger count and latest time in the message body.
[0111] This application archives all early warnings from the previous day to the historical alarm database at midnight every day. The archived fields include the closed-loop processing results and the time consumed, which are used for subsequent performance evaluation and adaptive training of model thresholds. Through the closed-loop process of register maintenance, horizontal proportion detection, hierarchical mapping and multi-channel push, the system can continuously analyze the data quality trend of the medical consortium and output easy-to-execute early warning instructions to ensure early detection of anomalies, early assignment of responsibilities, and traceable handling.
[0112] For example, after the platform completes the storage of server node operation status logs, it will periodically perform statistical analysis on the node operation status. Specifically, the system can read the heartbeat information of each server node every 30 seconds and record the node's response status. If a node has three consecutive missing heartbeat messages or no response, it is considered a node failure sample. The system uses this to calculate the node failure rate, which can be done by dividing the number of node failure samples within the statistical period by the theoretical total number of heartbeats, and finally presenting it as a percentage, rounded to two decimal places.
[0113] The abnormal resource utilization rate of a node is determined based on periodically sampling the node's CPU, memory, and network bandwidth usage. Specific sampling thresholds can be set as follows: CPU utilization exceeding 85%, memory utilization exceeding 80%, and bandwidth utilization exceeding 70%. If any indicator exceeds these thresholds in a single sampling, it is considered an abnormal sampling. At the end of each statistical period, the system divides the number of abnormal samplings within that period by the total number of samplings to obtain the abnormal resource utilization rate.
[0114] For example, the system comprehensively considers the health status of nodes and the execution status of data acquisition tasks to accurately determine the degradation trend of data quality. First, the platform classifies the health status of nodes into three levels: normal, warning, and fault, based on the node failure rate and abnormal resource utilization rate. When the health status of a server node is marked as faulty, and there are failure records of the data acquisition tasks it undertakes within the same period, the system will directly determine that the degradation trend of the task belongs to "node fault triggering", without the need for additional slope judgment of trend change parameters.
[0115] If a node's health status is at the warning level, and the data trend parameters of the data acquisition task on that node show a deteriorating trend (e.g., a positive slope) within the same period, the system will automatically increase the weight of that task in the deterioration trend detection, with an example increase of 1.5 times. In this way, abnormal fluctuations in node status can be more clearly reflected in the trend determination process of task data quality, strengthening the correlation between node operating status and task quality.
[0116] For example, the management terminal has detailed task management functions, including task start / stop control, task priority adjustment, and task retry policy management. The management terminal implements task start / stop through a predefined REST interface. For instance, by calling the interface / task / {id} / control and setting the action to "pause" or "resume" in the request parameters, the specified data acquisition task can be started / stopped in real time.
[0117] For task priority adjustment, the management terminal adjusts task priorities in real time via the interface / task / {id} / priority. The system accepts a priority range of 0 to 10 for example. After each adjustment, the new priority value is immediately written to the metadata table of the task scheduling center, triggering a real-time reordering of the task scheduling queue so that high-priority tasks are executed first.
[0118] Furthermore, the retry strategy for tasks is dynamically configured via the management terminal. The system supports setting the maximum number of retries (maxRetry) and the retry backoff interval (backoffSeconds). When a task's running status changes to failure, the scheduler will automatically reschedule the task to a backup server node within the same medical institution or a nearby server node in an exponential backoff manner, according to the configured backoff strategy, until the task is successfully completed or the predetermined maximum number of retries is reached. When a task ultimately fails, the system will automatically generate a corresponding failure alarm on the management terminal for administrators to manually intervene and further process.
[0119] As an optional implementation, see [link to implementation details]. Figure 2The flowchart of the method for generating degradation trend detection results provided in this application includes steps S201 to S205, wherein:
[0120] S201: Generate corresponding single-institution judgment result information based on the trend change parameters of each target medical institution;
[0121] S202: The single-institution judgment result information is used as multi-source evidence information input, and evidence theory is used for synthesis processing to obtain a comprehensive confidence result for the overall data collection situation;
[0122] S203: Calculate the conflict degree of the comprehensive confidence result, and perform weight adjustment on the conflict evidence information when the conflict degree exceeds the preset range to obtain the corrected comprehensive confidence result;
[0123] S204: Compare the corrected comprehensive confidence result with the preset confidence threshold to determine the abnormal confidence level at the medical consortium level;
[0124] S205: Based on the aforementioned abnormal confidence level and the single-institution judgment results of each target medical institution, generate the degradation trend detection results.
[0125] The single-institution judgment result information refers to the judgment entries formed by a single medical institution based on its trend change parameters within a specific time window. Each judgment entry includes fields such as institution code, indicator type, trend level, trend confidence level, and evidence weight. Evidence theory synthesis processing refers to synthesizing multi-source judgment entries according to a modified Dempster rule or similar method to obtain the comprehensive confidence level allocation of the three propositions—overall deterioration, normal, and uncertain—within the current time window. Conflict degree represents the accumulated contradiction value when different judgment entries contradict each other in proposition attribution. When this contradiction value exceeds the range pre-set in this application, the relevant evidence information needs to be downweighted. The abnormal confidence level, by comparing the deterioration confidence level with a preset confidence threshold, ultimately classifies the overall risk into high, medium, or low levels for subsequent early warning decisions.
[0126] In its implementation, this application first constructs a basic confidence assignment based on each individual institution's judgment result. The trend confidence level, combined with the evidence weight, is mapped to either a "deteriorated" or "normal" proposition. The unassigned portion is categorized as an "uncertain" proposition to explicitly indicate incomplete information or ambiguous trends. Multiple basic confidence assignments are sequentially input into the evidence theory synthesis module, and conflict levels are calculated in real time. If the conflict level between a new judgment item and the current synthesis result exceeds a predefined limit, the system reduces the evidence weight of that item to prevent extreme or low-coverage institutions from exerting unreasonable influence on the overall judgment. If the high conflict level still cannot be eliminated, the remaining conflicting portion is placed under the uncertain proposition, pausing subsequent synthesis, thus preserving local conflict information without undermining the consensus of most institutions. The final synthesis result includes three synthesis confidence levels: deteriorated, normal, and uncertain. The system further compares the deterioration confidence level with a credibility threshold to classify abnormal confidence levels at the medical consortium level as high, medium, or low.
[0127] For example, after collecting trend change parameters from three medical institutions, the system generates three single-institution judgment results. If one institution's trend level is determined to be accelerated deterioration with a high confidence level and a coverage rate of over 90%, it will provide strong support for the "deterioration" proposition during synthesis. If another institution's trend level is determined to be slight improvement with a low coverage rate, its evidence weight will be reduced, resulting in weaker support for the "normal" proposition during synthesis. Once the cumulative conflict between these two opposing pieces of evidence exceeds a set limit, the system will further reduce the evidence with relatively low confidence and poor coverage. After several reductions, if the conflict level is still higher than the limit, the remaining conflicting part is classified under an uncertain proposition to preserve and isolate this local information and avoid interfering with the overall judgment. Upon completion of synthesis, this application obtains a relatively high overall confidence level for the "deterioration" proposition, thus concluding that the abnormal confidence level is high. Simultaneously, the conflict reduction process is recorded in the system log so that maintenance personnel can further investigate the cause and reliability of this potentially extreme evidence item.
[0128] This approach offers significant advantages for scenarios involving multiple institutions and diverse data sources within a medical consortium environment. By employing dynamic weight adjustment and uncertain proposition allocation strategies during the synthesis phase, it avoids extreme misjudgments that can result from simple voting or weighted averaging methods. It also automatically weakens evidence items with low coverage or insufficient confidence when multiple sources of evidence conflict. The synthesis result ultimately provides a quantitative indicator of whether the overall data collection quality has significantly deteriorated, presented with high, medium, and low anomaly confidence levels. This facilitates different emergency measures for medical consortium managers based on varying levels of risk. Compared to methods relying solely on a single threshold or simple statistical comparisons, this evidence-based synthesis approach more accurately identifies overall trends, reduces noise interference, and improves the precision and reliability of data governance within medical consortia.
[0129] As an optional implementation, see [link to implementation details]. Figure 3 The flowchart of the method for obtaining a comprehensive confidence result for the overall data collection situation provided in this application includes steps S301 to S304, wherein:
[0130] S301: Construct an initial basic confidence allocation based on the single-institutional assessment results of each target medical institution;
[0131] S302: Calculate the degree of conflict between the initial basic confidence assignments, and when the degree of conflict exceeds a preset range, adjust the weight of the conflict evidence information to update the basic confidence assignments;
[0132] S303: Perform iterative synthesis processing on the updated basic confidence assignment to obtain a first synthesized confidence level and a second synthesized confidence level; wherein, the first synthesized confidence level is used to characterize that the data acquisition quality is in a deteriorating trend, and the second synthesized confidence level is used to characterize that the data acquisition quality is in a normal trend;
[0133] S304: Determine the uncertainty entropy value based on the first composite confidence level and the second composite confidence level; generate a comprehensive confidence result based on the first composite confidence level, the second composite confidence level, and the uncertainty entropy value; wherein, the uncertainty entropy value is used to characterize the consistency of the overall evidence of the medical consortium.
[0134] This application introduces a multi-iteration comprehensive processing method when globally synthesizing the single-institution judgment results of each target medical institution. To ensure timely correction of multi-source evidence when the degree of conflict exceeds the limit, while maintaining compatibility with uncertain information, this application first assigns an initial basic confidence allocation to each single-institution judgment result. The basic confidence allocation is used to characterize the degree of support that the entry provides for the two propositions: "data collection quality is on a deteriorating trend" and "data collection quality is on a normal trend," and reserves a portion of the probability space for uncertain propositions, mainly reflecting information gaps when the coverage of medical institutions is insufficient or the trend direction is ambiguous.
[0135] After obtaining all initial basic confidence assignments, this application calculates the conflict degree for any two basic confidence assignments. To identify high-conflict situations caused by institutional judgments with significant differences in coverage or extreme directions, the system accumulates the parts of the two assignments that are completely opposed to the deterioration and normal propositions as conflict values. If the conflict degree exceeds the threshold range set in the deployment phase, the system will lock the most recently included entry in the conflict degree iteration calculation process that is clearly opposed to other evidence and adjust its corresponding weights.
[0136] For example, if an institution's weight is originally high, but its coverage is insufficient to support its high credibility, then during the conflict degree iteration and merging, this item will be preferentially downweighted or partially transferred to an uncertain proposition. After one or more rounds of adjustment, when the conflict degree falls back to an acceptable range, the system will use the merging result as the output of the current iteration and enter the next round of synthesis.
[0137] In the final output of the multiple iterative synthesis, this application focuses on two key confidence levels: the first composite confidence level and the second composite confidence level. The first composite confidence level characterizes the strength of support for the proposition that the medical consortium as a whole is in a "deteriorating trend" after the information from all institutions is synthesized. If this value is higher than a certain confidence threshold, it means that the system's judgment of the overall deteriorating state is relatively certain. The second composite confidence level is used to quantitatively characterize the system's degree of confidence in the proposition of a "normal trend" under the combined effect of the judgment results of the same batch of institutions.
[0138] By introducing a dynamic weight adjustment mechanism in each iteration, extreme conflict evidence can be gradually weakened or removed from polar attribution, thus presenting a more reasonable overall trend assessment in the synthetic output.
[0139] After obtaining the first and second composite confidence levels, this application derives the uncertainty entropy value from the degree of inconsistency between the two and the amount of information not assigned to any explicit trend proposition. A high entropy value means that there is still a large range of opinion dispersion or conflict margin among multiple sources of evidence; a low entropy value indicates that most evidence can relatively consistently point to a certain trend state, and there is little residual evidence after dynamic conflict processing.
[0140] Based on the first composite confidence level, the second composite confidence level, and the uncertainty entropy value, this application integrates them into a final comprehensive confidence result. This result can be a combination of three elements, such as presented as three values: "deteriorated confidence level," "normal confidence level," and "uncertainty entropy." It can also be further converted into a level, color, or warning level in the backend for easy reference by management terminals or visualization panels.
[0141] In this way, when faced with heterogeneous and partially conflicting trend assessment results from medical institutions, this application can more robustly absorb noise from extreme or low-coverage evidence. Based on the design of uncertain propositions and entropy values, it also enables medical consortium managers to clearly understand the degree of disagreement in overall trend assessments, providing a more reasonable decision-making reference for subsequent operational interventions. Therefore, this application not only achieves global support calculation for the two propositions of "deterioration" and "normal," but also significantly reduces the risk of conflicting evidence interfering with the final results while retaining uncertain information, thereby improving the accuracy and interpretability of medical consortium data quality trend detection.
[0142] As an optional implementation, it also includes:
[0143] A geographic proximity matrix is generated based on the geographic location information between the target medical institutions.
[0144] Based on the geographic proximity matrix, the single-institution determination results of each target medical institution are subjected to proximity-based weighted processing.
[0145] The weighted single-agency determination results are used to generate the initial basic confidence assignment.
[0146] In practice, the first step is to collect address data for each target medical institution during the deployment phase. Then, third-party geospatial services, such as Amap, Baidu Maps, or OpenStreetMap's offline API, are used to convert the addresses to latitude and longitude coordinates. This application will write the latitude and longitude results, along with the confidence score returned by the parsing, into the institution location information database, and simultaneously register auxiliary fields such as hospital code, grade, and administrative division code in the entries.
[0147] During operation, when the system needs to calculate the geographical proximity of two medical institutions, it retrieves the corresponding coordinates from the location information database and calculates the spatial distance between them using a predetermined engineered distance algorithm, such as spherical distance or city-level road network lookup. If the distance is significantly less than a certain kilometer threshold and they belong to the same municipal administrative region, the system classifies them as having high proximity.
[0148] Multiple distance thresholds can be set in the configuration file. For example, the proximity score can be above 0.9 within 5 kilometers, between 0.8 and 0.5 between 5 and 30 kilometers, and below 0.4 for distances above 30 kilometers. For organizations experiencing address resolution failures or excessively low resolution confidence, the system will invoke a predefined fuzzy positioning scheme, such as using city-level or county-level center coordinates to replace the actual location, and automatically reduce the offset score by 0.1 to 0.2 when calculating proximity to reflect positioning uncertainty.
[0149] After obtaining the proximity scores for all institution pairs, this application generates a geographic proximity matrix, with rows and columns corresponding to unique institution identifiers, and matrix cells recording the proximity between the two institutions. To support rapid querying in subsequent engineering steps, this application can store this matrix in a distributed cache or relational database, indexed by "institution pair" to "proximity score," ensuring that the proximity between any two institutions can still be quickly retrieved even in large-scale medical consortia.
[0150] Before synthesizing evidence from the single-institution judgment results of each target medical institution, this application reads the proximity of each judgment in the proximity matrix and adjusts the weights according to the actual positional relationship between the institutions. Specifically, if the proximity of an institution to its neighboring institutions in the matrix is in a high range (e.g., above 0.85), the weight of the judgment entry can be multiplied by an amplification factor, for example, increasing it from 1.0 to 1.2; if the proximity is in a medium range (e.g., 0.6–0.7), the original weight remains unchanged; if the proximity is below 0.4, the weight of the corresponding entry is set to 0.8. If the proximity further decreases, the weight can be further reduced, for example, to 0.6 or lower. This application updates these calculated weights to the single-institution judgment result information and then uses this to generate the initial basic confidence assignment. In implementation, a field can be added to the judgment entry structure to store the final weight after proximity weighting.
[0151] In real-world scenarios, if a township health center and several other health centers within the same county frequently share data with the central hospital via the same VPN, their proximity should be assessed as high. If these institutions collectively exhibit high missing rates or latency anomalies within the same time window, this application will give its criteria entries greater influence in subsequent synthesis stages, thereby enhancing the ability to identify regional faults or network bottlenecks. Conversely, if a specialist hospital is too far from most other institutions, or does not use a unified network, it will not strongly influence the evidentiary information of these institutions.
[0152] Through this well-defined engineering process, this application can flexibly incorporate geographical location differences into the weighting calculation of single-institutional judgment results, providing more realistic basic data support for subsequent evidence synthesis and trend early warning.
[0153] As an optional implementation, generating a geographic proximity matrix based on the geographic location information between the target medical institutions includes:
[0154] Geographic parsing is performed on the address information of each target medical institution to obtain the corresponding institution location information parameters;
[0155] When the address of any target medical institution cannot be successfully resolved or the confidence level of the corresponding institution location information parameter is lower than a preset threshold, the institution location information parameter is supplemented based on a predefined auxiliary location information generation rule.
[0156] Based on the location information parameters of the institutions, the geographical proximity between any two target medical institutions is calculated, and the geographical proximity matrix is constructed.
[0157] When generating the geographic proximity matrix, this application first performs geospatial resolution processing on the address information of each target medical institution to obtain institution location information parameters. Commonly available address resolution services can be used; for example, the registered addresses of hospitals, clinics, or health centers can be submitted to a third-party geocoding API to obtain additional data such as latitude and longitude coordinates, administrative division level information, and resolution confidence indicators. Resolution confidence is typically expressed as a numerical percentage or tiered level, reflecting the reliability of the geospatial service's reverse resolution results for that address.
[0158] In some cases, the address of the target medical institution may be non-standard or fail to be successfully identified by the geographic service, or it may be successfully resolved but the returned confidence level is lower than the threshold set during the deployment of this application. To address these situations where the accurate latitude and longitude cannot be directly obtained, this application sets predefined auxiliary location information generation rules for completion.
[0159] For example, auxiliary information can be a pre-entered administrative division comparison table, or a mapping relationship between hospital level and district / county code maintained by the medical consortium management system. If the system returns a "missed" or "low confidence" message at the geographic service interface, it will directly refer to the comparison table or code mapping file to perform fuzzy address positioning. For example, it may assign the unresolved township health center to the center coordinates of its county-level administrative region, or combine the hospital level with the city ID to estimate an alternative coordinate for subsequent proximity calculation.
[0160] After generating or completing the location information parameters of the institutions, this application calculates the geographical proximity of any two target medical institutions based on these latitude and longitude coordinates or alternative coordinates. To better reflect the regional distribution characteristics commonly found in medical consortia, the system typically adopts several combined strategies. The latitude and longitude of the two medical institutions can be first substituted into a spherical distance algorithm or a city road network apparent distance database can be used. Exemplary practices include: if the institutions are located in the same city, city zoning weights are used to reduce some of the distance's influence; if the institutions are located across provinces or are more than tens of kilometers apart, the proximity is set to a relatively low value.
[0161] You can also define grading rules, such as marking distances within five kilometers as "high proximity" with a proximity degree of 0.8 to 1.0, and marking distances exceeding thirty kilometers as "low proximity" with a proximity degree of 0.2 to 0.4. The specific range can be configured by the operations and maintenance personnel during the deployment phase.
[0162] The resulting geographic proximity matrix has rows and columns corresponding to medical institution identifiers, and the matrix cell values represent the proximity scores between the two institutions. If some institution addresses are completed using auxiliary rules, their rows and columns in the matrix will have special markings or weight adjustments. For example, an uncertainty coefficient may be added to the completed coordinates to slightly reduce their proximity scores, so as not to overemphasize the influence of guessed locations.
[0163] After construction, the geographic proximity matrix can be further used for proximity-weighted or regional aggregation analysis. When several institutions are geographically closely distributed and simultaneously exhibit high missing rates or formatting errors, the system will prioritize suspecting whether they belong to the same network node or the same operator's line, sharing similar failure risks. If the matrix shows that the location of individual institutions differs significantly from the main cluster, the evidence coupling between them and the main cluster can be reduced in subsequent evidence synthesis. Through this engineered address resolution, auxiliary completion, and graded proximity calculation, this application can more accurately reflect the true regional relationships of various medical institutions, providing a stable and reproducible foundation for subsequent regional anomaly identification and evidence weight adjustment.
[0164] As an optional implementation, after calculating the geographical proximity between any two target medical institutions based on the institution location information parameters, the method further includes:
[0165] Within the third preset period, business interaction logs of each target medical institution are collected. The business interaction logs record at least the interaction timestamp, the identifier of the other party, the interaction type identifier, and the amount of interaction data.
[0166] For any two target medical institutions, count the number of interactions and the amount of interaction data within the third preset period;
[0167] Based on the number of interactions, the amount of interaction data, and the geographical proximity, the corresponding proximity value is weighted and adjusted to obtain the updated proximity value.
[0168] Based on the updated proximity values, the geographic proximity matrix is constructed.
[0169] This application not only calculates geographical proximity based on the location information of the target medical institution, but also collects and analyzes business interaction logs within the third preset period, and superimposes the number of interactions and the amount of interaction data with the original geographical proximity to present a business-level similarity that is more in line with the actual application scenario in the final geographical proximity matrix.
[0170] This application first configures a third preset period during deployment to periodically retrieve business interaction logs between target medical institutions. An exemplary third preset period can be daily or weekly, depending on the operations and maintenance personnel's assessment of the speed of changes in cross-institutional business collaboration. The business interaction logs must include at least an interaction timestamp, the identifier of the interacting party, an interaction type identifier, and the amount of interaction data, used to characterize the data transmission or service call relationships between different medical institutions within that period. This data can be collected using common big data or log collection tools, such as recording traffic statistics at the edge nodes or VPN gateways of each institution, and then pushing the logs in batches to the central node via a log broker.
[0171] Upon receiving business interaction logs, this application will aggregate statistics for any two target medical institutions, such as accumulating the number of interactions and the total amount of interaction data within a period by connection pairs or traffic pairs. If the interactions between two hospitals or health centers are significantly more frequent or the data volume is larger, the system can consider them to have a higher degree of coupling at the business and network levels, thus giving additional weight to the proximity score. To prevent interference from extreme cases, operations and maintenance personnel can configure upper and lower limits for interaction volume and frequency thresholds during the deployment phase. For example, fewer than ten interactions or less than 100MB of interaction data per period can be considered low-intensity interactions, while more than several GB or hundreds of connections can be considered high-intensity interactions. These intervals can be mapped to the weighting rules in the system configuration file.
[0172] After the statistics are completed, the system first compares the interaction results in parallel based on the previously calculated geographical proximity. If the geographical proximity is high but the actual interaction frequency is low, it means that although the two institutions are physically close, they may not be on the same medical service link or network topology. If the geographical proximity is moderate but the number of interactions and the amount of data are significantly higher than the normal level, it means that although the two institutions are relatively geographically dispersed, they are closely cooperating on data, such as sharing a testing center or imaging.
[0173] To address the aforementioned differences, this application will perform a weighted adjustment on the proximity, which can be implemented using a simple multiplication coefficient or a segment translation method.
[0174] For example, if the original geographical proximity is calculated to be 0.7, but the two organizations have a large amount of interaction within the third preset period, the final proximity can be increased from 0.7 to 0.85; if the physical proximity reaches 0.9, but the business interaction is almost zero, the proximity can be reduced to 0.75. To allow operations and maintenance personnel to configure flexibly, these weighting rules can be stored in the database or configuration file in the form of "(interaction intensity range, geographical proximity range) to adjustment coefficient", and dynamically loaded during program runtime.
[0175] Finally, the system uses these weighted proximity values to reconstruct or update the geographical proximity matrix. Compared to proximity based solely on address resolution, this matrix better reflects the true distance within the medical consortium in terms of network bandwidth sharing, business collaboration, and service invocation, making subsequent detection of regional faults or local anomalies more effective. For example, if a group of medical institutions that are not geographically closest exhibit high interaction volume due to long-term business collaboration, they will be assessed as closer in the final proximity matrix. Before the system synthesizes the results, it can identify that "these institutions constitute a business chain group," making it easier to capture high missing rates or latency phenomena that occur synchronously within this group.
[0176] This optional implementation primarily addresses the "one-size-fits-all" problem that may arise from relying solely on geographic location information. Some geographically close institutions may not necessarily have frequent data interactions, while some geographically distant institutions may have closer business relationships due to specialized collaboration or a unified image cloud. By retrieving business interaction logs within a third preset period and weighting them with geographic proximity, this application achieves a more realistic and dynamic inter-institutional similarity assessment at the engineering level, further improving the accuracy of subsequent trend anomaly detection and the ability to identify cross-regional collaboration anomalies.
[0177] For example, firstly, deploy a set of log collection or traffic probe tools to record connection handshakes and packet statistics on VPN tunnels or message buses between organizations.
[0178] Then, the third preset cycle scheduler runs the summary script once an hour or once a day, storing the summary results in the interaction statistics table. The interaction statistics table may contain fields such as "source organization identifier, target organization identifier, interaction type, total number of times, total data volume, and sampling time range".
[0179] Next, the system references the original geographical proximity value for each pair of institutions, reads the interactive statistics table, and adjusts the proximity by combining preset intervals or function formulas, for example, by making fluctuations within the range of 0.6 to 1.0.
[0180] Finally, the updated proximity values are written back into the proximity matrix, providing a query interface for subsequent evidence synthesis or cluster detection services. This flexible dynamic weight adjustment mechanism ensures that the medical consortium can maintain high accuracy in identifying regional anomalies and providing trend warnings even when network deployments change frequently or new collaborating hospitals join.
[0181] As an optional implementation, the step of stacking the time-series data corresponding to each target medical institution according to the institution dimension to construct a joint feature matrix for the same data quality indicator includes:
[0182] Based on the geographic proximity matrix, target medical institutions with geographic proximity higher than a preset grouping threshold are divided into the same neighborhood cluster according to the proximity aggregation rule;
[0183] For each neighboring cluster, the anomaly count of the target medical institution within the cluster is counted at each time-series index, and a regional synchronization anomaly marker is generated when the anomaly count reaches a preset cluster threshold.
[0184] When the regional synchronization anomaly flag exists, the corresponding time-series data is merged at the cluster level according to the preset merging rules to obtain a cluster-level representative data column.
[0185] The cluster-level representative data column and the time-series data column of the target medical institution outside the cluster are stacked according to the institution dimension and filled into the joint feature matrix.
[0186] This application, when stacking data from multiple institutions for the same data quality indicator, can further divide the data into neighboring clusters based on a geographic proximity matrix and perform column-level merging on the data from each target medical institution. This allows for the identification of "regional synchronization anomalies" before merging, thereby avoiding fitting bias or matrix rank loss caused by multiple columns of completely similar data, and providing more accurate and scalable input for trend analysis in the subsequent "joint feature matrix". This optional implementation can effectively solve the problem that simple stacking will result in "completely linearly correlated multiple columns" when several geographically close medical institutions with similar network environments experience anomalies synchronously within the same time window.
[0187] In practice, after reading the geographic proximity matrix, this application will perform grouping operations according to the proximity aggregation rules.
[0188] For example, if the proximity is higher than 0.8, they are considered to be in the same neighboring cluster; if it is between 0.5 and 0.8, it is processed separately; and if it is lower than 0.5, it is considered to have virtually no direct association. Multiple proximity thresholds can be set in the configuration file or database to adapt to different sizes of medical consortia and geographical distributions. The system centrally processes target medical institutions within the same cluster, facilitating the identification of possible "regional clustering anomalies." If deployment requirements exist, common clustering algorithms, such as hierarchical clustering or DBSCAN, can also be used to automatically cluster the geographical proximity matrix. These algorithms can be used in tools such as Python Scikit-learn or SparkMLlib.
[0189] For each partitioned neighboring cluster, this application counts the anomalies of the target medical institutions within the cluster at each time-series index. Anomaly counting can be performed based on predefined criteria, such as a missing rate exceeding 15%, a format error rate exceeding 3%, or a data acquisition delay exceeding 30 minutes. If multiple institutions within a cluster simultaneously trigger these thresholds at a certain time point or time slice, an anomaly is accumulated. When the accumulated value reaches a preset cluster threshold, the system generates a "regional synchronization anomaly flag," indicating that the neighboring cluster is in a state of synchronization degradation within that time-series index. The cluster threshold can be set to 50% or 60% or more of the total number of institutions in the cluster; the specific value depends on the actual cluster size within the medical consortium and can be optimized by operations and maintenance personnel through configuration files during the deployment phase.
[0190] When the system detects that a neighboring cluster has triggered a regional synchronization anomaly flag, it invokes a preset merging rule to further merge the time-series data column corresponding to that cluster at the column level. For example, in column-level merging, the average or median of multiple columns within a cluster can be calculated separately, or the maximum or minimum value can be selected as the "cluster-level representative," depending on the anomaly sensitivity requirements of the business operation and maintenance scenario. For instance, if a conservative risk is to be highlighted, selecting the maximum value can more sensitively capture degradation situations; if a stable estimate of the mean or median is desired, the impact of extreme values on the overall picture can be weakened. After merging, this application forms a completely new "cluster-level representative data column," which replaces the original multiple columns or adds a new column to the matrix to represent the data of that cluster, avoiding redundant calculations on highly overlapping columns during subsequent multivariate fitting.
[0191] Finally, this application stacks the cluster-level representative data column and the time-series data column of the target medical institutions outside the cluster according to the institution dimension, and fills them into the joint feature matrix. For institutions that are not in the same cluster or have not achieved synchronized anomaly marking, they are still kept as their own independent columns.
[0192] The advantage of this approach is that it improves the ability to identify regional synchronization anomalies while avoiding excessive repetition of rows and columns in engineering, thus reducing redundancy in the fitting process. This approach achieves a good balance between efficiency and analytical accuracy in practical deployment. If dozens of institutions within a medical consortium are closely adjacent and share the same network, and all of them experience a simultaneous increase in synchronization error rate, without merging, the joint feature matrix will contain many almost identical columns. This method allows for the aggregation of these strongly correlated columns, with representative values written into a single column. Subsequent multivariate fitting performs a trend calculation on this merged column, accurately identifying regional fluctuations while significantly reducing the computational and storage overhead of the algorithm.
[0193] For example, the time-series data can be clustered and statistically analyzed using Python scripts or Spark SQL: All pairs of institutions with a proximity of ≥0.8 are identified in the proximity table, and a "neighborhood cluster" index is constructed. For each neighboring cluster, the time-series data table is traversed, and the number of anomalous institutions within the cluster at any given time t is counted. If this number exceeds a certain percentage threshold of the cluster size, such as 50%, a "regional synchronization anomaly marker" is generated at time t. If a synchronization anomaly marker appears, the columns within that cluster are merged: a "column average" or "column maximum value" strategy can be adopted, writing the result to a new cluster representative column, and discarding or retaining the original columns as reference columns in the matrix. These are then finally stacked together with other institution columns that did not trigger synchronization anomalies to form a new joint feature matrix.
[0194] This clustering process primarily addresses the severe multicollinearity problem caused by simultaneous anomalies in clustered institutions within a large-scale medical consortium setting, and also enhances the analytical capabilities at the regional level. If it is indeed necessary to retain fine-grained information, the original columns and representative columns can exist in parallel, with weight reduction or masking flags set during the fitting phase. Flexibility is controlled through the system configuration file. In summary, this optional implementation allows medical consortia to focus on regionally clustered faults during monitoring, reducing the waste of resources from repeatedly analyzing highly homogeneous data, and improving the timeliness and accuracy of trend analysis.
[0195] As an optional implementation, stacking the cluster-level representative data column and the time-series data column of the target medical institution outside the cluster along the institution dimension to fill the joint feature matrix includes:
[0196] For each neighboring cluster, the cluster-level representative data column is used to perform differential processing on the time-series data column of the member institutions within the cluster to generate the corresponding residual data column;
[0197] The cluster-level representative data column and each of the residual data columns are stacked according to the institution dimension, and merged with the time-series data column of the target medical institution outside the cluster to fill the joint feature matrix.
[0198] In an optional embodiment of this application, before stacking the cluster-level representative data column with the time-series data column of the target medical institution outside the cluster, the time-series data column of the member institutions within the cluster is first differentially processed to generate a residual data column. This can effectively solve the matrix rank loss problem caused by high correlation of multiple columns, while retaining the independent deviation information of each member institution relative to the cluster-level representative.
[0199] This application first obtains a cluster-level representative data column for each neighboring cluster. This representative column can be a "cluster-level" time series trajectory generated by the mean, maximum, or other summarization methods, used to summarize the average or extreme value trend of the cluster as a whole. For the time series data column of each member institution within the cluster, this application subtracts the corresponding time value of the representative column point by point, obtaining a residual column. For example, if the representative column records the average missing rate of the cluster under each time index, and the missing rate of a certain institution is significantly higher than the average at the same index, the residual is positive, and its value reflects the specific magnitude by which the missing rate of that institution is higher than the cluster-level representative value at that time. Through this differencing process, the complete linear dependency between the member column and the representative column can be eliminated, reducing the impact of collinearity problems in subsequent multivariate fitting.
[0200] In practical implementation, the system can use distributed data processing tools, such as Spark DataFrame or batch update scripts in traditional databases, to perform row-by-row differencing operations on each column within a cluster. If the representative column is stored in a "cluster_col" field and the member columns are stored in an "institution_col" field, this application can calculate "residual_col = institution_col – cluster_col" and write the result to the same table or a newly created table. To avoid unnecessary growth of long columns, the system can perform the differencing only once within a short window, store the differencing result and the representative column after the calculation is complete, and then release the temporary table to save storage space.
[0201] Subsequently, this application stacks the cluster-level representative data column and each residual data column along the institution dimension, and then merges them with the time-series data columns of other institutions outside the cluster to form a new joint feature matrix. In this matrix, the cluster-level representative column reflects the overall regional trend, while each residual column shows the degree of independent deviation of each institution relative to this trend. Since the linear correlation coefficient between the residual column and the representative column is theoretically significantly reduced, the matrix has higher discernibility, and the fitting algorithm can avoid secondary convergence on the repeated column information of institutions within the same cluster when capturing overall fluctuations.
[0202] This method addresses the situation in large-scale medical consortia where multiple entities are geographically close and exhibit nearly synchronized trends. Directly maintaining highly similar data columns can easily lead to problems such as unidentifiable coefficients and unstable model values during subsequent multivariate regression or iterative fitting. By first generating cluster-level representative columns and then differencing to generate residual columns, this approach not only preserves both the overall trend of the region and the differences between individual institutions but also prevents insufficient rank in the model during training or fitting due to identical values in multiple columns. In other words, this difference-stacking strategy addresses both the detection requirements of global and local anomalies in engineering practice, helping to quickly pinpoint whether the problem is "overall regional anomaly" or "deviation from individual institutions."
[0203] For example, for a neighboring cluster containing five community health centers, this application first uses their average missing rate column to construct a "cluster-level representative data column," and then subtracts each institution's column in turn to obtain five residual data. If the missing rate of one institution is frequently 2 to 3 percentage points higher than the average, the residual column will show a "positive deviation" at multiple points in time.
[0204] In the subsequent joint feature matrix, researchers or fitting algorithms can either treat the representative column as the overall level of the cluster, or use the residual column to discern the persistent error of individual institutions compared to the cluster mean. This allows for faster identification of whether an institution is simply following the overall market trend or exhibiting individual problems during subsequent feature modeling or anomaly diagnosis.
[0205] In this way, while maintaining the ability to summarize regional trends, this implementation method also reduces the modeling difficulty of generating multiple columns of homogeneous information through differential methods, enabling the system to identify local and overall trend problems within the medical consortium more efficiently and accurately based on the joint feature matrix.
[0206] As an optional implementation, a multivariate fitting algorithm is performed based on the joint feature matrix to obtain trend change parameters for each target medical institution, including:
[0207] The joint feature matrix is partitioned according to a preset sliding time window to obtain multiple window feature sub-matrices;
[0208] For each of the aforementioned window feature submatrices, a weighted regression fitting is performed to obtain the corresponding candidate trend parameters and residual confidence values.
[0209] The candidate trend parameters and residual confidence values of each window are used as multi-source evidence information to construct a window-level basic confidence assignment.
[0210] The window-level basic confidence assignment is processed by evidence theory synthesis, and the weight of conflicting evidence information is dynamically adjusted when the degree of conflict exceeds a preset range, generating comprehensive trend change parameters for each target medical institution.
[0211] Furthermore, this application captures dynamic evolution features more finely by partitioning the joint feature matrix according to a preset sliding time window, and performs weighted regression fitting within each time window to obtain candidate trend parameters and residual confidence scores. Subsequently, this application uses these windowed results as multi-source evidence information input into the evidence theory model, dynamically adjusting the weights of evidence information with excessive conflict, and finally outputting the comprehensive trend change parameters for each target medical institution. This method is engineering-compatible with various data processing and machine learning frameworks, such as using Python's pandas DataFrame or Spark RDD partitioning operations to perform sliding window partitioning of the matrix, or leveraging third-party regression modules such as statsmodels or scikit-learn to perform weighted fitting within each window.
[0212] In its implementation, this application first partitions the joint feature matrix according to a preset sliding time window. An exemplary approach is to set a window length and step size in the configuration file, such as stepping 10 data points out of 30 to achieve sliding, thus obtaining multiple "window feature sub-matrices". If the medical consortium's business cycle is long, timestamps can also be used for partitioning: intervals of one hour, one day, or one week can be flexibly configured to adapt to trend observations of different granularities. In handling large-scale medical consortium scenarios, the system can use a distributed computing engine to partition the matrix and distribute it to worker nodes for parallel regression calculations.
[0213] Within each window's feature submatrix, this application performs weighted regression fitting to obtain candidate trend parameters and residual confidence values. The weighting method can be based on the number of valid samples within the window, coverage rate, or anomaly percentage, or it can refer to the weighting results of geographical proximity, incorporating them into the regression loss function. Once the regression model converges iteratively, this application outputs a candidate trend parameter vector (for all medical institutions or all columns) and the corresponding residual confidence values. The residual confidence value can be used to measure the extent to which the window's fit matches the real data; a higher value generally indicates a better interpretability of the current window's data. A low residual confidence value indicates significant abnormal fluctuations or potential conflicts within the window.
[0214] After completing regressions across multiple windows, this application treats the candidate trend parameters and residual confidence scores generated by each window as multi-source evidence information, constructing window-level basic confidence assignments. Specifically, the trend parameters are broken down into support levels for propositions such as "deterioration" and "normal." Alternatively, parameter ranges (e.g., high slope, low slope, or near zero) can be mapped to different propositions. The confidence values of the corresponding propositions are then amplified or reduced based on the residual confidence scores, while retaining a margin for unmapped propositions. This forms a set of basic confidence assignments covering all windows.
[0215] To integrate all window results, this application will multiply the basic confidence assignments at the window level one by one and accumulate the conflict degree during the evidence theory synthesis process. If a certain window confidence assignment significantly contradicts the mainstream result, causing the conflict degree to exceed the preset range, the system will automatically reduce the weight corresponding to that window.
[0216] For example, with an initial weight of 1, when a conflict occurs, the weight can be reduced to 0.5 for another synthesis. If the conflict is still too large, it can be further reduced to 0.3. If the conflict remains high after multiple rounds of reduction, the system will treat the window result as a serious anomaly and allocate most of its probability to uncertain propositions. After the synthesis is completed, each target medical institution can obtain a comprehensive trend change parameter vector, representing the overall trend estimate after aggregation across multiple windows.
[0217] This alternative implementation primarily addresses the issue of excessive interference from outlier windows in a single global fit. In actual medical consortium operations, certain time windows may differ significantly from other windows due to network jitter, sudden drops in business volume during holidays, or unexpected equipment failures. If regression is performed directly on the full dataset, extreme windows will impose illogical biases on the overall model. By synthesizing candidate trend parameters and residual confidence scores from multiple windows using evidence theory, this application can dynamically adjust window-level outlier results in engineering practice, preserving temporal locality while suppressing the contamination of the global model by extreme windows. As a result, the final generated comprehensive trend change parameters better represent the true situation of most windows within the medical consortium, and provide visualized conflict levels or uncertain proposition assignments for individual outlier windows, enabling targeted and in-depth investigations during subsequent audits or maintenance.
[0218] Based on the same inventive concept, this disclosure also provides a data management system for medical consortia that corresponds to the data management method for medical consortia. Since the principle of the system in this disclosure for solving problems is similar to the data management method for medical consortia described above, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be described again.
[0219] Reference Figure 4 The diagram shown is a schematic of a data management system for medical consortia provided in an embodiment of this application. The system includes:
[0220] The data acquisition module 10 is used to acquire data acquisition task logs of each target medical institution according to a first preset period. The data acquisition task logs include: acquisition task identifier, acquisition time, number of acquisition items, acquisition result and acquisition status.
[0221] The first processing module 20 is used to statistically analyze the data quality indicators of each target medical institution based on the data acquisition task log, and generate time-series data of each data quality indicator according to a second preset period. The data quality indicators include: data missing rate, format error rate and data acquisition delay.
[0222] The second processing module 30 is used to stack the time series data corresponding to each target medical institution according to the institution dimension for the same data quality index, construct a joint feature matrix, and perform a multivariate fitting algorithm based on the joint feature matrix to obtain the trend change parameters of each target medical institution.
[0223] The detection module 40 is used to generate a degradation trend detection result based on the trend change parameter and a preset judgment condition; based on the degradation trend detection result, it generates a trend warning information and outputs the trend warning information to the management terminal.
[0224] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
Claims
1. A data management method for medical consortia, characterized in that: include: According to the first preset cycle, acquire the data collection task logs of each target medical institution and the corresponding server node operation status logs; The server node's running status log includes: node identifier, node heartbeat information, node resource utilization, and node availability status; Based on the data acquisition task log and the server node's operation status log, the data quality indicators and server node operation status indicators of each target medical institution are statistically analyzed, and time-series data of each data quality indicator and operation status indicator are generated according to a second preset period; the operation status indicators include: node failure rate and abnormal resource occupancy rate. For the same data quality indicator, the time series data corresponding to each target medical institution are stacked according to the institution dimension to construct a joint feature matrix, and a multivariate fitting algorithm is executed based on the joint feature matrix to obtain the trend change parameters of each target medical institution. Based on the trend change parameters and the operating status indicators of the server nodes, a degradation trend detection result is generated according to preset judgment conditions; based on the degradation trend detection result, a trend warning information is generated and the trend warning information is output to the management terminal.
2. The data management method for medical consortia according to claim 1, characterized in that, The management terminal is also used for: It displays the real-time running status information of the corresponding server nodes of each target medical institution and the execution status information of the data collection tasks; The operational status information includes: node availability status, node resource utilization rate, and node fault indication; the execution status information includes the current task status, task execution progress, and task execution result. The management terminal is also used to: provide task management functions based on the execution status information of the data acquisition task and the running status information of the server node; the task management functions include task start / stop control, task priority adjustment and task retry strategy management.
3. The data management method for medical consortia according to claim 1, characterized in that, The data quality indicators include: data missing rate, format error rate, and data acquisition delay; the data acquisition task log includes: acquisition task identifier, acquisition time, number of acquisition items, acquisition results, and acquisition status. Based on the aforementioned trend change parameters and preset judgment conditions, a degradation trend detection result is generated, including: Based on the trend change parameters of each target medical institution, generate corresponding single-institution judgment result information; The single-institution judgment result information is used as multi-source evidence information input, and evidence theory is used for synthesis processing to obtain a comprehensive confidence result for the overall data collection situation; The degree of conflict of the comprehensive confidence result is calculated, and when the degree of conflict exceeds a preset range, the weight of the conflicting evidence information is adjusted to obtain the corrected comprehensive confidence result. By comparing the corrected overall confidence result with the preset confidence threshold, the abnormal confidence level at the medical consortium level is determined; The degradation trend detection results are generated based on the abnormal confidence level and the single-institution judgment results of each target medical institution.
4. The data management method for medical consortia according to claim 3, characterized in that, The comprehensive confidence result obtained for the overall data collection situation includes: Based on the individual assessment results of each target medical institution, an initial basic confidence allocation is constructed; Calculate the degree of conflict between the initial basic confidence assignments, and when the degree of conflict exceeds a preset range, adjust the weight of conflict evidence information to update the basic confidence assignments; An iterative synthesis process is performed on the updated basic confidence assignment to obtain a first synthesized confidence level and a second synthesized confidence level; wherein, the first synthesized confidence level is used to characterize that the data acquisition quality is in a deteriorating trend, and the second synthesized confidence level is used to characterize that the data acquisition quality is in a normal trend. Based on the first composite confidence level and the second composite confidence level, the uncertainty entropy value is determined; Based on the first composite confidence level, the second composite confidence level, and the uncertainty entropy value, a comprehensive confidence result is generated; The uncertainty entropy value is used to characterize the degree of consistency of the overall evidence of the medical consortium.
5. The data management method for medical consortia according to claim 4, characterized in that, Also includes: A geographic proximity matrix is generated based on the geographic location information between the target medical institutions. Based on the geographic proximity matrix, the single-institution determination results of each target medical institution are subjected to proximity-based weighted processing. The weighted single-institution determination results are used to generate the initial basic confidence assignment. The generation of the geographic proximity matrix based on the geographic location information between each target medical institution includes: Geographic parsing is performed on the address information of each target medical institution to obtain the corresponding institution location information parameters; When the address of any target medical institution cannot be successfully resolved or the confidence level of the corresponding institution location information parameter is lower than a preset threshold, the institution location information parameter is supplemented based on a predefined auxiliary location information generation rule. Based on the location information parameters of the institutions, the geographical proximity between any two target medical institutions is calculated, and the geographical proximity matrix is constructed.
6. The data management method for medical consortia according to claim 5, characterized in that, After calculating the geographical proximity between any two target medical institutions based on the aforementioned institutional location information parameters, the method further includes: Within the third preset period, business interaction logs of each target medical institution are collected. The business interaction logs record at least the interaction timestamp, the identifier of the other party, the interaction type identifier, and the amount of interaction data. For any two target medical institutions, count the number of interactions and the amount of interaction data within the third preset period; Based on the number of interactions, the amount of interaction data, and the geographical proximity, the corresponding proximity value is weighted and adjusted to obtain the updated proximity value. Based on the updated proximity values, the geographic proximity matrix is constructed.
7. The data management method for medical consortia according to claim 6, characterized in that, The method of stacking the time-series data corresponding to each target medical institution according to the institution dimension to construct a joint feature matrix for the same data quality indicator includes: Based on the geographic proximity matrix, target medical institutions with geographic proximity higher than a preset grouping threshold are divided into the same neighborhood cluster according to the proximity aggregation rule; For each neighboring cluster, the anomaly count of the target medical institution within the cluster is counted at each time-series index, and a regional synchronization anomaly marker is generated when the anomaly count reaches a preset cluster threshold. When the regional synchronization anomaly flag exists, the corresponding time-series data is merged at the cluster level according to the preset merging rules to obtain a cluster-level representative data column. The cluster-level representative data column and the time-series data column of the target medical institution outside the cluster are stacked according to the institution dimension and filled into the joint feature matrix.
8. The data management method for medical consortia according to claim 7, characterized in that, Stacking the cluster-level representative data column and the time-series data column of the target medical institution outside the cluster according to the institution dimension, and filling them into the joint feature matrix includes: For each neighboring cluster, the cluster-level representative data column is used to perform differential processing on the time-series data column of the member institutions within the cluster to generate the corresponding residual data column; The cluster-level representative data column and each of the residual data columns are stacked according to the institution dimension, and merged with the time-series data column of the target medical institution outside the cluster to fill the joint feature matrix.
9. The data management method for medical consortia according to claim 8, characterized in that, Based on the joint feature matrix, a multivariate fitting algorithm is executed to obtain the trend change parameters of each target medical institution, including: The joint feature matrix is partitioned according to a preset sliding time window to obtain multiple window feature sub-matrices; For each of the aforementioned window feature submatrices, a weighted regression fitting is performed to obtain the corresponding candidate trend parameters and residual confidence values. The candidate trend parameters and residual confidence values of each window are used as multi-source evidence information to construct a window-level basic confidence assignment. The window-level basic confidence assignment is processed by evidence theory synthesis, and the weight of conflicting evidence information is dynamically adjusted when the degree of conflict exceeds a preset range, generating comprehensive trend change parameters for each target medical institution.
10. A data management system for medical consortia, characterized in that: include: The data acquisition module is used to acquire data acquisition task logs and corresponding server node operation status logs for each target medical institution according to a first preset cycle. The server node's running status log includes: node identifier, node heartbeat information, node resource utilization, and node availability status; The first processing module is used to calculate the data quality indicators and server node operation status indicators of each target medical institution based on the data acquisition task log and the server node operation status log, and generate time-series data of each data quality indicator and operation status indicator according to a second preset period; the operation status indicators include: node failure rate and abnormal resource occupancy rate. The second processing module is used to stack the time series data corresponding to each target medical institution according to the institution dimension for the same data quality index, construct a joint feature matrix, and perform a multivariate fitting algorithm based on the joint feature matrix to obtain the trend change parameters of each target medical institution. The detection module is used to generate a degradation trend detection result based on the trend change parameters and the operating status indicators of the server node, according to preset judgment conditions; based on the degradation trend detection result, generate trend warning information, and output the trend warning information to the management terminal.
Citation Information
Patent Citations
Medical management system based on diagnosis and treatment logs
CN116344012A
Management method and system for medical poeia and medical data of medical conjunct platform
CN118093196A
Management method and device of distributed acquisition system facing supercomputing Internet, and storage medium
CN120371614A
Adaptive data processing optimization method and device, equipment and medium
CN120743705A