Newborn disease risk assessment method based on big data analysis
By constructing an event-driven risk assessment map, integrating multi-source heterogeneous data and performing real-time calibration, the static and isolated problems of neonatal disease risk assessment systems are solved, enabling coherent tracking and personalized intervention of neonatal health risks, and improving the accuracy of assessment and the adaptability of intervention.
Patent Information
- Application Number
- CN202610024752.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-02-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing neonatal disease risk assessment systems rely on isolated data points and static thresholds, which fail to form a coherent understanding of the patient's overall condition and lack the ability to make real-time dynamic adjustments, resulting in insufficient assessment accuracy and intervention adaptability.
We construct an event-driven risk assessment map, integrate multi-source heterogeneous data, identify stable and evolving risk clusters through a risk tracking process, derive hierarchical risk management plans, and integrate clinical feedback at the end of each cycle to calibrate the map and dynamically reorganize risk management measures.
It enables continuous tracking and adaptive management of newborn health risks, improves the timeliness of risk identification and the accuracy of intervention, and can identify complex risk patterns earlier and adapt to individualized changes in risk situation.
Smart Images

Figure CN121506503A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of neonatal medical big data risk assessment technology, specifically a method for neonatal disease risk assessment based on big data analysis. Background Technology
[0002] Neonatal disease risk assessment primarily relies on the monitoring and analysis of single or limited data sources, such as independently interpreting imaging reports or processing structured laboratory indicators. The multimodal data generated by various medical information systems is fragmented, and assessments are often based on isolated data points or static threshold triggers, failing to form a coherent understanding of the patient's overall condition. This analytical model lags behind actual clinical progress, making it difficult to capture the relationships between different risk factors and their dynamic evolution over time.
[0003] Existing risk warning models are mostly built based on historical data, and their rules and parameters are usually fixed after deployment. These systems cannot incorporate real-time clinical feedback to correct their judgments, leading to a disconnect between the model and actual clinical progress. Corresponding intervention plans are also mostly pre-set static plans, lacking the ability to dynamically adjust based on real-time changes in individual patient risk. Static systems have limitations in the accuracy of their assessments and the adaptability of their interventions when facing the complex evolution of neonatal conditions.
[0004] The purpose of this invention is to solve the above-mentioned problems by constructing a dynamic analysis framework that can integrate multi-source heterogeneous data and continuously learn, so as to achieve coherent tracking and adaptive management of newborn health risks. Summary of the Invention
[0005] The purpose of this invention is to provide a method for assessing neonatal disease risk based on big data analysis, so as to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides a method for neonatal disease risk assessment based on big data analysis, the method comprising:
[0007] Initial multimodal health records of newborns are obtained from multiple heterogeneous medical data sources. These initial multimodal health records include image files, time-series vital sign streams, structured laboratory reports, and unstructured medical texts.
[0008] The initial multimodal health records are fused and analyzed to construct an event-driven risk assessment graph, which consists of entity nodes, relation edges, and multidimensional state vectors attached to the entity nodes.
[0009] Initiate a continuous risk tracking process, using the aforementioned risk assessment map to perform hierarchical risk scanning, and identify stable risk clusters, evolving risk clusters, and isolated risk signals;
[0010] Based on the spatial distribution and temporal evolution patterns of the stable risk clusters and the evolving risk clusters, a hierarchical risk management plan is derived, which includes an active monitoring layer, a preparatory response layer, and an immediate handling layer.
[0011] At the end of each cycle of the risk tracking process, the latest clinical treatment feedback and monitoring readings are integrated to perform reverse calibration on the entity node status and relation edge strength in the risk assessment map.
[0012] Based on the reverse-calibrated risk assessment map, the implementation sequence of measures and resource allocation in the hierarchical risk management plan are dynamically reorganized.
[0013] Preferably, the specific implementation of fusing and analyzing the initial multimodal health records to construct an event-driven risk assessment map includes:
[0014] The time-series vital sign stream is segmented, and the continuous waveform data is divided into multiple physiological event segments based on physiological event markers, and a waveform feature vector is generated for each physiological event segment.
[0015] Named entity recognition technology is used to extract clinical entities and medical events from unstructured medical texts. The extracted clinical entities and medical events are then aligned and associated with fields in structured test reports to form text feature vectors.
[0016] The image file is input into a pre-trained deep learning network to extract image feature vectors;
[0017] A unified event index is created, and waveform feature vectors, text feature vectors, and image feature vectors that occur within the same time window are concatenated and their dimensions reduced to generate a fused event vector for the time window.
[0018] The clinical entities, medical events, and physiological event fragments are used as entity nodes, and the causal relationships, temporal relationships, and statistical symbiotic relationships between entity nodes are used as relation edges.
[0019] The fused event vector is used as the initial multidimensional state vector of the core entity node within the corresponding time window to complete the construction of the event-driven risk assessment map.
[0020] Preferably, the specific implementation of initiating a continuous risk tracking process and using the risk assessment map for hierarchical risk scanning includes:
[0021] On the risk assessment graph, a graph traversal of a specified depth is performed, starting from an entity node with an abnormal multidimensional state vector.
[0022] During the traversal, all visited entity nodes and their connected relationship edges are collected to form a candidate risk subgraph;
[0023] The candidate risk subgraph is subjected to graph structure metric calculation to obtain the clustering coefficient, average path length and node degree distribution of the candidate risk subgraph;
[0024] The graph structure metric of the candidate risk subgraph is matched with a predefined graph pattern template, which includes star diffusion pattern, chain transmission pattern and network interweaving pattern.
[0025] If a candidate risk subgraph successfully matches a star-shaped diffusion pattern, the candidate risk subgraph is marked as a stable risk cluster;
[0026] If a candidate risk subgraph matches a chain-like transmission pattern, the candidate risk subgraph is marked as an evolutionary risk cluster.
[0027] If a candidate risk subgraph cannot match any predefined graph pattern template, but the outlier value of the multidimensional state vector of the entity node it contains exceeds the independent threshold, then the entity node is marked as an isolated risk signal.
[0028] All identified stable risk clusters, evolutionary risk clusters, and isolated risk signals are added to the risk tracking list, and each entry is appended with its corresponding spectral structure identifier and anomaly intensity summary.
[0029] Preferably, the specific steps for deriving a hierarchical risk management plan based on the spatial distribution and temporal evolution patterns of the stable risk cluster and the evolving risk cluster include:
[0030] For each stable risk cluster in the risk tracking list, extract the historical multidimensional state vector sequence of all entity nodes contained therein;
[0031] Time series analysis is performed on the historical multidimensional state vector sequence to fit its changing trend, and the state prediction vector for the next period is extrapolated.
[0032] Compare the state prediction vector with the preset warning thresholds of each clinical indicator to generate a list of future risk sites;
[0033] Based on the spatiotemporal density of the future risk site list, a deployment scheme for the active monitoring layer is generated, which specifies the monitoring points, monitoring frequency, and data transmission protocol.
[0034] For each evolutionary risk cluster in the risk tracking list, analyze its graph structure identifier and identify the risk transmission path with the fastest growth in relation edge strength;
[0035] Simulate the propagation process of abnormal signals along the risk transmission path and calculate the estimated time for the signal to reach key physiological nodes;
[0036] Based on the estimated urgency of the time, different levels of early warning response rules and personnel dispatch plans are configured in the prepared response layer;
[0037] By integrating the deployment scheme of the active monitoring layer with the early warning response rules of the preparatory response layer, an immediate handling layer plan is formed, which includes specific operation instructions, triggering conditions and execution entities, together constituting a hierarchical risk management plan.
[0038] Preferably, the specific implementation of integrating the latest clinical treatment feedback and monitoring readings at the end of each cycle of the risk tracking process to perform reverse calibration of the entity node status and relation edge strength in the risk assessment map includes:
[0039] At the end of the preset risk assessment period, collect all clinical treatment records and equipment monitoring readings generated during the implementation of the hierarchical risk management plan within the risk assessment period;
[0040] Analyze clinical treatment records to extract treatment measures, target entity nodes, and record effect evaluations;
[0041] Convert equipment monitoring readings into updated waveform feature vectors or text feature vectors;
[0042] Locate the target entity node in the risk assessment map and quantify the recorded effect evaluation into an effect intensity coefficient;
[0043] The effect intensity coefficient is used to weight and adjust the current multidimensional state vector of the target entity node to reduce the value of the abnormal feature dimension.
[0044] Traverse the neighboring entity nodes directly connected to the target entity node through relational edges, and propagate the effect strength coefficient attenuation according to the type and strength of the relational edges, and fine-tune the multidimensional state vector of the neighboring entity nodes accordingly.
[0045] The updated waveform feature vector or text feature vector is fused with the multidimensional state vector of the corresponding entity node after effect propagation adjustment to generate a new multidimensional state vector of the corresponding entity node, thereby completing the reverse calibration of node state and edge strength in the risk assessment map.
[0046] Preferably, the specific implementation of dynamically reorganizing the measure execution sequence and resource allocation in the hierarchical risk management plan based on the reverse-calibrated risk assessment map includes:
[0047] After the reverse calibration is completed, based on the new multidimensional state vector of the entity nodes, the anomaly intensity summary of all stable risk clusters and evolutionary risk clusters is recalculated.
[0048] Compare the recalculated anomaly intensity summary with the values from the previous period, and calculate its rate of change and direction of change;
[0049] For risk clusters whose abnormal intensity summary weakens beyond the success threshold, reduce their priority in the hierarchical risk management plan, lower the monitoring frequency of their corresponding active monitoring layer deployment plan, or lower the early warning level of the preparatory response layer.
[0050] For abnormal intensity summary enhancements or newly emerging risk clusters, their priority is increased, their corresponding operation instructions are placed in advance in the immediate handling layer plan, and additional computing and storage resources are allocated to the risk clusters for deep map analysis.
[0051] Based on the adjusted priorities, all measures to be implemented in the tiered risk management plans are reordered to form a new sequence of measures to be implemented.
[0052] Based on the new measures execution sequence and the resource allocation results of each risk cluster, the resource scheduling instructions and task distribution list for the next cycle are generated.
[0053] Preferably, the specific steps of inputting the image file into a pre-trained deep learning network to extract image feature vectors include:
[0054] A convolutional neural network pre-trained on a large medical image dataset is used as the backbone network for feature extraction;
[0055] The input neonatal image files are subjected to standardized preprocessing, including grayscale normalization, size scaling, and region of interest cropping.
[0056] The preprocessed image data is input into the pre-trained convolutional neural network, and high-dimensional feature tensors are extracted from specific intermediate layers of the network during the forward propagation process.
[0057] The extracted high-dimensional feature tensor is subjected to global average pooling to compress it into a fixed-length image feature vector.
[0058] The image feature vector is compared with the waveform feature vector and text feature vector obtained from the same newborn at the same time point at the feature level. If a significant contradiction is found, the feature review process is triggered to re-evaluate the quality and annotation of the image file.
[0059] Preferably, the specific steps for calculating the graph structure metric of the candidate risk subgraph include:
[0060] Calculate the shortest path length between all entity nodes in the candidate risk subgraph, and take the arithmetic mean of these shortest path lengths as the average path length of the candidate risk subgraph.
[0061] The number of directly connected neighbor nodes of each entity node in the candidate risk subgraph is counted to obtain the node degree, and a histogram of the node degree distribution is plotted.
[0062] For each entity node in the candidate risk subgraph, calculate the ratio of the actual number of relational edges between its neighboring nodes to the maximum possible number of relational edges, and then take the average of the ratios for all entity nodes to obtain the clustering coefficient of the candidate risk subgraph.
[0063] The calculated average path length, node degree distribution, and clustering coefficient are packaged together with the number of entity nodes and relation edges of the candidate risk subgraph to form a graph structure metric data package for the candidate risk subgraph, which is used for subsequent graph pattern matching.
[0064] Preferably, the specific rules for attenuating and propagating the effect intensity coefficient include:
[0065] Identify the type of relationship edge between the connected entity node and its neighboring entity nodes, wherein the relationship edge type includes strong causality, weak correlation, and temporal accompaniment;
[0066] A propagation attenuation factor is preset for each type of relation edge. Relation edges with strong causality correspond to a larger propagation attenuation factor, while relation edges with weak correlation correspond to a smaller propagation attenuation factor.
[0067] Starting from the entity node of the target, obtain its original effect intensity coefficient;
[0068] Multiply the original effect strength coefficient by the propagation attenuation factor corresponding to the relationship edge type to obtain the attenuated effect strength coefficient propagated to the neighboring entity node.
[0069] The attenuation effect strength coefficient is used to numerically adjust the dimension in the multidimensional state vector of the neighboring entity node that is associated with the abnormal features of the target entity node.
[0070] If a neighboring entity node further connects to other nodes, the attenuated effect strength coefficient after propagation to the neighboring entity node is used as the new starting strength. Based on the new relationship edge type and the corresponding propagation attenuation factor, multi-level attenuation propagation is carried out until the effect strength coefficient is lower than the propagation termination threshold.
[0071] Preferably, the specific implementation of using the clinical entities, medical events, and physiological event fragments as entity nodes, and using the causal relationships, temporal relationships, and statistical symbiotic relationships between entity nodes as relation edges includes:
[0072] Clinical entities and medical events extracted from the unstructured diagnostic text, as well as physiological event fragments segmented from the time-series vital signs stream, are instantiated as unique entity nodes in the risk assessment graph, respectively.
[0073] Evidence of pairwise relationships between entity nodes is extracted from field associations in structured lab reports, semantic co-occurrence in unstructured medical texts, and timestamp alignment in multimodal data.
[0074] Based on the causal rules of the medical knowledge base, causal relationship edges are created for paired entity nodes with causal logic, and the causal relationship edges are assigned initial weights that reflect their causal determinism.
[0075] Based on the timestamp order in the unified event index, time sequence relationship edges are created for pairs of entity nodes that are related in the time series, and the time sequence relationship edges are assigned an initial weight reflecting their temporal proximity.
[0076] Based on big data statistical analysis of historical multimodal health records, statistical co-occurrence relationship edges are created for pairs of entity nodes that frequently co-occur statistically but have no clear causal or temporal constraints, and the statistical co-occurrence relationship edges are assigned initial weights that reflect their co-occurrence frequency.
[0077] The created causal relationship edges, temporal relationship edges, and statistical symbiotic relationship edges are associated with the corresponding source entity nodes and target entity nodes to construct a risk assessment graph network structure that includes heterogeneous entities and multiple relationships.
[0078] Compared with the prior art, the beneficial effects of the present invention are:
[0079] By acquiring and fusing information from multiple heterogeneous data sources, an event-driven risk assessment atlas was constructed. This atlas consists of entity nodes, relational edges, and multidimensional state vectors attached to the nodes. This allows previously scattered image features, temporal signal trends, textual descriptions of clinical findings, and laboratory values to be correlated and quantified within a unified semantic network. The event-driven mechanism ensures that the atlas is updated in sync with clinical diagnostic and treatment activities, connecting discrete data points into an event sequence reflecting pathophysiological processes. This enables deep semantic fusion and alignment of cross-modal information, elevating risk assessment from judging isolated abnormal indicators to dynamically depicting the overall state evolution pattern based on clinical event chains. Risk assessment is transformed from static, lagging snapshot-style analysis to a continuous, accompanying understanding of the child's clinical progress, enabling earlier identification of complex risk patterns caused by the interplay of multiple factors.
[0080] In the risk tracking process, the latest clinical treatment feedback and monitoring readings are integrated to perform reverse calibration on the entity node status and relation edge strength in the risk assessment map. A closed-loop learning pathway from clinical practice feedback to the risk model is established. The actual effects of clinical interventions are fed back in the form of data, continuously revising the estimates of risk entity status and the weights of the correlation strength between entities in the map, enabling the model's knowledge representation to continuously evolve with actual case data. Based on the reverse-calibrated dynamic map, the implementation sequence of measures and resource allocation in the hierarchical risk management plan are reorganized. This makes risk response strategies no longer fixed scripts, but rather real-time calculations and outputs based on the model's latest and most individualized understanding of the current risk structure, intensity, and potential evolution. The system has the ability to self-optimize based on actual effects, and its output plans can closely adapt to the patient's real-time and personalized risk situation, improving the timeliness and accuracy of interventions. Attached Figure Description
[0081] Figure 1 This is a schematic diagram illustrating the working principle of the neonatal disease risk assessment method based on big data analysis described in this invention.
[0082] Figure 2 A flowchart for constructing the integrated analysis and risk assessment map;
[0083] Figure 3 A flowchart for deriving a tiered risk management plan;
[0084] Figure 4 A bar chart showing the relationship between early warning levels and trigger thresholds in neonatal disease risk assessment;
[0085] Figure 5 This is a chart comparing the number of measures and resource allocation at different levels in the neonatal disease risk management plan. Detailed Implementation
[0086] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0087] Please see Figure 1This invention provides a method for neonatal disease risk assessment based on big data analysis. The method includes: acquiring initial multimodal health records of target newborns from multiple heterogeneous medical data sources such as hospital information systems, monitoring equipment, and laboratory information systems. These records encompass image files, continuous time-series vital sign streams, structured laboratory test data, and unstructured doctor's medical records. The acquired initial multimodal health records are fused and analyzed to construct an event-driven risk assessment graph. The graph's structure consists of nodes representing various medical entities, relational edges representing inter-entity relationships, and multidimensional state vectors attached to the nodes, characterizing the entity's state. After the graph is constructed, a continuous, cyclical risk tracking process is initiated, using the constructed risk assessment graph to perform hierarchical risk scanning, identifying risk clusters with stable structures, risk clusters exhibiting evolutionary trends, and independent abnormal risk signals. Based on the spatial and temporal patterns of the identified stable and evolving risk clusters, the system derives a hierarchical risk management plan. This plan typically includes an active monitoring layer for routine monitoring, a preparatory response layer for preparedness, and an immediate intervention layer for emergency intervention. At the end of each pre-set risk assessment cycle, clinical treatment feedback and newly added monitoring data generated during the cycle are integrated to perform reverse calibration on the state vectors of entity nodes and the connection strength of relation edges in the risk assessment map. Based on the reverse-calibrated risk assessment map reflecting the latest health status, the execution sequence and corresponding resource allocation of various measures in the hierarchical risk management plan are dynamically reorganized.
[0088] In one embodiment of the present invention, see [reference] Figure 2The system segments the continuous temporal vital signs stream, dividing the waveform data into multiple independent physiological event segments based on predefined physiological event markers, and generating a waveform feature vector for each segment. Named entity recognition technology is used to process unstructured medical text, extracting mentioned clinical entities and medical events. These extraction results are then aligned and associated with specific fields in structured laboratory reports to form text feature vectors. Newborn image files are input into a pre-trained deep learning network model, from which image feature vectors are extracted. A unified event index timeline is created, concatenating waveform feature vectors, text feature vectors, and image feature vectors occurring within the same time window, and using dimensionality reduction techniques to generate a fused event vector representing the overall situation of that time window. The clinical entities and medical events extracted from the text, as well as the physiological event segments segmented from the vital signs stream, are instantiated as unique entity nodes in the risk assessment atlas. Evidence of pairwise relationships between entity nodes is extracted from field associations in structured lab reports, semantic co-occurrence information in unstructured medical texts, and alignment relationships of multimodal data on timestamps. Based on causal rules defined in the medical knowledge base, causal relationship edges are created for entity node pairs with causal logical connections, and these edges are assigned an initial weight reflecting their causal certainty. Based on the timestamp order in the unified event index, temporal relationship edges are created for entity node pairs with a clear chronological order in the time series, and these edges are assigned an initial weight reflecting their temporal proximity. Based on large-scale statistical analysis results of historical multimodal health records, statistical co-occurrence relationship edges are created for entity node pairs that frequently co-occur statistically but lack clear causal or strict temporal constraints, and these edges are assigned an initial weight reflecting their co-occurrence frequency. The generated fused event vectors are used as the initial multidimensional state vectors of the core entity nodes within the corresponding time window, thus completing the construction of an event-driven risk assessment graph network containing heterogeneous entity nodes, multi-dimensional relationship edges, and node state vectors.
[0089] In practical implementation, taking a newborn with respiratory distress symptoms as an example, the initial multimodal health record of the newborn is obtained from the hospital information system. The initial multimodal health record includes a 24-hour continuous electrocardiogram (ECG) vital signs stream, a chest X-ray image file, a structured blood routine test report, and an unstructured medical text record. In practice, the continuous ECG waveform is segmented into multiple cardiac event segments based on cardiac beat markers, and heart rate variability index and waveform morphology parameters are calculated for each cardiac event segment to generate a waveform feature vector. From the unstructured medical text record, named entity recognition technology is used to extract clinical entities such as "increased respiratory rate" and "decreased blood oxygen saturation" and medical events such as "suspected lung infection". These extraction results are aligned and associated with the "white blood cell count" field in the structured blood routine test report to form a text feature vector containing numerical and semantic information. The chest X-ray image file is input into a pre-trained convolutional neural network, and high-dimensional feature tensors are extracted from the middle layer of the network and generated into an image feature vector through global average pooling. In some embodiments, the waveform feature vector has a dimension of 128, the text feature vector has a dimension of 64, and the image feature vector has a dimension of 256. These dimensions are determined through training with historical data. Data comparison shows that the ECG waveform feature vectors of different newborns have significant differences in disease states. When extracting clinical entities and medical events from unstructured medical records, the named entity recognition model adopts a combined architecture of bidirectional long short-term memory network and conditional random field. A unified event index is created, and time windows are divided in minutes. The waveform feature vector, text feature vector, and image feature vector occurring within the same minute time window are concatenated to obtain a 448-dimensional concatenated vector. Subsequently, principal component analysis is used to reduce the dimension to 100 to generate a fused event vector for the time window. The formula for generating the fused event vector is:
[0090] ;
[0091] in: Represents the fusion event vector. This represents the dimensionality reduction function in principal component analysis. Represents the waveform feature vector. Represents the text feature vector. Represents the image feature vector, symbol This indicates a vector concatenation operation.
[0092] Using extracted clinical entities "increased respiratory rate," medical events "suspected lung infection," and cardiac event fragments as entity nodes, this study extracts the statistical association between "white blood cell count" and "lung infection" from the field associations of structured lab reports, the co-occurrence relationship between "increased respiratory rate" and "decreased blood oxygen saturation" from the semantic co-occurrence of unstructured medical texts, and the temporal sequence relationship between cardiac event fragments and clinical entities from the timestamp alignment of multimodal data. Based on causal rules from a medical knowledge base, a causal relationship edge is created between "lung infection" and "increased respiratory rate" with an initial weight of 0.8. Based on the timestamp sequence of a unified event index, a temporal relationship edge is created between "abnormal cardiac activity" and "decreased blood oxygen saturation" with an initial weight of 0.6. Based on big data statistical analysis of historical multimodal health records, a statistical co-occurrence relationship edge is created between "elevated white blood cell count" and "lung infection" with an initial weight of 0.7. In some embodiments, the initial weights of the relation edges are calculated using a Bayesian network. Data comparison shows that the weights calculated using a Bayesian network better reflect the true risk associations than simple frequency statistics. The structure of the Bayesian network is built based on expert knowledge, and the parameters are learned from a dataset containing 100,000 historical records. The fused event vector is used as the initial multidimensional state vector for the core entity node "suspected lung infection" within the corresponding time window, completing the construction of an event-driven risk assessment graph. The risk assessment graph contains multiple entity nodes, relation edges, and multidimensional state vectors attached to the entity nodes. Optionally, the unique identifier of the entity node adopts the unified medical language system vocabulary encoding to ensure that each clinical entity, medical event, and physiological event fragment has global uniqueness in the risk assessment graph. Data comparison shows that the unique identifier avoids entity duplication and ambiguity.
[0093] In one embodiment of the invention, a convolutional neural network pre-trained on a large publicly available medical image dataset is used as the backbone network for feature extraction. Standardized preprocessing is performed on the input neonatal image files. This preprocessing includes normalizing the image grayscale values to a specific range, scaling the image size to the input size required by the network model, and cropping key regions of interest in the image based on prior knowledge or automatic detection algorithms. The preprocessed image data is then input into the pre-trained convolutional neural network for forward propagation computation. High-dimensional feature tensors are extracted from the output of specific intermediate layers in the network architecture, such as after the last convolutional layer and before the fully connected layer. Global average pooling is performed on the extracted high-dimensional feature tensors to compress them along the spatial dimension, ultimately generating a fixed-length, one-dimensional image feature vector. This image feature vector is compared with waveform feature vectors and text feature vectors obtained from the same neonate at the same time point for correlation analysis at the feature level. If significant contradictions or conflicts are found between different modal features, a feature review process is triggered, which re-evaluates the quality of the original image files and the accuracy of their annotation information.
[0094] In the specific implementation, the step of extracting image feature vectors is performed on a single neonatal brain ultrasound image file. A ResNet-50 convolutional neural network, pre-trained on the large medical image dataset ImageNet, is used as the backbone network for feature extraction. The input neonatal brain ultrasound image file undergoes standardization preprocessing, including normalizing the pixel grayscale values to the range [0,1], scaling the image size to 224 pixels by 224 pixels to meet the input requirements of the convolutional neural network, and automatically cropping the ventricle region of interest based on anatomical landmarks identified in the ultrasound image. In the specific implementation, the pre-processed image data is input into the pre-trained ResNet-50 convolutional neural network. During the forward propagation of the convolutional neural network, a high-dimensional feature tensor is extracted from the global average pooling layer after the last residual block in the ResNet-50 network architecture. The dimensions of the high-dimensional feature tensor are 7 x 7 x 2048. Data comparison shows that the high-dimensional feature tensor extracted from the last convolutional layer contains richer spatial structure information than the features extracted from the fully connected layer, resulting in an average accuracy increase of 3.5 percentage points in subsequent disease classification tasks. Global average pooling is performed on the extracted high-dimensional feature tensor, compressing the 7x7 feature map in each channel into a single scalar. This compresses the 2048-channel 7x7 feature tensor into a fixed-length 2048-dimensional image feature vector. The image feature vector generation process can be formally represented as follows:
[0095] ;
[0096] in: This represents the generated image feature vector. This represents the global average pooling function. This represents a high-dimensional feature tensor extracted from a convolutional neural network.
[0097] In some embodiments, region of interest (ROI) cropping is achieved by training an auxiliary lightweight region proposal network (RPN). The RPN takes the preprocessed image as input and outputs bounding box coordinates. Data comparison shows that cropping using the RPN is superior to fixed-region cropping. The generated image feature vector is correlated with the ECG waveform feature vector and diagnostic text feature vector obtained from the same newborn at the same time point at the feature level. The cosine similarity between the image feature vector and the waveform feature vector is calculated, as well as the cosine similarity between the image feature vector and the embedded vectors describing the brain in the text feature vector. If both similarity values are lower than a preset consistency threshold, a feature review process is triggered. In specific implementations, the feature review process includes reloading the original image file, checking the image file's acquisition parameters and quality indicators, and re-evaluating the image file's annotations using a second opinion. It is understood that correlation verification, as a kind of integrity check, aims to discover multimodal data inconsistencies caused by image acquisition artifacts, annotation errors, or data asynchrony. Optionally, the consistency threshold is determined by analyzing the multimodal feature similarity distribution of normal cases in historical data, and is set to the 5th percentile of the distribution. It is understandable that the feature review process can correct some feature biases caused by data quality issues. Data comparison shows that after introducing the review process, the number of error risk assessment alarms caused by image quality issues decreased by 25%. Optionally, for image files confirmed to be of poor quality after review, the system will ignore the corresponding image feature vectors and mainly rely on other modal data for risk assessment, while recording the event in the system log.
[0098] In one embodiment of the present invention, see [reference] Figure 3On the constructed risk assessment graph, starting with entity nodes whose multidimensional state vectors are determined to be anomalous, a graph traversal of a specified depth is performed. During the traversal, all visited entity nodes and the relational edges connecting these nodes are collected, thus forming a candidate risk subgraph. A graph structure metric is calculated on the candidate risk subgraph, calculating the shortest path length between all pairs of entity nodes and taking the arithmetic mean of these lengths as the average path length of the subgraph. The number of directly connected neighbor nodes of each entity node in the candidate risk subgraph is counted to obtain the degree of each node, and a histogram of node degree distribution is plotted based on the degrees of all nodes. For each entity node in the candidate risk subgraph, the ratio of the actual number of relational edges between its neighbor nodes to the theoretically maximum number of relational edges is calculated, and the average of this ratio is taken for all entity nodes to obtain the clustering coefficient of the subgraph. The calculated average path length, node degree distribution, and clustering coefficient, along with the total number of entity nodes and the total number of relational edges in the subgraph, are packaged into a graph structure metric data package. This graph structure metric data package is matched against a set of predefined graph pattern templates, including star-shaped diffusion patterns, chain-like transmission patterns, and mesh-like interweaving patterns. If the metric data of a candidate risk subgraph successfully matches a star-shaped diffusion pattern, the candidate risk subgraph is marked as a stable risk cluster. If the metric data of a candidate risk subgraph successfully matches a chain-like transmission pattern, the candidate risk subgraph is marked as an evolutionary risk cluster. If a candidate risk subgraph cannot match any predefined graph pattern template, but the outlier value of the multidimensional state vector of one of its entity nodes exceeds an independently set threshold, the entity node is marked as an isolated risk signal. All identified stable risk clusters, evolutionary risk clusters, and isolated risk signals are added to a continuously maintained risk tracking list, and each entry in the list is appended with its corresponding graph structure identifier and anomaly intensity summary information.
[0099] In practical implementation, when initiating the continuous risk tracking process, the risk assessment graph includes a multi-dimensional state vector displaying an entity node with an outlier: "Blood oxygen saturation consistently below 90%". Starting from this entity node, a graph traversal of depth 3 is performed. During the traversal, all visited entity nodes, including "increased heart rate", "respiratory distress", and "metabolic acidosis", along with the edges connecting these nodes, are collected, forming a candidate risk subgraph containing 4 entity nodes and 5 edges. A graph structure metric is calculated on the candidate risk subgraph, determining the shortest path length between all entity nodes. The shortest path length from entity node "Blood oxygen saturation consistently below 90%" to entity node "metabolic acidosis" is 2, and the shortest path length from entity node "increased heart rate" to entity node "respiratory distress" is 1. The arithmetic mean of these shortest path lengths is calculated, resulting in an average path length of 1.4 for the candidate risk subgraph. The number of directly connected neighbor nodes for each entity node in the candidate risk subgraph was counted. The node degree of the entity node "blood oxygen saturation consistently below 90%" was 3, the node degree of the entity node "increased heart rate" was 2, the node degree of the entity node "respiratory distress" was 2, and the node degree of the entity node "metabolic acidosis" was 1. A histogram of node degree distribution showed that there was 1 node with a degree of 3, 2 nodes with a degree of 2, and 1 node with a degree of 1. For each entity node in the candidate risk subgraph, the ratio of the actual number of relational edges between its neighbor nodes to the maximum possible number of relational edges was calculated. The entity node "blood oxygen saturation consistently below 90%" had 1 actual relational edge between its neighbor nodes, and the maximum possible number of relational edges was 3, resulting in a ratio of 0.333. The average of the ratios for all entity nodes yielded a clustering coefficient of 0.208 for the candidate risk subgraph. The calculated average path length (1.4), node degree distribution histogram, and clustering coefficient (0.208) are packaged together with the number of entity nodes (4) and relation edges (5) of the candidate risk subgraphs to form a graph structure metric data package for the candidate risk subgraphs. Data comparison shows that in a dataset containing 100 different risk subgraphs, the average time to calculate the above metrics is less than 50 milliseconds, meeting the real-time requirements.
[0100] In some embodiments, the depth parameter of the graph traversal can be configured according to the urgency of the risk assessment. For high-risk alarm starting point entity nodes, the depth parameter is set to 4 to cover a wider range of associations. Data comparison shows that depth 4 can find 15% more associated entity nodes than depth 3. The graph structure metric data package of the candidate risk subgraph is matched with a predefined graph pattern template. The predefined star-shaped diffusion pattern template requires a clustering coefficient of less than 0.3 and a central node with a node degree significantly higher than other nodes. The chain-like transmission pattern template requires an average path length to node number ratio greater than 0.8 and a node degree distribution concentrated at 2. The candidate risk subgraph has a clustering coefficient of 0.208, which is less than 0.3, and a central node with a node degree of 3, "blood oxygen saturation consistently below 90%", which matches the star-shaped diffusion pattern successfully. Therefore, the candidate risk subgraph is marked as a stable risk cluster. In a specific implementation, if another candidate risk subgraph has an average path length of 3.2, a node number of 4, and a node degree of 2 for all nodes, it matches the chain-like transmission pattern successfully and is marked as an evolutionary risk cluster. Optionally, the graph pattern template is defined by inductively analyzing the risk assessment graph structure of historical confirmed cases, and data comparison shows that the inductively defined template is used for matching. If a candidate risk subgraph contains the entity node "hyperbilirubinemia" and its multidimensional state vector outlier exceeds the independent threshold of 18.5 mg / dL, but the graph structure metric of the candidate risk subgraph cannot match any predefined template, then the entity node "hyperbilirubinemia" is marked as an isolated risk signal.
[0101] It is understandable that graph structure metrics provide a quantitative basis for pattern matching, such as the clustering coefficient. The calculation formula is:
[0102] ;
[0103] in: Represents the clustering coefficient. The number of entity nodes in the candidate risk subgraph. Representative entity node The actual number of relational edges between neighboring nodes. Representative entity node The node degree. In some embodiments, the anomaly strength summary is generated by extracting the top-3 anomaly dimension names and their values from the multidimensional state vectors of all anomaly entity nodes within the risk cluster. Optionally, the risk tracking list is implemented using a priority queue data structure, allowing for rapid adjustment of the entry order based on dynamic changes in the anomaly strength summary.
[0104] In one embodiment of the present invention, for each stable risk cluster in the risk tracking list, the historical multidimensional state vector sequence of all entity nodes contained therein is extracted. Time series analysis is performed on these historical multidimensional state vector sequences to fit their trend curves, and the state prediction vector for the next assessment cycle is extrapolated based on this trend. The obtained state prediction vector is compared with preset warning thresholds for various clinical indicators to generate a list of future risk sites containing time points and entity nodes that may exceed the thresholds in the future. Based on the temporal and spatial density distribution of the future risk site list, a specific deployment plan for the active monitoring layer is generated. This plan clearly specifies the physiological locations requiring enhanced monitoring, the recommended monitoring frequency, and the protocol used for data transmission. For each evolving risk cluster in the risk tracking list, its spectral structure identifier is analyzed to identify the risk transmission path with the fastest growth in relation edge strength. The propagation process of abnormal signals along this risk transmission path is simulated, and the estimated time required for the signal to propagate from the starting node to the key physiological node is calculated. Based on the urgency of the time estimate, different levels of early warning triggering rules and personnel dispatch plans corresponding to the early warning levels are configured in the preparatory response layer. By integrating the deployment plan of the proactive monitoring layer with the early warning response rules of the preparatory response layer, an immediate response layer plan is formed, which includes specific operational instructions, clear triggering conditions, and designated execution entities. The above three layers of plans together constitute a hierarchical and complete risk management plan.
[0105] In practice, the risk tracking list includes an entry labeled as a stable risk cluster, named "Hypoglycemia-Cluster." This stable risk cluster contains entity nodes "blood glucose concentration," "heart rate variability," and "drowsy state." For the stable risk cluster "Hypoglycemia-Cluster," a historical multidimensional state vector sequence of the entity node "blood glucose concentration" over the past 6 hours is extracted. This historical multidimensional state vector sequence contains vector data at 36 time points, with each vector containing three dimensions: blood glucose value, rate of decline, and volatility. Time series analysis is performed on the blood glucose value dimension of the historical multidimensional state vector sequence, and a linear regression model is used to fit its changing trend. The state prediction vector for the next 2 hours is then extrapolated using the formula:
[0106] ;
[0107] in: This represents a state prediction vector at a future point in time. Represents a matrix of historical multidimensional state vector sequences. This represents the regression coefficient vector obtained by fitting using the least squares method. This represents the random error term. The blood glucose value dimension in the state prediction vector is compared with the preset neonatal blood glucose concentration warning threshold of 2.6 mmol / L to generate a future risk site list. This list includes three time points and risk site combinations predicted to result in blood glucose concentrations below the threshold at specific time points. Based on the spatiotemporal density of the future risk site list, a deployment scheme for the active monitoring layer is generated. The spatiotemporal density is obtained by calculating the number of risk sites and their spatial distribution dispersion per unit time. The deployment scheme specifies monitoring finger-prick blood glucose every 15 minutes, continuous monitoring of electrocardiograms, and transmission of vital sign data to the central server every second via the HL7 protocol. See Table 1 for the future risk site list and its spatiotemporal density calculation.
[0108] Table 1: List of Future Risk Sites and Spatiotemporal Density Table Time point (minutes) Risk sites Predicted blood glucose level (mmol / L) Preset threshold (mmol / L) 30 blood glucose concentration 2.7 2.6 60 blood glucose concentration 2.5 2.6 90 blood glucose concentration 2.4 2.6
[0109] Data comparison shows that the average absolute error in time between the list of future risk sites extrapolated by linear regression and the subsequent actual hypoglycemic events is 12 minutes, while the method based on simple threshold alarms cannot provide time prediction.
[0110] In practice, the risk tracking list also includes an entry labeled as an evolutionary risk cluster, identified as "Sepsis-Chain". Analysis of the graph structure of the "Sepsis-Chain" cluster reveals it to be a chain-like structure composed of four entity nodes: "increased body temperature", "increased inflammatory markers", "decreased blood pressure", and "prolonged capillary refill time". The risk transmission path with the fastest increasing relationship edge strength is identified as the edge from "increased inflammatory markers" to "decreased blood pressure", with the strength of this edge increasing from 0.5 to 0.9 in the past hour. Simulation of the abnormal signal propagation process along the risk transmission path is performed, setting the average propagation delay between adjacent entity nodes to 10 minutes based on historical data. The estimated time for the signal to travel from the initial node "increased body temperature" to the key physiological node "prolonged capillary refill time" is calculated to be 30 minutes. Based on the urgency of the estimated time, different levels of early warning response rules are configured in the prepared response layer. A red alert is triggered when the estimated time is less than or equal to 30 minutes. The response rules for a red alert include immediately notifying the intensive care team and preparing intravenous infusion lines and vasoactive drugs. An orange alert is triggered when the estimated time is between 31 and 60 minutes. The response rules for an orange alert include notifying the attending physician and increasing the frequency of vital sign monitoring.
[0111] See Figure 4This is a bar chart showing the correspondence between warning levels and trigger thresholds in neonatal disease risk assessment. It primarily displays the abnormal intensity trigger thresholds corresponding to different warning levels. This chart is a core configuration tool for the "preparatory response layer" in the risk management plan. Different warning levels correspond to differentiated response rules, and the trigger thresholds ensure the accuracy of the response. The gradient design of the thresholds enables tiered risk management, avoiding excessive resource waste or insufficient response. This type of visualization helps medical staff quickly identify the treatment standards corresponding to risk levels and is a key reference for the dynamic management of neonatal disease risks.
[0112] In one embodiment of the present invention, at the end of a preset risk assessment period, all clinical treatment records and equipment monitoring readings generated during the period from the implementation of the hierarchical risk management plan are collected. The clinical treatment records are analyzed to extract the recorded treatment measures, the entity nodes affected by the measures, and the recorded effect evaluations. Newly collected equipment monitoring readings are converted into updated waveform feature vectors or text feature vectors. The entity nodes affected by the treatment measures are located in the risk assessment map, and the recorded effect evaluations are quantified into an effect intensity coefficient. This effect intensity coefficient is used to weight and adjust the current multidimensional state vector of the affected entity node, reducing the values of abnormal feature dimensions in its vector. The algorithm iterates through all neighboring entity nodes directly connected to the target entity node via relational edges. Based on the type and strength of the connection edges, it propagates the effect intensity coefficient through attenuation. It identifies the types of relational edges, including strong causality, weak correlation, and time-series association, and pre-determines a propagation attenuation factor for each type. Strong causality corresponds to a larger attenuation factor, while weak correlation corresponds to a smaller attenuation factor. Starting from the original effect intensity coefficient of the target entity node, it multiplies it by the propagation attenuation factor corresponding to the relational edge type to obtain the attenuated effect intensity coefficient propagated to neighboring nodes. The attenuated coefficients are then used to fine-tune the associated dimensions in the neighboring node's state vector. The updated waveform feature vector or text feature vector is fused with the multidimensional state vector of the corresponding entity node after effect propagation adjustment to generate a new multidimensional state vector for that entity node, completing the reverse calibration of the graph node state and edge strength. After reverse calibration, based on the new state vector of the entity node, the anomaly intensity summary of all stable risk clusters and evolutionary risk clusters is recalculated. The recalculated anomaly intensity summary is compared with the value of the previous period to calculate its rate of change and direction of change. For risk clusters whose anomaly intensity summaries weaken by more than a preset success threshold, their priority in the hierarchical risk management plan is reduced, the monitoring frequency in their corresponding active monitoring layer deployment plan is lowered, or their warning level in the preparatory response layer is lowered. For risk clusters whose anomaly intensity summaries strengthen or newly emerge, their priority is increased, their corresponding operational instructions are prioritized in the immediate response layer plan, and additional computing and storage resources are allocated to these risk clusters for deep graph analysis. Based on the adjusted priorities of all risk clusters, all pending measures in the hierarchical risk management plan are reordered to form a new measure execution sequence. Based on the new measure execution sequence and the resource allocation results for each risk cluster, resource scheduling instructions and task distribution lists for the next assessment cycle are generated.
[0113] In practical implementation, taking the operation at the end of a risk assessment cycle for a case diagnosed with neonatal respiratory distress syndrome as an example, at the end of the preset 4-hour risk assessment cycle, all clinical treatment records and equipment monitoring readings generated during the risk assessment cycle from the implementation of the tiered risk management plan are collected. The clinical treatment record includes a medical order record of "providing continuous positive airway pressure support," and the equipment monitoring readings include tidal volume and respiratory rate waveform sequences exported from the ventilator and values recorded by the blood oxygen saturation monitor. The clinical treatment record is analyzed, and the treatment measure is extracted as "continuous positive airway pressure support," the target entity node is "lung compliance," and the recorded effect evaluation is "relief of respiratory distress symptoms." The tidal volume waveform readings monitored by the ventilator are transformed into an updated waveform feature vector through feature extraction. The target entity node "lung compliance" is located in the risk assessment atlas, and the recorded effect evaluation "relief of respiratory distress symptoms" is quantified into an effect intensity coefficient of 0.7 according to a predefined semantic-numerical mapping table. The current multidimensional state vector of the target entity node "Lung Compliance" is weighted and adjusted using an effect strength coefficient of 0.7. The multidimensional state vector includes a dimension of "Lung Compliance Deterioration" with a value of 0.9, and the adjustment is performed using the following formula:
[0114] ;
[0115] in: This represents a new multidimensional state vector for the entity node. The multidimensional state vector representing the current state of the entity node. Represents the intensity coefficient of the effect. Represents a with A binary mask vector of the same dimension is used to identify which dimensions belong to the anomalous feature dimensions to be reduced, with the sign... This indicates element-wise multiplication, and after calculation, the value of the "low lung compliance" dimension decreased to 0.27. Data comparison shows that using a weighted adjustment method to update node states reflects the gradual effects of interventions more smoothly than directly replacing them with observed values, reducing drastic fluctuations in state values by 35% in the simulated data.
[0116] The algorithm iterates through the neighboring entity nodes of the target entity node "Lung Compliance" directly connected via relational edges. These neighboring entity nodes include "Blood Oxygen Saturation" and "Work of Respiration." Based on the type and strength of the relational edges, the effect intensity coefficient is attenuated and propagated. The relational edge connecting "Lung Compliance" and "Blood Oxygen Saturation" is identified as strongly causal, and a propagation attenuation factor of 0.8 is preset for this type. The relational edge connecting "Lung Compliance" and "Work of Respiration" is identified as weakly correlated, and a smaller propagation attenuation factor of 0.3 is preset for this type. Starting from the original effect intensity coefficient of 0.7 for the target entity node "Lung Compliance," the original effect intensity coefficient of 0.7 is multiplied by the propagation attenuation factor corresponding to the relational edge type, resulting in an attenuated effect intensity coefficient of 0.56 for propagation to the neighboring entity node "Blood Oxygen Saturation" and 0.21 for propagation to the neighboring entity node "Work of Respiration." The "Insufficient Saturation" dimension in the multidimensional state vector of the neighboring entity node's "Blood Oxygen Saturation" is numerically adjusted using a decayed effect strength coefficient of 0.56, which is associated with the abnormal feature of the "Lung Compliance" of the target entity node. The "Increased Work" dimension in the multidimensional state vector of the neighboring entity node's "Work of Respiration" is adjusted using a decayed effect strength coefficient of 0.21. If the neighboring entity node further connects to other nodes, the decayed effect strength coefficient propagated to the neighboring entity node is used as the new starting strength. Multi-level decay propagation is then performed based on the new relation edge type and the corresponding propagation decay factor until the effect strength coefficient falls below the propagation termination threshold of 0.1.
[0117] See Figure 5 This chart compares the number of measures and resource allocation at different levels in a neonatal disease risk management plan, primarily showcasing the resource investment and scale of measures at the active monitoring, preparedness response, and immediate treatment levels. This chart serves as a core basis for optimizing resources in the risk management plan. The preparedness response level has the highest resource allocation (approximately 45%), matching its core role of "risk transmission and interception," ensuring sufficient resources for key aspects. The active monitoring level has the most measures, corresponding to its need for "multi-dimensional, high-frequency monitoring," ensuring comprehensive capture of risk signals. This type of visualization helps balance the allocation of resources and measures at each level, avoiding resource waste or insufficient investment in key aspects, and is a guarantee of the efficiency of dynamic neonatal disease risk management.
[0118] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0119] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for assessing neonatal disease risk based on big data analysis, characterized in that, Includes the following steps: Initial multimodal health records of newborns are obtained from multiple heterogeneous medical data sources. These initial multimodal health records include image files, time-series vital sign streams, structured laboratory reports, and unstructured medical texts. The initial multimodal health records are fused and analyzed to construct an event-driven risk assessment graph, which consists of entity nodes, relation edges, and multidimensional state vectors attached to the entity nodes. Initiate a continuous risk tracking process, using the aforementioned risk assessment map to perform hierarchical risk scanning, and identify stable risk clusters, evolving risk clusters, and isolated risk signals; Based on the spatial distribution and temporal evolution patterns of the stable risk clusters and the evolving risk clusters, a hierarchical risk management plan is derived, which includes an active monitoring layer, a preparatory response layer, and an immediate handling layer. At the end of each cycle of the risk tracking process, the latest clinical treatment feedback and monitoring readings are integrated to perform reverse calibration on the entity node status and relation edge strength in the risk assessment map. Based on the reverse-calibrated risk assessment map, the implementation sequence of measures and resource allocation in the hierarchical risk management plan are dynamically reorganized.
2. The neonatal disease risk assessment method based on big data analysis according to claim 1, characterized in that, The specific implementation of fusing and analyzing the initial multimodal health records to construct an event-driven risk assessment map includes: The time-series vital sign stream is segmented, and the continuous waveform data is divided into multiple physiological event segments based on physiological event markers, and a waveform feature vector is generated for each physiological event segment. Named entity recognition technology is used to extract clinical entities and medical events from unstructured medical texts. The extracted clinical entities and medical events are then aligned and associated with fields in structured test reports to form text feature vectors. The image file is input into a pre-trained deep learning network to extract image feature vectors; Create a unified event index, and concatenate and reduce the dimensionality of waveform feature vectors, text feature vectors, and image feature vectors that occur within the same time window to generate a fused event vector for the time window. The clinical entities, medical events, and physiological event fragments are used as entity nodes, and the causal relationships, temporal relationships, and statistical symbiotic relationships between entity nodes are used as relation edges. The fused event vector is used as the initial multidimensional state vector of the core entity node within the corresponding time window to complete the construction of the event-driven risk assessment map.
3. The neonatal disease risk assessment method based on big data analysis according to claim 2, characterized in that, The specific implementation of initiating a continuous risk tracking process and using the risk assessment map for hierarchical risk scanning includes: On the risk assessment graph, a graph traversal of a specified depth is performed, starting from an entity node with an abnormal multidimensional state vector. During the traversal, all visited entity nodes and their connected relationship edges are collected to form a candidate risk subgraph; The candidate risk subgraph is subjected to graph structure metric calculation to obtain the clustering coefficient, average path length and node degree distribution of the candidate risk subgraph; The graph structure metric of the candidate risk subgraph is matched with a predefined graph pattern template, which includes star diffusion pattern, chain transmission pattern and network interweaving pattern. If a candidate risk subgraph successfully matches a star-shaped diffusion pattern, the candidate risk subgraph is marked as a stable risk cluster; If a candidate risk subgraph matches a chain-like transmission pattern, the candidate risk subgraph is marked as an evolutionary risk cluster. If a candidate risk subgraph cannot match any predefined graph pattern template, but the outlier value of the multidimensional state vector of the entity node it contains exceeds the independent threshold, then the entity node is marked as an isolated risk signal. All identified stable risk clusters, evolutionary risk clusters, and isolated risk signals are added to the risk tracking list, and each entry is appended with its corresponding spectral structure identifier and anomaly intensity summary.
4. The neonatal disease risk assessment method based on big data analysis according to claim 3, characterized in that, The specific steps for deriving a hierarchical risk management plan based on the spatial distribution and temporal evolution patterns of the stable and evolving risk clusters include: For each stable risk cluster in the risk tracking list, extract the historical multidimensional state vector sequence of all entity nodes contained therein; Time series analysis is performed on the historical multidimensional state vector sequence to fit its changing trend, and the state prediction vector for the next period is extrapolated. Compare the state prediction vector with the preset warning thresholds of each clinical indicator to generate a list of future risk sites; Based on the spatiotemporal density of the future risk site list, a deployment scheme for the active monitoring layer is generated, which specifies the monitoring points, monitoring frequency, and data transmission protocol. For each evolutionary risk cluster in the risk tracking list, analyze its graph structure identifier and identify the risk transmission path with the fastest growth in relation edge strength; Simulate the propagation process of abnormal signals along the risk transmission path and calculate the estimated time for the signal to reach key physiological nodes; Based on the estimated urgency of the time, different levels of early warning response rules and personnel dispatch plans are configured in the prepared response layer; By integrating the deployment scheme of the active monitoring layer with the early warning response rules of the preparatory response layer, an immediate handling layer plan is formed, which includes specific operation instructions, triggering conditions and execution entities, together constituting a hierarchical risk management plan.
5. The neonatal disease risk assessment method based on big data analysis according to claim 4, characterized in that, The specific implementation of integrating the latest clinical treatment feedback and monitoring readings at the end of each cycle of the risk tracking process to perform reverse calibration on the entity node status and relation edge strength in the risk assessment map includes: At the end of the preset risk assessment period, collect all clinical treatment records and equipment monitoring readings generated during the implementation of the hierarchical risk management plan within the risk assessment period; Analyze clinical treatment records to extract treatment measures, target entity nodes, and record effect evaluations; Convert equipment monitoring readings into updated waveform feature vectors or text feature vectors; Locate the target entity node in the risk assessment map and quantify the recorded effect evaluation into an effect intensity coefficient; The effect intensity coefficient is used to weight and adjust the current multidimensional state vector of the target entity node to reduce the value of the abnormal feature dimension. Traverse the neighboring entity nodes directly connected to the target entity node through relational edges, and propagate the effect strength coefficient attenuation according to the type and strength of the relational edges, and fine-tune the multidimensional state vector of the neighboring entity nodes accordingly. The updated waveform feature vector or text feature vector is fused with the multidimensional state vector of the corresponding entity node after effect propagation adjustment to generate a new multidimensional state vector of the corresponding entity node, thereby completing the reverse calibration of node state and edge strength in the risk assessment map.
6. The neonatal disease risk assessment method based on big data analysis according to claim 5, characterized in that, The specific implementation of dynamically reorganizing the action sequence and resource allocation in the hierarchical risk management plan based on the reverse-calibrated risk assessment map includes: After the reverse calibration is completed, based on the new multidimensional state vector of the entity nodes, the anomaly intensity summary of all stable risk clusters and evolutionary risk clusters is recalculated. Compare the recalculated anomaly intensity summary with the values from the previous period, and calculate its rate of change and direction of change; For risk clusters whose abnormal intensity summary weakens beyond the success threshold, reduce their priority in the hierarchical risk management plan, lower the monitoring frequency of their corresponding active monitoring layer deployment plan, or lower the early warning level of the preparatory response layer. For abnormal intensity summary enhancements or newly emerging risk clusters, their priority is increased, their corresponding operation instructions are placed in advance in the immediate handling layer plan, and additional computing and storage resources are allocated to the risk clusters for deep map analysis. Based on the adjusted priorities, all measures to be implemented in the tiered risk management plans are reordered to form a new sequence of measures to be implemented. Based on the new measures execution sequence and the resource allocation results of each risk cluster, the resource scheduling instructions and task distribution list for the next cycle are generated.
7. The neonatal disease risk assessment method based on big data analysis according to claim 2, characterized in that, The specific steps for inputting the image file into the pre-trained deep learning network to extract the image feature vector include: A convolutional neural network pre-trained on a large medical image dataset is used as the backbone network for feature extraction; The input neonatal image files are subjected to standardized preprocessing, including grayscale normalization, size scaling, and region of interest cropping. The preprocessed image data is input into the pre-trained convolutional neural network, and high-dimensional feature tensors are extracted from specific intermediate layers of the network during the forward propagation process. The extracted high-dimensional feature tensor is subjected to global average pooling to compress it into a fixed-length image feature vector. The image feature vector is compared with the waveform feature vector and text feature vector obtained from the same newborn at the same time point at the feature level. If a significant contradiction is found, the feature review process is triggered to re-evaluate the quality and annotation of the image file.
8. The neonatal disease risk assessment method based on big data analysis according to claim 3, characterized in that, The specific steps for calculating the graph structure metric of the candidate risk subgraph include: Calculate the shortest path length between all entity nodes in the candidate risk subgraph, and take the arithmetic mean of these shortest path lengths as the average path length of the candidate risk subgraph. The number of directly connected neighbor nodes of each entity node in the candidate risk subgraph is counted to obtain the node degree, and a histogram of the node degree distribution is plotted. For each entity node in the candidate risk subgraph, calculate the ratio of the actual number of relational edges between its neighboring nodes to the maximum possible number of relational edges, and then take the average of the ratios for all entity nodes to obtain the clustering coefficient of the candidate risk subgraph. The calculated average path length, node degree distribution, and clustering coefficient are packaged together with the number of entity nodes and relation edges of the candidate risk subgraph to form a graph structure metric data package for the candidate risk subgraph, which is used for subsequent graph pattern matching.
9. The neonatal disease risk assessment method based on big data analysis according to claim 5, characterized in that, The specific rules for attenuating and propagating the effect intensity coefficient include: Identify the type of relationship edge between the connected entity node and its neighboring entity nodes, wherein the relationship edge type includes strong causality, weak correlation, and temporal accompaniment; A propagation attenuation factor is preset for each type of relation edge. Relation edges with strong causality correspond to a larger propagation attenuation factor, while relation edges with weak correlation correspond to a smaller propagation attenuation factor. Starting from the entity node of the target, obtain its original effect intensity coefficient; Multiply the original effect strength coefficient by the propagation attenuation factor corresponding to the relationship edge type to obtain the attenuated effect strength coefficient propagated to the neighboring entity node. The attenuation effect strength coefficient is used to numerically adjust the dimension in the multidimensional state vector of the neighboring entity node that is associated with the abnormal features of the target entity node. If a neighboring entity node further connects to other nodes, the attenuated effect strength coefficient after propagation to the neighboring entity node is used as the new starting strength. Based on the new relationship edge type and the corresponding propagation attenuation factor, multi-level attenuation propagation is carried out until the effect strength coefficient is lower than the propagation termination threshold.
10. The neonatal disease risk assessment method based on big data analysis according to claim 2, characterized in that, The specific implementation of using the clinical entities, medical events, and physiological event fragments as entity nodes, and the causal relationships, temporal relationships, and statistical symbiotic relationships between entity nodes as relation edges includes: Clinical entities and medical events extracted from the unstructured diagnostic text, as well as physiological event fragments segmented from the time-series vital signs stream, are instantiated as unique entity nodes in the risk assessment graph, respectively. Evidence of pairwise relationships between entity nodes is extracted from field associations in structured lab reports, semantic co-occurrence in unstructured medical texts, and timestamp alignment in multimodal data. Based on the causal rules of the medical knowledge base, causal relationship edges are created for paired entity nodes with causal logic, and the causal relationship edges are assigned initial weights that reflect their causal determinism. Based on the timestamp order in the unified event index, time sequence edges are created for pairs of entity nodes that are related in the time series, and the time sequence edges are assigned an initial weight that reflects their temporal proximity. Based on big data statistical analysis of historical multimodal health records, statistical co-occurrence relationship edges are created for pairs of entity nodes that frequently co-occur statistically but have no clear causal or temporal constraints, and the statistical co-occurrence relationship edges are assigned initial weights that reflect their co-occurrence frequency. The created causal relationship edges, temporal relationship edges, and statistical symbiotic relationship edges are associated with the corresponding source entity nodes and target entity nodes to construct a risk assessment graph network structure that includes heterogeneous entities and multiple relationships.
Citation Information
Cited By
Infection risk assessment method and system based on nursing
CN121938640A
A nursing-based infection risk assessment method and system
CN121938640B