Big data driven patient health monitoring method
Through a big data-driven health monitoring method, using improved self-organized mapping algorithms and intelligent analysis models, the shortcomings of existing systems in identifying rare health parameters and providing personalized health management suggestions are solved, and efficient health risk prediction and intervention are achieved in patients with chronic diseases or patients with rare diseases.
Patent Information
- Application Number
- CN202510145868.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing health monitoring system has many problems in data collection, integration and analysis, and it is difficult to identify rare health parameters and provide personalized health management advice, especially for patients with chronic diseases or patients with rare diseases, with limited ability to predict and intervene health risks.
Through a big data-driven patient health monitoring method, multi-source health data is integrated, and complex data patterns are mined using improved self-organizing mapping algorithms, and personalized health risk assessment reports are generated in real time with intelligent analysis and prediction models.
It realizes high sensitivity identification of rare health parameters, improves the accuracy and practicality of health risk prediction, and provides accurate and dynamic health management services.
Smart Images

Figure CN120148815A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of health monitoring, and particularly to a big data-driven patient health monitoring method. Background Art
[0002] With the progress of medical health technology, health monitoring has gradually shifted from traditional regular physical examinations to real-time and continuous dynamic monitoring. However, there are still many problems in the data collection, integration, and analysis of existing health monitoring systems. Traditional methods have insufficient ability to identify rare health parameters (such as arrhythmia and premature beats), and potential health risks are easily missed.
[0003] In addition, the existing systems have limited capabilities in personalized health assessment and are difficult to provide targeted health management suggestions based on patients' historical health data and real-time monitoring results. Especially for patients with chronic diseases or rare diseases, the current monitoring systems lack sufficient accuracy and flexibility to achieve accurate health risk prediction and intervention. Summary of the Invention
[0004] The present invention provides a big data-driven patient health monitoring method. By integrating multi-source health data, using an improved self-organizing mapping algorithm to mine complex data patterns, and combining intelligent analysis and prediction models, a personalized health risk assessment report is generated in real time to provide accurate and dynamic health management services for patients.
[0005] The big data-driven patient health monitoring method includes the following steps:
[0006] S1, Multi-source health data collection and integration: Collect patients' health data through medical devices and electronic medical record systems, including conventional health parameters and rare health parameters;
[0007] S2, Unsupervised data pattern mining: Based on an improved self-organizing mapping algorithm, introduce a neighborhood dynamic adjustment mechanism to update the neighborhood function, mine patterns from health data, and identify potential abnormal signal patterns;
[0008] S3, Automated data annotation: Combine known health data templates and the results of unsupervised data pattern mining, and generate annotated data through rule-based reasoning and fuzzy matching techniques to achieve automatic annotation of rare health parameters;
[0009] S4, Personalized health risk assessment and feedback: According to the annotated data, combined with the patient's historical health data, use a big data analysis prediction model to generate a personalized health risk assessment report.
[0010] Optionally, the S1 specifically includes:
[0011] S11, Medical device data collection: Using an electrocardiograph, pulse oximeter, and sphygmomanometer to collect patients' routine health parameters in real time, including heart rate, blood oxygen, and blood pressure;
[0012] Using an electrocardiograph to collect electrocardiogram data over a long period of time, and based on a feature extraction algorithm, identifying rare health parameters, including arrhythmia, premature beats, or atrial fibrillation.
[0013] Optionally, the specific steps of S2 include:
[0014] S21, Health data preprocessing: Normalize the collected health data set X = {x 1 , x 2 ,..., x n} to ensure that the data ranges of features in different dimensions are consistent. Use the principal component analysis method to reduce the dimension of high-dimensional data and extract the principal feature subset X'. The principal feature subset X' includes several data points x'.
[0015] S22, Construction of an improved self-organizing mapping algorithm: Initialize the self-organizing mapping network, define the two-dimensional grid structure of neurons, including multiple network nodes. Each network node is represented by a weight vector W j = {w j1 , w j2 ,..., w jm}, where W j represents the weight vector of the jth network node, m represents the number of features, which is the same as the number of columns of X', and introduce a neighborhood dynamic adjustment mechanism to update the neighborhood function h ij (t);
[0016] S23, Improved dynamic learning rate during training: In each iteration, map the input data point X' to the network node BMU (Best Matching Unit) closest to it. The weight vector of the network node BMU satisfies: BMU = argmin j ‖x' - W j ‖, where ‖x' - W j ‖ represents the Euclidean distance between the feature data point and the weight of the network node. Update the weights of the network node BMU and the nodes within its neighborhood;
[0017] S24, Abnormal pattern mining: Through the trained self-organizing mapping network, map the health data points to a two-dimensional topological grid and calculate the activation frequency f j of each node: where f j is the activation frequency of network node j, N j is the number of data points mapped to network node j, N is the total number of data points, and identify the nodes with abnormal activation frequencies as abnormal pattern nodes. The judgment condition is: fj <μ f -k·σ f , where μ f and σ f are the mean and standard deviation of the activation frequencies of all nodes respectively, k is the adjustment coefficient for anomaly recognition, and its value is 2;
[0018] S25. Mark the input data points mapped to the anomaly pattern nodes as anomaly data points Output their feature vectors and corresponding time indices for subsequent automatic annotation and risk assessment.
[0019] Optionally, in S22, update the neighborhood function h ij (t) is expressed as:
[0020] where R i and R j are the positions of network nodes i and j, ‖R i -R j ‖ represents the Euclidean distance between network nodes, and σ(t) represents the neighborhood radius that decays with time t, where σ 0 is the initial radius, τ σ is the time decay constant of the radius, and the time decay constant of the radius is proportional to the number of iterations: t max is the maximum number of training iterations.
[0021] Optionally, in S23, update the weights of the network node BMU and the nodes within its neighborhood. The weight update rule is:
[0022] W j (t + 1) = W j (t) + η(t)·h ij (t)·.x′ - W j (t) / , where η(t) is the dynamically adjusted learning rate, which is defined as: η 0 is the initial learning rate, set to 0.1, τη is the time decay constant of the learning rate, and the time decay constant of the learning rate: τ η = t max / 2, and the decay rate of the learning rate is synchronized with the middle stage of the training process.
[0023] Optionally, S3 specifically includes:
[0024] S31. Template construction: Based on historical data, create a healthy data template, including the feature ranges and descriptions of common and rare health parameters;
[0025] S32, Initial Matching: Calculate the similarity between the input abnormal data points and the template features to determine the matching degree with each healthy data template;
[0026] S33, Rule Reasoning: Apply rule logic based on the matching results to determine whether the data points conform to the healthy data template feature range and generate initial annotations;
[0027] S34, Fuzzy Matching: For abnormal data points that cannot be matched, use fuzzy logic to calculate the fuzzy matching membership degree to determine their association strength with the healthy data template;
[0028] S35, Annotation Output: For the results of matching and fuzzy matching, label the rare health parameter categories and output the annotation information, including the matching similarity score and the feature vector.
[0029] Optionally, in the initial matching of S32, the input abnormal data points are initially matched with the features in template T, and the similarity score S of each abnormal data point with template T p is calculated as follows: p :
[0030] where represents the q-th feature dimension of the abnormal data point , T pq represents the reference value of the healthy data template T p in the q-th feature dimension, σ pq represents the standard deviation of the q-th feature dimension, which is used for fuzzy matching, and m1 is the total number of feature dimensions.
[0031] Optionally, in the rule reasoning of S33, based on the initial similarity score S p , use rule-based reasoning technology to generate annotation data, and the rule form is: Label as where represents the th rule, τ s is the similarity threshold for judging the matching degree, is the healthy parameter category (arrhythmia, premature beat or atrial fibrillation) corresponding to template T p , and if the abnormal data point matches multiple templates, label based on the maximum similarity priority rule.
[0032] Optionally, in S33, for abnormal data points that cannot match the template, use fuzzy logic to generate annotation data for rare health parameters, and label the data points with a fuzzy matching membership degree higher than the threshold μ minThe data points are labeled with the corresponding rare health parameter categories, and the labeling results are output, including the labeled categories, similarity scores, membership degrees, and the corresponding health parameter feature vectors, and the threshold μ min The value ranges from 0.6 to 0.8, and the default value is 0.7, indicating that at least 70% of the matching degree is required to be considered a reasonable association.
[0033] Optionally, the S4 specifically includes:
[0034] S41, integration of labeled data and historical data: Integrate the rare health parameter categories generated by labeling and their corresponding feature vectors X label with the patient's historical health data H, where the historical health data H includes the patient's physiological index trends (heart rate change curve, blood pressure record), past disease records, and treatment plan data;
[0035] S42, feature extraction and data standardization: Extract features from the integrated data, including time series features (trend slope, fluctuation range), and pattern features (abnormal distribution);
[0036] S43, construction of a big data analysis prediction model: Use historical health data and labeled data to train a big data analysis prediction model, including a classification model based on support vector machines for classifying health risk categories, and the model input features F = {X label , H}, and output the patient's health risk level.
[0037] Advantages of the present invention:
[0038] In the present invention, by combining a known health data template and an unsupervised pattern mining technology, an efficient labeling and classification system is constructed. In the labeling link, rule reasoning and fuzzy matching technologies are used to achieve high-sensitivity recognition of rare health parameters (such as arrhythmia, premature beats, atrial fibrillation), overcoming the deficiencies of traditional methods in rare data recognition. By using a support vector machine classification model to grade and predict health risk categories, it can effectively capture the risk patterns hidden in complex health data, improving the accuracy and practicality of prediction.
[0039] In the present invention, the improved self-organizing mapping algorithm dynamically adjusts the neighborhood function, enabling the model to focus on the overall data distribution in the initial stage of training and gradually converge to the fine learning stage of local features. This mechanism ensures that potential patterns in healthy data (such as abnormal signal distributions) can be accurately identified, especially suitable for high-dimensional and diverse data scenarios. The dynamic neighborhood adjustment mechanism flexibly adjusts the neighborhood radius according to the number of iterations, enabling the algorithm to capture global patterns in the early stage and focus on the optimization of local details in the later stage. This method avoids the problems of premature convergence or pattern loss that may be caused by traditional fixed neighborhood strategies, improving the robustness and adaptability of healthy data pattern mining. After introducing the dynamic neighborhood adjustment mechanism, the improved self-organizing mapping algorithm can achieve efficient feature mapping and pattern clustering with limited computing resources. This method shortens the training time of the model by gradually reducing the neighborhood range, while ensuring that the model achieves the best balance between accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention;
[0042] Figure 2 It is a schematic diagram of automated data annotation according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The present invention will be described in detail below in conjunction with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; and the drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.
[0044] As Figure 1 - Figure 2 shown, the big data-driven patient health monitoring method includes the following steps:
[0045] S1, Multi-source health data collection and integration: Collect the patient's health data through medical devices and electronic medical record systems, including conventional health parameters (heart rate, blood oxygen, blood pressure) and rare health parameters (special heart rhythm types);
[0046] S2, Unsupervised Data Pattern Mining: Based on the improved self-organizing mapping algorithm, introduce a neighborhood dynamic adjustment mechanism to update the neighborhood function, mine patterns from health data, and identify potential abnormal signal patterns;
[0047] S3, Automated Data Annotation: Combine the known health data template and the results of unsupervised data pattern mining, generate annotation data through rule-based reasoning and fuzzy matching techniques, and achieve automatic annotation of rare health parameters;
[0048] S4, Personalized Health Risk Assessment and Feedback: According to the annotation data, combined with the patient's historical health data, use the big data analysis prediction model to generate a personalized health risk assessment report.
[0049] S1 specifically includes:
[0050] S11, Medical Device Data Acquisition: Use electrocardiographs, pulse oximeters, and sphygmomanometers to collect the patient's routine health parameters in real time, including heart rate, blood oxygen, and blood pressure;
[0051] Use an electrocardiograph to collect long-term electrocardiogram data, and based on the feature extraction algorithm (QRS complex detection), identify rare health parameters, including arrhythmia, premature beats, or atrial fibrillation.
[0052] S2 specifically includes:
[0053] S21, Health Data Preprocessing: Normalize the collected health data set X = {x 1 ,x 2 ,...,x n} to ensure that the data ranges of different dimensional features are consistent. Use the principal component analysis method to reduce the dimension of high-dimensional data and extract the principal feature subset X'. The principal feature subset X' includes several data points x'. The data point x' is a row vector in X', corresponding to a certain data point after dimension reduction. The data point x' is a single instance in the principal feature subset X' after normalization and dimension reduction, and is used as input for further processing by the self-organizing mapping algorithm. The principal component analysis method is expressed as:
[0054] X' = X · V, where X represents the normalized original data matrix, with a size of n × m, n is the number of data points, m is the number of features, X' is the data matrix after dimension reduction, and V represents the eigenvector matrix of the principal component analysis, representing the principal component direction;
[0055] S22, Construction of the Improved Self-Organizing Mapping Algorithm: Initialize the self-organizing mapping network, define the two-dimensional grid structure of neurons, including multiple network nodes, where each network node consists of a weight vector W j = {w j1 ,w j2 ,...,wjm} indicates that, where W j represents the weight vector of the j-th network node, m represents the number of features, which is the same as the number of columns of X′, and a neighborhood dynamic adjustment mechanism is introduced to update the neighborhood function h ij (t);
[0056] S23, improved dynamic learning rate during the training process: In each iteration, map the input data point X′ to the network node BMU that is closest to it. BMU is the best matching unit, and the weight vector of the network node BMU satisfies: BMU = argmin j ‖x′ - W j ‖, where ‖x′ - W j ‖ represents the Euclidean distance between the feature data point and the network node weight, and update the weights of the network node BMU and the nodes within its neighborhood;
[0057] S24, abnormal pattern mining: Through the trained self-organizing mapping network, map the healthy data points to a two-dimensional topological grid, and calculate the activation frequency f j : where f j is the activation frequency of network node j, N j is the number of data points mapped to network node j, N is the total number of data points, identify the nodes with abnormal activation frequencies as abnormal pattern nodes, and the judgment condition is: f j < μ f - k·σ f , where μ f and σ f are the mean and standard deviation of the activation frequencies of all nodes respectively, and k is the adjustment coefficient for abnormal identification, with a value of 2;
[0058] S25, mark the input data points mapped to the abnormal pattern nodes as abnormal data points Output their feature vectors and corresponding time indices for subsequent automatic annotation and risk assessment.
[0059] The specific logic is as follows:
[0060] 1. The self-organizing mapping network optimizes the distribution of network nodes through node weight initialization, dynamic learning rate, and neighborhood function, so that similar data points are clustered into the same network node.
[0061] 2. The node activation frequency is used to identify abnormal patterns, and abnormal network nodes are determined through statistical analysis.
[0062] 3. The corresponding data points output by the abnormal network nodes constitute potential abnormal signal patterns for subsequent annotation and analysis.
[0063] In S22, update the neighborhood function hij (t) is expressed as:
[0064] Wherein, R i and R j are the positions of network nodes i and j, and ‖R i -R j ‖ represents the Euclidean distance between network nodes, and σ(t) represents the neighborhood radius that decays with time t. Wherein, σ 0 is the initial radius, τ σ is the time decay constant of the radius, and the time decay constant of the radius is proportional to the number of iterations: t max is the maximum number of training iterations.
[0065] In S23, update the weights of the network node BMU and the nodes within its neighborhood, and the weight update rule is:
[0066] W j (t + 1) = W j (t) + η(t)·h ij (t)·.x′ - W j (t) / , wherein, η(t) is the dynamically adjusted learning rate, which is defined as: η 0 is the initial learning rate, set to 0.1, τη is the time decay constant of the learning rate, and the time decay constant of the learning rate: τ η = t max / 2, and the decay speed of the learning rate is synchronized with the middle stage of the training process.
[0067] S3 specifically includes:
[0068] S31, template construction: Based on historical data, create a healthy data template, including the feature ranges and descriptions of common and rare health parameters, and construct a known healthy data template T = {T 1 , T 2 ,..., T p}, wherein, T p represents the pth type of healthy data template, and each healthy data template T p includes the reference value range [L pq , U pq of common health parameters and the description formula of abnormal features;
[0069] S32, preliminary matching: Calculate the similarity between the input abnormal data points and the template features, and judge their matching degree with each healthy data template;
[0070] S33, Rule Reasoning: Apply rule logic based on the matching result to determine whether the data point conforms to the feature range of the healthy data template and generate an initial annotation;
[0071] S34, Fuzzy Matching: For abnormal data points that cannot be matched, calculate the fuzzy matching membership degree using fuzzy logic to determine the association strength with the healthy data template;
[0072] S35, Annotation Output: For the results of matching and fuzzy matching, annotate the rare health parameter categories and output the annotation information, including the matching similarity score and the feature vector.
[0073] In the preliminary matching of S32, the feature vector of the input abnormal data point is preliminarily matched with the features in template T, and the similarity score S of each abnormal data point with template T p is calculated as follows: p :
[0074] where, represents the q-th feature dimension of the abnormal data point , T pq represents the reference value of the healthy data template T p in the q-th feature dimension, σ pq represents the standard deviation of the q-th feature dimension, which is used for fuzzy matching, and m1 is the total number of feature dimensions.
[0075] In the rule reasoning of S33, based on the preliminary similarity score S p , annotation data is generated using rule-based reasoning technology, and the rule form is: Annotated as where, represents the th rule, τ s is the similarity threshold for judging the matching degree, is the healthy parameter category (arrhythmia, premature beat or atrial fibrillation) corresponding to template T p . If an abnormal data point matches multiple templates, it is annotated based on the maximum similarity priority rule.
[0076] In S33, for abnormal data points that cannot match the template, fuzzy logic is used to generate annotation data for rare health parameters, and the membership function of fuzzy matching is:
[0077] where Δ pq =U pq -L pq represents the span of the reference value range, and the membership degree represents the abnormal data point Degree of fuzzy matching with the template; data points with a fuzzy matching membership degree higher than the threshold μ min are labeled as the corresponding rare health parameter categories, and the labeling results are output, including the labeled category, similarity score, membership degree, and the corresponding health parameter feature vector, threshold μ min The value ranges from 0.6 to 0.8, and the default value is 0.7, indicating that at least 70% of the matching degree is required to be considered a reasonable association.
[0078] S4 specifically includes:
[0079] S41, integration of labeled data and historical data: Integrate the rare health parameter categories generated by labeling and their corresponding feature vectors X label with the patient's historical health data H, where the historical health data H includes the patient's physiological index trends (heart rate change curve, blood pressure record), past disease records, and treatment plan data;
[0080] S42, feature extraction and data standardization: Extract features from the integrated data, including time series features (trend slope, fluctuation range), pattern features (abnormal distribution);
[0081] S43, construction of big data analysis prediction model: Use historical health data and labeled data to train a big data analysis prediction model, including a classification model based on support vector machine for classifying health risk categories, and the model input features F = {X label , H}, and output the patient's health risk level.
[0082] The classification model based on support vector machine is specifically as follows:
[0083] 1. Input data preparation: Use the integrated feature set F = {X label , H} as the input of the support vector machine, where;
[0084] 2. Classification target definition: Define the health risk category y, and the categories are:
[0085] Low risk (y = 0): Good health condition, no significant abnormalities;
[0086] Medium risk (y = 1): There are certain potential health problems, but the risk is relatively low;
[0087] High risk (y = 2): There are significant health problems or high-risk signals.
[0088] 3. Training data construction: Combine the labeled data and historical data to form a training sample set {(F i , y i )}, where: F i is the feature vector of the i-th sample, y iIt is the corresponding health risk category label. Part of the dataset is used as the training set, and the remaining data is used as the test set to verify the classification effect.
[0089] 4. Support Vector Machine Model Training: Select the Gaussian kernel function suitable for health data. The kernel function determines how to calculate the similarity between data points in the feature space. The goal of the model is to find an optimal classification boundary that maximizes the margin between classifications while allowing a certain degree of error to accommodate noise or anomalies in the data. The influence of the error is controlled by the penalty parameter to ensure the balance of the model. During the training process, the optimization algorithm adjusts the model parameters step by step according to the input training data, including the direction and position of the classification boundary and the bias value used for classification. Through repeated optimization, an optimal classification model that can distinguish health risk categories is finally obtained. The Support Vector Machine model will use a small number of key data points (i.e., support vectors) to define the classification boundary. These support vectors play a key role in the classification ability of the model, while the influence of other data points is relatively small. After training, the model can classify new data.
[0090] 5. Classification Decision: For the newly input data point x, the classification decision is calculated by the following formula:
[0091] where α i is the Lagrange multiplier, representing the contribution of the support vector to the decision, and b is the bias, which determines the position of the classification boundary.
[0092] 6. Classification Output:
[0093] a) If f(x) ∈ [0, 1], then it is classified as low risk;
[0094] b) If f(x) ∈ [0, 1], then it is classified as medium risk;
[0095] c) If f(x) > 2, then it is classified as high risk.
[0096] The present invention covers any substitutions, modifications, equivalent methods, and solutions made within the essence and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention even without these detailed descriptions. Additionally, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.
[0097] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A big data driven patient health monitoring method, characterized in that: The following steps are involved: S1, multi-source health data collection and integration: collect patients’ health data through medical devices and electronic medical record systems, including routine health parameters and rare health parameters; S2, unsupervised data pattern mining: Based on the improved self-organizing map algorithm, a neighborhood dynamic adjustment mechanism is introduced to update the neighborhood function, perform pattern mining on health data, and identify potential abnormal signal patterns; S3, Automated Data Annotation: Combining known health data templates and unsupervised data pattern mining results, the annotation data is generated through rule-based reasoning and fuzzy matching technology to achieve automatic annotation of rare health parameters; S4, Personalized health risk assessment and feedback: Based on the labeled data and combined with the patient's historical health data, a personalized health risk assessment report is generated using a big data analysis and prediction model.
2. The big data driven patient health monitoring method according to claim 1, characterized in that: The S1 specifically includes: S11, medical equipment data collection: using electrocardiographs, pulse oximeters, and sphygmomanometers to collect patients’ routine health parameters in real time, including heart rate, blood oxygen, and blood pressure; An electrocardiograph is used to collect ECG data over a long period of time, and based on feature extraction algorithms, rare health parameters including arrhythmias, premature beats or atrial fibrillation are identified.
3. The big data driven patient health monitoring method according to claim 1, characterized in that: The S2 specifically includes: S21, health data preprocessing: collect the health data set X = {x1, x2, ..., x n }Perform normalization processing, use principal component analysis method to reduce the dimension of high-dimensional data, extract the main feature subset X′, and the main feature subset X′ includes several data points x′; S22, improved self-organizing map algorithm construction: Initialize the self-organizing map network, define the two-dimensional grid structure of neurons, including multiple network nodes, each of which is represented by a weight vector W j ={w j1 ,w j2 ,...,w jm } indicates that, where W j represents the weight vector of the jth network node, m represents the number of features, and introduces a neighborhood dynamic adjustment mechanism to update the neighborhood function h according to the number of iterations t. ij (t); S23, improved dynamic learning rate during training: In each iteration, the input data point X′ is mapped to the closest network node BMU, which is the best matching unit. The weight vector of the network node BMU satisfies: BMU = arg min j ‖x′-W j ‖, where ‖x′-W j ‖ represents the Euclidean distance between the feature data point and the network node weight, and updates the weights of the network node BMU and its neighborhood nodes; S24, abnormal pattern mining: through the trained self-organizing map network, the healthy data points are mapped into a two-dimensional topological grid and the activation frequency f of each node is calculated. j : Among them, f j is the activation frequency of network node j, N j is the number of data points mapped to network node j, N is the total number of data points, and nodes with abnormal activation frequency are identified as abnormal mode nodes. The judgment condition is: j <μ f -k·σ f , where μ f and σ f are the mean and standard deviation of the activation frequency of all nodes, respectively, and k is the adjustment coefficient for anomaly identification; S25, marking the input data points mapped to the abnormal pattern nodes as abnormal data points Output its feature vector and corresponding time index.
4. The big data driven patient health monitoring method according to claim 3, characterized in that: In S22, the neighborhood function h is updated. ij (t) is expressed as: Among them, R i and R j is the location of network nodes i and j, ‖R i -R j ‖ represents the Euclidean distance between network nodes, and σ(t) represents the neighborhood radius that decays with time t.
5. The big data driven patient health monitoring method according to claim 4, characterized in that: In S23, the weights of the network node BMU and the nodes in its neighborhood are updated, and the weight update rule is: W j (t+1)=W j (t)+η(t)·h ij (t)·.x′-W j (t) / , where η(t) is the dynamically adjusted learning rate, defined as: η0 is the initial learning rate, τη is the time decay constant of the learning rate.
6. The big data driven patient health monitoring method according to claim 5, characterized in that: The S3 specifically includes: S31, Template Construction: Create health data templates based on historical data, including characteristic ranges and descriptions of common and rare health parameters; S32, preliminary matching: calculating the similarity between the input abnormal data points and the template features to determine the matching degree between them and each healthy data template; S33, rule reasoning: applying rule logic according to the matching results, determining whether the data point meets the feature range of the health data template, and generating initial annotations; S34, fuzzy matching: for abnormal data points that cannot be matched, fuzzy logic is used to calculate the fuzzy matching membership to determine the strength of its association with the health data template; S35, labeling output: for the matching and fuzzy matching results, label the rare health parameter categories and output the labeling information, including the matching similarity score and feature vector.
7. The big data driven patient health monitoring method according to claim 6, characterized in that: In the preliminary matching of S32, the input abnormal data points are The feature vector of is preliminarily matched with the features in the template T, and each abnormal data point is calculated With template T p The similarity score S p : in, Indicates abnormal data points The qth feature dimension, T pq Represents the health data template T p The reference value in the qth feature dimension, σ pq represents the standard deviation of the qth feature dimension, which is used for fuzzy matching, and m1 is the total feature dimension.
8. The big data driven patient health monitoring method according to claim 7, characterized in that: In the rule reasoning of S33, according to the preliminary similarity score S p , the rule-based reasoning technology is used to generate the labeled data, and the rule form is: Marked as in, Indicates rules, τ s is the similarity threshold, used to judge the degree of matching. Template T p For the corresponding health parameter categories, if the abnormal data point matches multiple templates, it is labeled based on the maximum similarity priority rule.
9. The big data driven patient health monitoring method according to claim 8, characterized in that: In S33, for abnormal data points that cannot match the template, fuzzy logic is used to generate annotated data of rare health parameters, and the fuzzy matching membership is higher than the threshold μ min The data points are labeled as corresponding rare health parameter categories, and the labeling results are output, including the labeling category, similarity score, membership degree and the corresponding health parameter feature vector.
10. The big data driven patient health monitoring method according to claim 1, characterized in that: The S4 specifically includes: S41, integration of annotated data and historical data: the rare health parameter categories generated by annotation and their corresponding feature vectors X label Integrate with the patient's historical health data H, which includes the patient's physiological indicator trends, previous disease records and treatment plan data; S42, feature extraction and data standardization: extract features from the integrated data, including time series features and pattern features; S43, Construction of big data analysis prediction model: Use historical health data and annotated data to train big data analysis prediction model, including a classification model based on support vector machine, which is used to classify health risk categories. The model input feature F = {X label ,H}, output the patient’s health risk level.