Disease condition early warning method and system based on multi-modal data fusion, and medium
By integrating multimodal data fusion and adaptive machine learning, clinical physiological parameters, text medical records, and imaging data are combined to solve the problems of single data, poor dynamic adaptability, and coarse grading granularity in existing technologies, thus achieving efficient and accurate disease risk assessment and personalized treatment recommendations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING GENERAL HOSPITAL NANJING MILLITARY COMMAND P L A
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-17
AI Technical Summary
Existing medical risk warning technologies suffer from problems such as limited data dimensions, poor dynamic adaptability, low integration, and coarse grading granularity, resulting in inaccurate assessment results and weak decision-making guidance.
A multimodal data fusion approach is adopted to integrate clinical physiological parameters, text medical records, and imaging data. Through cleaning, normalization, and feature extraction, an adaptive machine learning algorithm is used to construct a dynamic risk assessment model and generate targeted decision recommendations.
It has improved the comprehensiveness and accuracy of disease risk assessment, enhanced the timeliness and dynamic adaptability of early warning, provided clear diagnostic and treatment references, and optimized the allocation of medical resources.
Smart Images

Figure CN121885154A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information management, and in particular to a method, system, and medium for disease early warning based on multimodal data fusion. Background Technology
[0002] This invention relates to the field of medical data processing and risk warning technology, specifically a hierarchical early warning method based on multimodal data fusion and dynamic risk assessment. This method is applicable to the early diagnosis, condition assessment and risk warning of critical illnesses such as acute pancreatitis, acute kidney injury and infectious diseases in the field of critical care medicine.
[0003] Currently, mainstream risk warning technologies on the market have certain limitations in medical data processing and risk warning. Most solutions conduct assessments based on single types of data or static data. In the medical field, a common practice is to over-rely on single biomarkers, such as judging a patient's condition solely based on blood glucose or blood lipid levels; or to conduct risk assessments only by referring to test results at fixed time points. Although some solutions introduce basic machine learning algorithms, such as logistic regression and single decision trees, these algorithms fail to fully integrate multi-source and multi-type dynamic data when processing complex medical data, making it difficult to comprehensively and accurately reflect the patient's condition and risk status.
[0004] The existing technology has the following main drawbacks: 1. Limited Data Dimensions and Insufficient Coverage: Single data types cannot comprehensively reflect risk-related factors. Relying solely on biomarkers for assessment fails to capture the impact of patients' lifestyle habits (such as diet, exercise, and sleep patterns) and real-time physiological states (such as heart rate variability and blood oxygen saturation) on their condition. These factors play crucial roles in the occurrence, development, and prognosis of the disease; ignoring them leads to inaccurate risk assessment results.
[0005] 2. Static assessment is lagging and lacks dynamic adaptability: Most treatment plans are based on fixed-point data or historical data modeling, which cannot track data change trends in real time. In medical settings, static examination results can only reflect a patient's physical condition at a specific moment and are difficult to present the dynamic progression of the patient's condition. For example, a patient's condition may change drastically in a short period of time, but static assessment-based plans cannot detect this change in time, thus easily missing the optimal intervention window and affecting treatment outcomes.
[0006] 3. Low data integration and low information utilization: Data from different sources and in different formats (such as clinical text medical records and numerical test results) are not effectively integrated, resulting in the loss of correlation information between data. Medical data comes from a wide range of sources, including electronic medical records, laboratory reports, and imaging data. These data are diverse in format and contain rich information. However, current technologies have failed to deeply integrate this multimodal data, making it impossible for models to uncover the synergistic value between different data, reducing information utilization, and consequently affecting the accuracy of risk warnings.
[0007] 4. The grading and early warning system is coarse-grained and lacks guidance for decision-making: Existing solutions mostly only output a binary "high / low risk" result, without classifying the risk according to its severity or rate of development, and thus cannot provide differentiated response strategies for different scenarios. For example, when facing patients with different risk levels, it is impossible to distinguish between "high-risk requiring urgent intervention" and "medium-risk requiring close monitoring," making it difficult for medical staff to develop personalized treatment plans based on specific circumstances, thereby reducing the scientific validity and effectiveness of decision-making. Summary of the Invention
[0008] Therefore, it is necessary to provide a disease early warning method, system, and medium based on multimodal data fusion to address the problems of single data dimension, poor dynamic adaptability, low integration degree, and coarse hierarchical granularity mentioned above.
[0009] A disease early warning method based on multimodal data fusion includes: Acquire multimodal data information, including the patient's clinical physiological parameters, text medical record data, and image data; The multimodal data information is cleaned, denoised, and normalized to obtain standardized data information; Feature extraction is performed on the standardized data information to obtain feature data information, and the feature data information is fused to obtain the fused potential risk factors; The fusion decision results are imported into the constructed dynamic risk assessment model to obtain and output risk assessment results, wherein the dynamic risk assessment model satisfies:
[0010] In the above formula, Based on the risk assessment results, This refers to the vector information corresponding to multimodal data. As a potential risk factor, For bias terms, The weights of the aforementioned potential risk factors, It is the Sigmoid activation function.
[0011] In one preferred embodiment, the multimodal data information is normalized, including: The multimodal data information is imported into a normalized data model, which satisfies the following:
[0012] in, For the imported multimodal data information, The maximum value of the imported data. The minimum value of the imported data. This is data that has undergone normalization.
[0013] In one preferred embodiment, the step of extracting features from the standardized data information to obtain feature data information, and fusing the feature data information to obtain fused potential risk factors, includes: The mean, variance, and peak value time-domain features and power spectral density frequency-domain features of clinical physiological parameters are extracted to obtain clinical physiological parameter feature vectors. The TF-IDF method is used to extract features from text data to obtain text feature vectors. Convolutional neural networks are used to extract features from image data to obtain image feature vectors. The fusion based on the feature data information satisfies:
[0014] in, To integrate decision-making results, For clinical physiological parameter feature vectors, The weights of the feature vector of clinical physiological parameters. For text feature vectors, For text feature vector weights, For image feature vectors, The weights are the image feature vectors. Import the aforementioned feature data into the risk knowledge base to obtain potential risk factors.
[0015] In one preferred embodiment, the acquisition of multimodal data information, including the patient's clinical physiological parameters, text medical record data, and image data, includes: The clinical physiological parameters are collected by establishing a communication connection with the electrocardiogram monitor, blood pressure monitor, and pulse oximeter via the HL7 protocol. The image data is acquired by interfacing with the PACS system via the DICOM protocol. The collected medical records are obtained from the HIS system through database queries or interface calls, and the unstructured text is converted into structured data using a named entity recognition algorithm based on natural language processing.
[0016] In one preferred embodiment, the denoising of the multimodal data information includes: The clinical physiological parameters are denoised using Kalman filtering or median filtering, the text medical record data is denoised using syntax checking and semantic analysis, and the image data is denoised using Gaussian filtering or mean filtering.
[0017] In one preferred embodiment, the step of extracting features from the standardized data information to obtain feature data information, and fusing the feature data information to obtain a fusion decision result, includes: Using the fused high-dimensional feature vectors as sample features and the disease diagnosis and risk level results confirmed by clinical physician consultation as sample labels, and taking into account the characteristics of high dimensionality and imbalanced sample categories of multimodal data, a dynamic risk assessment model is constructed using a machine learning algorithm that combines feature selection capability and classification accuracy.
[0018] In one preferred embodiment, the step of outputting risk assessment results based on the dynamic risk assessment model includes: The risk assessment results are classified into risk levels to obtain risk level early warning information; Based on the aforementioned risk level warning information, decision-making recommendations are generated.
[0019] In one preferred embodiment, the step of generating decision suggestion information based on the risk level warning information includes: If the risk level warning information is a low-risk warning information, the decision suggestion information is routine monitoring; If the risk level warning information is a medium-risk warning information, the decision recommendation information is to strengthen observation; If the risk level warning information is a high-risk warning information, the decision suggestion information is to recommend that the patient undergo further examination; If the risk level warning information is an extremely high risk warning information, the decision suggestion information is to recommend emergency intervention.
[0020] The method disclosed in the above embodiments of the present invention integrates three core medical data types: clinical physiological parameters, text medical records, and imaging data. It fully explores the complementary information of different types of data, avoids the one-sidedness of early warning caused by a single data source, and improves the comprehensiveness and accuracy of disease risk assessment. It uses an adaptive machine learning algorithm to construct a dynamic risk assessment model, which can adjust the assessment parameters in real time according to the dynamic changes of the patient's disease data, adapt to the individual differences of different patients and the evolution of the disease, and improve the timeliness and dynamic adaptability of early warning. It has strong clinical applicability: it generates targeted decision-making suggestions through a risk grading mechanism, providing clinicians with clear and operable diagnostic and treatment references. The differentiated suggestions for low-risk routine monitoring and extremely high-risk emergency intervention help optimize the allocation of medical resources, improve the efficiency of diagnosis and treatment and the prognosis of patients.
[0021] A disease early warning system based on multimodal data fusion, characterized in that the system comprises: The data acquisition module is used to acquire multimodal data information, including the patient's clinical physiological parameters, text medical record data, and image data. The data processing module is used to clean, denoise, and normalize the multimodal data information to obtain standardized data information. The fusion decision module is used to extract features from the standardized data information to obtain feature data information, and to fuse the feature data information to obtain a fusion decision result. The evaluation output module is used to import the fused decision results into the constructed dynamic risk assessment model to obtain and output the risk assessment results, wherein the dynamic risk assessment model satisfies:
[0022] In the above formula, Based on the risk assessment results, This refers to the vector information corresponding to multimodal data. As a potential risk factor, For bias terms, The weights of the aforementioned potential risk factors, It is the Sigmoid activation function.
[0023] The system disclosed in the above embodiments of the present invention integrates three core medical data types: clinical physiological parameters, text medical records, and imaging data. It fully leverages the complementary information of different data types, avoiding the biased nature of early warnings caused by a single data source, and improving the comprehensiveness and accuracy of disease risk assessment. It employs an adaptive machine learning algorithm to construct a dynamic risk assessment model, which can adjust assessment parameters in real time according to the dynamic changes in patient disease data, adapting to individual differences and disease progression among different patients, thus improving the timeliness and dynamic adaptability of early warnings. It has strong clinical applicability: through a risk grading mechanism, it generates targeted decision-making suggestions, providing clinicians with clear and actionable diagnostic and treatment references. Differentiated suggestions for low-risk routine monitoring and extremely high-risk emergency intervention help optimize the allocation of medical resources, improve diagnostic and treatment efficiency, and enhance patient prognosis.
[0024] A storage medium containing computer-executable instructions, which, when executed by a computer processor, implement the above-described disease early warning method based on multimodal data fusion.
[0025] The storage medium of this invention integrates three core medical data types—clinical physiological parameters, text medical records, and imaging data—through the disclosed method. It fully leverages the complementary information of different data types, avoiding the biased nature of early warnings caused by a single data source, and improving the comprehensiveness and accuracy of disease risk assessment. An adaptive machine learning algorithm is used to construct a dynamic risk assessment model, which can adjust assessment parameters in real time according to the dynamic changes in patient disease data, adapting to individual differences and disease progression among different patients, thus improving the timeliness and dynamic adaptability of early warnings. It has strong clinical applicability: through a risk grading mechanism, it generates targeted decision-making suggestions, providing clinicians with clear and actionable diagnostic and treatment references. Differentiated suggestions for low-risk routine monitoring and extremely high-risk emergency intervention help optimize the allocation of medical resources, improve diagnostic and treatment efficiency, and enhance patient prognosis. Attached Figure Description
[0026] Figure 1 This is a flowchart of a disease early warning method based on multimodal data fusion disclosed in the first preferred embodiment of the present invention; Figure 2 This is a schematic diagram of the module of the disease early warning method based on multimodal data fusion disclosed in the second preferred embodiment of the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0028] It should be noted that when an element is referred to as being "set on" another element, it can be directly on the other element or there may be an intervening element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or there may be an intervening element. The terms "vertical," "horizontal," "left," "right," and similar expressions used herein are for illustrative purposes only and do not represent the only possible implementation.
[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0030] like Figure 1 As shown, the first preferred embodiment of the present invention discloses a disease early warning method based on multimodal data fusion, the method comprising: S10: Acquire multimodal data information, which includes the patient's clinical physiological parameters, text medical record data, and image data.
[0031] Specifically, in step S10, the aforementioned clinical physiological parameters establish communication connections with the electrocardiogram monitor, blood pressure monitor, and pulse oximeter via the HL7 protocol to collect parameters such as the patient's heart rate, blood pressure, and blood oxygen saturation in real time. The aforementioned text medical record data is obtained from the HIS system through database queries or interface calls. A named entity recognition algorithm using natural language processing is employed to transform the unstructured medical record text into structured data, extracting key information such as the patient's basic information, past medical history, medication history, and symptom descriptions. The aforementioned imaging data interfaces with the PACS system via the DICOM protocol to collect the patient's CT, MRI, X-ray, and other medical imaging data.
[0032] This step integrates three core types of medical data: clinical physiological parameters, text medical records, and imaging data. It fully leverages the complementary information from different data types, avoiding the biased warnings caused by a single data source, and improving the comprehensiveness and accuracy of disease risk assessment. S20: The multimodal data information is cleaned, denoised, and normalized to obtain standardized data information.
[0033] In this embodiment, step S20 first requires structuring the multimodal data information, specifically including: The multimodal data is structured and classified to clarify the physical meaning, value range, and data type of each sub-item. For example, in this embodiment, the classification results of the multimodal data are shown in the table below:
[0034] In step S20 above, the cleaning of multimodal data information is mainly to remove missing values, outliers and duplicate data in the data, so as to ensure the integrity and consistency of the data.
[0035] Specifically, the removal of missing values, outliers, and duplicate data from the data of different modalities mentioned above includes, for the aforementioned clinical physiological parameters, a dual filtering method using the "3σ principle and clinical common sense threshold" is employed. For example, the normal range for systolic blood pressure is 60-230 mmHg; data exceeding this range or deviating from the mean by more than 3σ are considered outliers and directly removed. For the aforementioned text medical record data, in this embodiment, if structured data (such as the length of medical history) exceeds a reasonable range (e.g., ...), the data is removed. (120 years), which is determined to be an input error. The median of this parameter is used to fill in the data before participating in the extreme value statistics. In response to the above-mentioned data, the box plot method is used to identify outliers. Pixel values or feature values that are out of range are replaced with the upper or lower limit values.
[0036] Secondly, denoising the aforementioned multimodal data mainly includes using Kalman filtering or median filtering to remove acquisition noise from clinical physiological parameters; using syntax checking and semantic analysis techniques to correct grammatical errors and remove meaningless information from text medical record data; and using Gaussian filtering or mean filtering to reduce image noise and improve image quality from image data.
[0037] Finally, the multimodal data information is normalized, mainly including: The multimodal data information is imported into a normalized data model, which satisfies the following:
[0038] in, For the imported multimodal data information, The maximum value of the imported data. The minimum value of the imported data. This is data that has undergone normalization.
[0039] In this embodiment, during the normalization process of the multimodal data information, extreme value statistics are performed on the historical clinical data. Specifically, based on the historical clinical data, the minimum value after removing outliers is calculated separately for each data sub-item. and maximum value An initial extreme value dictionary is established; a sliding window mechanism is used to periodically incorporate newly collected valid data and recalculate the values of each sub-item. and Update the extreme value dictionary to ensure that normalization adapts to the distribution changes of clinical data; in addition, for parameters with clear clinical standard ranges, use the extreme values of the clinical standard range to avoid unreasonable extreme values due to sample bias.
[0040] More specifically, regarding the normalization processing of the aforementioned clinical physiological parameters, for continuous numerical parameters of clinical physiological parameters with clear physical meaning, such as heart rate, blood pressure, and blood oxygen saturation, the cleaned and denoised clinical physiological parameters are read. (e.g., a patient's systolic blood pressure) =140mmHg); retrieve the corresponding value for this parameter from the extreme value dictionary. (such as systolic blood pressure) =90) and (such as systolic blood pressure) =180); Applying the above normalized data model: systolic blood pressure =140; Boundary check: If the calculation result exceeds the [0,1] interval (e.g., outliers are not completely removed), then it is forcibly truncated to 0 ( (time) or 1 ( In this embodiment, batch processing is supported: all clinical physiological parameters are normalized one by one according to the above steps, and the standardized parameter matrix is output.
[0041] The features corresponding to the above-mentioned text medical record data are normalized, mainly including: after the text medical records are converted into structured data through named entity recognition, they are divided into two types of features: numerical structured features and semantic features; the above-mentioned numerical structured features (such as the length of medical history and duration of medication) are normalized according to the normalization process of clinical physiological parameters, for example, the length of medical history x=10 years ( =0, =50), after normalization; if the original semantic feature values are already distributed in the [0,1] interval, only consistency check needs to be performed to ensure that there are no values outside the range. If some values exceed [0,1] due to keyword weight calculation deviation (e.g., 1.05), then the formula is used to re-normalize, using the global TF-IDF value. and As a benchmark (e.g.) =0, =1.05), compress the data to [0,1]; normalize the categorical variables after encoding (such as discrete variables such as disease type, medication type, etc.): first use one-hot encoding to convert into binary vectors (such as "hypertension" encoded as [1,0,0], "diabetes" encoded as [0,1,0]); count the extreme values of the encoded vectors by column (each category) and perform normalization (because the value after encoding is 0 or 1, it remains unchanged after normalization, only to unify the data scale).
[0042] The image data is normalized based on its corresponding features: After denoising, the image data is divided into two categories: original pixel values and deep feature values. The processing methods for original pixel values and deep feature values are as follows: The original pixel values are normalized by separately calculating the extreme values according to the image modality: For example, the pixel value range of CT images is usually -1000 (air) to 500 (soft tissue). First, the negative pixel values are shifted to the non-negative interval (x'=x+1000). =0, =1500; Apply the normalization formula: map pixel values to [0,1]; Channel images (e.g., RGB ultrasound images): Calculate the extreme values for each channel separately and normalize them to ensure consistent data scale across channels; Deep feature normalization (e.g., texture features and lesion features extracted by CNN): Deep features may exhibit a normal distribution (mean ≠ 0, variance ≠ 1), first calculate the feature's performance in the training set. and (such as texture features) =-5, =15); after normalization using the formula, if the characteristic distribution is too concentrated (e.g., variance... If the value is 0.01, a combination strategy of "normalization + standardization" is adopted: first normalize to [0,1], then perform Z-score standardization to improve feature discrimination; targeted normalization of lesion areas: for lesion areas that have been marked in the image, extract the pixel values of the area separately, count the local extrema (rather than the global extrema), perform normalization, strengthen the difference between lesion features and normal tissue, and improve the effectiveness of subsequent feature extraction.
[0043] S30: Extract features from the standardized data information to obtain feature data information, and fuse the feature data information to obtain the fused potential risk factors.
[0044] Specifically, feature extraction is performed on the aforementioned standard spoken data information to obtain feature data information, which is then fused to obtain potential risk factors. Mean, variance, and peak value time-domain features and power spectral density frequency-domain features are extracted from clinical physiological parameters to obtain clinical physiological parameter feature vectors. TF-IDF is used to extract features from text data to obtain text feature vectors; convolutional neural networks are used to extract features from image data to obtain image feature vectors. The fusion based on the feature data information satisfies:
[0045] in, To integrate decision-making results, For clinical physiological parameter feature vectors, The weights of the feature vector of clinical physiological parameters. For text feature vectors, For text feature vector weights, For image feature vectors, The weights are the image feature vectors. Next, the aforementioned feature data information is imported into the risk knowledge base to obtain potential risk factors. Specifically, the above steps in this embodiment extract features from the aforementioned clinical physiological parameters. Generally, clinical physiological parameters (such as heart rate, blood pressure, blood oxygen saturation, etc.) are continuous numerical data, containing two types of core information: time domain and frequency domain. In this embodiment, a combined extraction strategy of "time domain features + frequency domain features" is adopted.
[0046] More specifically, the aforementioned time-domain features include standardized data output from step S20, grouped by parameter type (e.g., circulatory system parameters: heart rate, blood pressure; respiratory system parameters: blood oxygen saturation, respiratory rate). Each group of data is used to construct a separate time-series matrix (dimension: number of timestamps × number of parameters) to ensure targeted feature extraction. Next, focusing on the numerical variation patterns of the parameters, the following key features are extracted, including basic statistical features, trend features, and abnormal correlation features. The basic statistical features include: mean (reflecting the overall level of the parameter, such as the 24-hour average heart rate), variance (reflecting the degree of parameter fluctuation, such as a larger variance in blood pressure indicating poorer blood pressure stability), standard deviation, median, and range (the difference between the maximum and minimum values); trend features include: slope (reflecting the trend of parameter changes, such as the upward / downward slope of heart rate over time), and number of inflection points (such as the number of sudden drops in blood oxygen saturation, indicating a risk of respiratory abnormalities); abnormal correlation features include: peak values (such as peak systolic blood pressure), trough values (such as trough diastolic blood pressure), and the percentage of abnormal values (the percentage of values exceeding the clinically normal range, such as heart rate). The percentage of time spent at 100 times per minute), and the duration of continuous abnormalities (such as blood oxygen saturation). (90% of the duration). Example: For a patient's 24-hour heart rate data (after standardization), extract the mean = 0.62, variance = 0.08, peak value = 0.85, and number of inflection points = 3, and construct a time-domain feature subvector. .
[0047] The frequency domain feature extraction described above converts the time-domain signal into a frequency-domain signal using Fourier Transform (FFT) to extract features related to physiological rhythms. Specifically, frequency domain features include power spectral density (PSD): calculating the power proportion of different frequency ranges, such as the power proportion of the low-frequency band (0.04-0.15Hz) (reflecting sympathetic nerve activity) and the power proportion of the high-frequency band (0.15-0.4Hz) (reflecting parasympathetic nerve activity) of the heart rate signal; dominant frequency, dominant frequency power, and frequency band energy ratio. Example: After performing an FFT on the heart rate signal, extract the low-frequency power proportion = 0.65, the high-frequency power proportion = 0.35, and the dominant frequency = 0.08Hz to construct a frequency domain feature sub-vector. .
[0048] The time-domain feature vectors and frequency-domain feature vectors are concatenated to form a complete clinical physiological parameter feature vector. The formula is as follows:
[0049] In the example above, Here, the dimension is 7, and the specific dimension is adjusted according to the number of features extracted.
[0050] Feature extraction for the aforementioned text medical record data includes processing the data in step S2 into structured numerical data (e.g., years of medical history) and unstructured text keywords (e.g., symptoms, diagnosis). A combined strategy of "direct extraction of structured features + extraction of semantic features from unstructured text" is employed. The structured text feature extraction involves directly using the structured data (e.g., age, years of medical history, duration of medication, number of surgeries, etc.) after name entity recognition transformation as basic features, sorted by "clinical relevance" to form structured feature sub-vectors. Example: Patient age = 55 years (standardized = 0.55), history of hypertension = 0 years (standardized = 0.2), duration of medication = 8 years (standardized = 0.16), number of surgeries = 0. Construct... . The above-mentioned extraction of unstructured text semantic features includes extracting keyword semantic features from unstructured texts such as symptom descriptions, diagnosis results, and doctor's orders using the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm. The steps are as follows: Text preprocessing: Tokenize the structured text (using a dedicated tokenization tool in the medical field) and remove stop words (eliminating meaningless words such as "de" [的], "le" [了], etc., and general medical vocabulary such as "patient", "visit", etc.); Keyword screening: Based on the medical risk knowledge base, preset a keyword dictionary related to diseases (such as cardiovascular disease keywords: chest tightness, chest pain, palpitations, hypertension, coronary artery stenosis, etc.), and only retain the keywords in the dictionary to participate in feature calculation; TF-IDF calculation: Term frequency (TF): The number of occurrences of a keyword in the current medical record / the total number of words in the current medical record; Inverse document frequency (IDF): The total number of medical records / the number of medical records containing the keyword, which is used to reduce the weight of general keywords. Keyword weight: ; Construction of semantic feature sub-vectors: Select the keywords with the top-N TF-IDF values (N is adjusted according to the length of the medical record text, usually N = 20), and sort them by weight to form semantic feature sub-vectors. Example: After processing the medical record text of a certain patient, the keywords "chest tightness" ( ), "hypertension" ( ), "aggravated after activity" ( ) are used to construct .
[0051] Finally, construct the text medical record feature vector, which specifically includes concatenating the structured feature sub-vector and the semantic feature sub-vector to form a complete text medical record feature vector. The formula is as follows: , Example: . The above dimension is 7, and the specific dimension in this embodiment can be adjusted according to the number of structured features and the number of keywords.
[0052] Feature extraction is performed on the aforementioned image data, including manual feature extraction and feature extraction based on a CNN model. The manual feature extraction described above is a shallow feature extraction. In this embodiment, manual feature extraction specifically includes extracting morphological features, density / grayscale features, and texture features from the segmented lesion region. The morphological features include: lesion area (number of pixels), perimeter, roundness, major axis / minor axis ratio, and boundary roughness. The density / grayscale features include: average grayscale value of the lesion region, grayscale variance, grayscale entropy (reflecting the uniformity of grayscale distribution), and peak position of the grayscale histogram. The texture features include: extracting energy (reflecting texture uniformity), contrast (reflecting texture clarity), correlation (reflecting the degree of association between texture pixels), and entropy (reflecting texture complexity) using the Gray-Level Co-occurrence Matrix (GLCM). Example: For a lesion region in a patient's coronary artery CT scan, the extracted area = 0.72 (after standardization), roundness = 0.65, average grayscale value = 0.58, and texture energy = 0.81, constructing a manual feature subvector. The above feature extraction based on the CNN model includes using a pre-trained CNN model and fine-tuning it based on a medical image dataset to extract deep features. In this embodiment, the CNN model mentioned above may include ResNet50, VGG16, etc. Finally, the hand-crafted feature vectors are concatenated with the deep semantic feature vectors to form a complete image data feature vector. The formula is as follows:
[0053] In this embodiment, the above steps employ a combined strategy of "weighted summation fusion + feature selection optimization" to... , , The three types of feature vectors are fused into a unified fusion decision result, ensuring the simplicity and discriminative power of the fused features.
[0054] Specifically, this step determines the weights of clinical physiological parameters based on the clinical importance and feature discriminative power of each modality feature, using a weighting method of "clinical experience weights + model adaptive weights". Text feature weights Image feature weights The above weights satisfy:
[0055] In this embodiment, an initial weight allocation can be performed on the above formula. For example, the initial weight can be set based on clinical diagnosis and treatment logic. For instance, in cardiovascular disease early warning, imaging features (such as the degree of vascular stenosis) have the highest weight. =0.5), followed by clinical physiological parameters ( =0.3), text features again ( =0.2); then, adaptive weight optimization is performed: using "classification accuracy of fused features" as the objective function, the weight values are adjusted through grid search or particle swarm optimization (PSO) algorithms. For example, if the contribution of "history of myocardial infarction" in the text features to risk assessment is higher than initially expected, the weights are automatically increased. To 0.25, correspondingly reduced Set the weight to 0.25 to ensure that the weight matches the actual value of the feature.
[0056] Finally, the feature data information is fused to satisfy:
[0057] in, To integrate decision-making results, For clinical physiological parameter feature vectors, The weights of the feature vector of clinical physiological parameters. For text feature vectors, For text feature vector weights, For image feature vectors, The weights are the image feature vectors. Finally, the aforementioned feature data is imported into a risk knowledge base to obtain potential risk factors. In this embodiment, the aforementioned potential risk factors refer to key factors directly related to the patient's condition that may lead to disease progression or poor prognosis, such as a history of hypertension, coronary artery stenosis, and chest tightness symptoms. These are obtained by matching the fusion features with the risk knowledge base. In this embodiment, the aforementioned risk knowledge base is a structured database containing association rules of "feature—risk factor—disease type—risk level". The fusion features are matched with the knowledge base, and the filtered fusion features are semantically mapped, converting the feature items into standard terms that the knowledge base can recognize (e.g., "heart rate variance" is mapped to "heart rate variability"). Then, matching is performed according to "modal priority": first, image features (lesion-related), then clinical physiological parameter features (physiological abnormality-related), and finally text features (medical history, symptom-related). The top N core potential risk factors are output for subsequent parameter adjustment of the dynamic risk assessment model.
[0058] S40: Import the fused decision results into the constructed dynamic risk assessment model to obtain and output the risk assessment results, wherein the dynamic risk assessment model satisfies:
[0059] In the above formula, Based on the risk assessment results, This refers to the vector information corresponding to multimodal data. As a potential risk factor, For bias terms, The weights of the aforementioned potential risk factors, It is the Sigmoid activation function.
[0060] Specifically, in this step, the input data is integrated: the filtered and merged decision results output from step S3 are combined. With potential risk factors The model input matrix X is formed through structured integration; potential risk factors are encoded using one-hot encoding to convert discrete risk factors into binary vectors. In this embodiment, the above model adopts a three-stage architecture of "feature enhancement layer + adaptive classification layer + dynamic update layer" to adapt to the characteristics of high dimensionality and imbalanced sample classes in multimodal data.
[0061] In this embodiment, the above-mentioned standardized data information is used to extract features to obtain feature data information, and the feature data information is fused to obtain a fusion decision result. This includes: using the fused high-dimensional feature vector as sample features and the disease diagnosis and risk level results confirmed by clinical physician consultation as sample labels, and combining the characteristics of high dimensionality and imbalanced sample categories of multimodal data, a machine learning algorithm with both feature screening capability and classification accuracy is used to construct a dynamic risk assessment model.
[0062] In this embodiment, the step of outputting risk assessment results based on the dynamic risk assessment model includes classifying the risk assessment results to obtain risk level early warning information; and generating decision-making suggestion information based on the risk level early warning information. Specifically, the decision-making suggestion information includes: if the risk level early warning information is low-risk, the decision-making suggestion is routine monitoring; if the risk level early warning information is medium-risk, the decision-making suggestion is enhanced observation; if the risk level early warning information is high-risk, the decision-making suggestion is to recommend further examination for the patient; and if the risk level early warning information is extremely high-risk, the decision-making suggestion is to recommend emergency intervention.
[0063] like Figure 2 As shown, the second preferred embodiment of the present invention discloses a disease early warning system 100 based on multimodal data fusion. The system 100 includes a data acquisition module 110, a data processing module 120, a fusion decision module 130, and an evaluation output module 140.
[0064] The aforementioned data acquisition module 110 is used to acquire multimodal data information, which includes the patient's clinical physiological parameters, text medical record data, and image data; Specifically, the aforementioned clinical physiological parameters of the data acquisition module 110 are communicated with the electrocardiogram monitor, blood pressure monitor, and pulse oximeter via the HL7 protocol to collect parameters such as the patient's heart rate, blood pressure, and blood oxygen saturation in real time. The aforementioned text medical record data is obtained from the HIS system through database queries or interface calls. A named entity recognition algorithm using natural language processing is employed to transform the unstructured medical record text into structured data, extracting key information such as the patient's basic information, past medical history, medication history, and symptom descriptions. The aforementioned imaging data is interfaced with the PACS system via the DICOM protocol to collect the patient's CT, MRI, X-ray, and other medical imaging data.
[0065] This system integrates three core medical data types: clinical physiological parameters, text medical records, and imaging data. It fully leverages the complementary information of different data types, avoids the one-sidedness of early warnings caused by a single data source, and improves the comprehensiveness and accuracy of disease risk assessment.
[0066] The aforementioned data processing module 120 is used to clean, denoise, and normalize the multimodal data information to obtain standardized data information.
[0067] In this embodiment, the data processing module 120 first needs to structure the multimodal data information, specifically including: The multimodal data is structured and classified to clarify the physical meaning, value range and data type of each sub-item. The data processing module 120 cleans the multimodal data information mainly to remove missing values, outliers and duplicate data in the data to ensure the integrity and consistency of the data.
[0068] Specifically, the removal of missing values, outliers, and duplicate data from the data of different modalities mentioned above includes, for the aforementioned clinical physiological parameters, a dual filtering method using the "3σ principle and clinical common sense threshold" is employed. For example, the normal range for systolic blood pressure is 60-230 mmHg; data exceeding this range or deviating from the mean by more than 3σ are considered outliers and directly removed. For the aforementioned text medical record data, in this embodiment, if structured data (such as the length of medical history) exceeds a reasonable range (e.g., ...), the data is removed. (120 years), which is determined to be an input error. The median of this parameter is used to fill in the data before participating in the extreme value statistics. In response to the above-mentioned data, the box plot method is used to identify outliers. Pixel values or feature values that are out of range are replaced with the upper or lower limit values.
[0069] Secondly, denoising the aforementioned multimodal data mainly includes using Kalman filtering or median filtering to remove acquisition noise from clinical physiological parameters; using syntax checking and semantic analysis techniques to correct grammatical errors and remove meaningless information from text medical record data; and using Gaussian filtering or mean filtering to reduce image noise and improve image quality from image data.
[0070] Finally, the multimodal data information is normalized, mainly including: The multimodal data information is imported into a normalized data model, which satisfies the following:
[0071] in, For the imported multimodal data information, The maximum value of the imported data. The minimum value of the imported data. This is data that has undergone normalization.
[0072] In this embodiment, during the normalization process of the multimodal data information, extreme value statistics are performed on the historical clinical data. Specifically, based on the historical clinical data, the minimum value after removing outliers is calculated separately for each data sub-item. and maximum value An initial extreme value dictionary is established; a sliding window mechanism is used to periodically incorporate newly collected valid data and recalculate the values of each sub-item. and Update the extreme value dictionary to ensure that normalization adapts to the distribution changes of clinical data; in addition, for parameters with clear clinical standard ranges, use the extreme values of the clinical standard range to avoid unreasonable extreme values due to sample bias.
[0073] More specifically, regarding the normalization processing of the aforementioned clinical physiological parameters, for continuous numerical parameters of clinical physiological parameters with clear physical meaning, such as heart rate, blood pressure, and blood oxygen saturation, the cleaned and denoised clinical physiological parameters are read. (e.g., a patient's systolic blood pressure) =140mmHg); retrieve the corresponding value for this parameter from the extreme value dictionary. (such as systolic blood pressure) =90) and (such as systolic blood pressure) =180); Applying the above normalized data model: systolic blood pressure =140; Boundary check: If the calculation result exceeds the [0,1] interval (e.g., outliers are not completely removed), then it is forcibly truncated to 0 ( (time) or 1 ( In this embodiment, batch processing is supported: all clinical physiological parameters are normalized one by one as described above, and the standardized parameter matrix is output.
[0074] The features corresponding to the above-mentioned text medical record data are normalized, mainly including: after the text medical records are converted into structured data through named entity recognition, they are divided into two types of features: numerical structured features and semantic features; the above-mentioned numerical structured features (such as the length of medical history and duration of medication) are normalized according to the normalization process of clinical physiological parameters, for example, the length of medical history x=10 years ( =0, =50), after normalization; if the original semantic feature values are already distributed in the [0,1] interval, only consistency check needs to be performed to ensure that there are no values outside the range. If some values exceed [0,1] due to keyword weight calculation deviation (e.g., 1.05), then the formula is used to re-normalize, using the global TF-IDF value. and As a benchmark (e.g.) =0, =1.05), compress the data to [0,1]; normalize the categorical variables after encoding (such as discrete variables such as disease type, medication type, etc.): first use one-hot encoding to convert into binary vectors (such as "hypertension" encoded as [1,0,0], "diabetes" encoded as [0,1,0]); count the extreme values of the encoded vectors by column (each category) and perform normalization (because the value after encoding is 0 or 1, it remains unchanged after normalization, only to unify the data scale).
[0075] The image data is normalized based on its corresponding features: After denoising, the image data is divided into two categories: original pixel values and deep feature values. The processing methods for original pixel values and deep feature values are as follows: The original pixel values are normalized by separately calculating the extreme values according to the image modality: For example, the pixel value range of CT images is usually -1000 (air) to 500 (soft tissue). First, the negative pixel values are shifted to the non-negative interval (x'=x+1000). =0, =1500; Apply the normalization formula: map pixel values to [0,1]; Channel images (e.g., RGB ultrasound images): Calculate the extreme values for each channel separately and normalize them to ensure consistent data scale across channels; Deep feature normalization (e.g., texture features and lesion features extracted by CNN): Deep features may exhibit a normal distribution (mean ≠ 0, variance ≠ 1), first calculate the feature's performance in the training set. and (such as texture features) =-5, =15); after normalization using the formula, if the characteristic distribution is too concentrated (e.g., variance... If the value is 0.01, a combination strategy of "normalization + standardization" is adopted: first normalize to [0,1], then perform Z-score standardization to improve feature discrimination; targeted normalization of lesion areas: for lesion areas that have been marked in the image, extract the pixel values of the area separately, count the local extrema (rather than the global extrema), perform normalization, strengthen the difference between lesion features and normal tissue, and improve the effectiveness of subsequent feature extraction.
[0076] The aforementioned fusion decision module 130 is used to extract features from the standardized data information to obtain feature data information, and to fuse the feature data information to obtain a fusion decision result.
[0077] Specifically, the standardized data information is subjected to feature extraction to obtain feature data information, which is then fused to obtain potential risk factors. Mean, variance, and peak value time-domain features and power spectral density frequency-domain features are extracted from clinical physiological parameters to obtain clinical physiological parameter feature vectors. TF-IDF is used to extract features from text data to obtain text feature vectors. Convolutional neural networks are used to extract features from image data to obtain image feature vectors.
[0078] The fusion based on the feature data information satisfies:
[0079] in, To integrate decision-making results, For clinical physiological parameter feature vectors, The weights of the feature vector of clinical physiological parameters. For text feature vectors, For text feature vector weights, For image feature vectors, The weights are the image feature vectors.
[0080] Next, the aforementioned feature data information is imported into the risk knowledge base to obtain potential risk factors. Specifically, in this embodiment, feature extraction is performed on the aforementioned clinical physiological parameters. Generally, clinical physiological parameters (such as heart rate, blood pressure, blood oxygen saturation, etc.) are continuous numerical data, containing two types of core information: time domain and frequency domain. In this embodiment, a combined extraction strategy of "time domain features + frequency domain features" is adopted.
[0081] More specifically, the aforementioned time-domain features include the output standardized data, grouped by parameter type (e.g., circulatory system parameters: heart rate, blood pressure; respiratory system parameters: blood oxygen saturation, respiratory rate). Each group of data is used to construct a separate time-series matrix (dimension: number of timestamps × number of parameters) to ensure targeted feature extraction. Next, focusing on the numerical variation patterns of the parameters, the following key features are extracted, including basic statistical features, trend features, and abnormal correlation features. The basic statistical features include: mean (reflecting the overall level of the parameter, such as the 24-hour average heart rate), variance (reflecting the degree of parameter fluctuation, such as a larger variance in blood pressure indicating poorer blood pressure stability), standard deviation, median, and range (the difference between the maximum and minimum values); trend features include: slope (reflecting the trend of parameter changes, such as the upward / downward slope of heart rate over time), and number of inflection points (such as the number of sudden drops in blood oxygen saturation, indicating a risk of respiratory abnormalities); abnormal correlation features include: peak values (such as peak systolic blood pressure), trough values (such as trough diastolic blood pressure), and the percentage of abnormal values (the percentage of values exceeding the clinically normal range, such as heart rate). The percentage of time spent at 100 times per minute), and the duration of continuous abnormalities (such as blood oxygen saturation). (90% of the duration). Example: For a patient's 24-hour heart rate data (after standardization), extract the mean = 0.62, variance = 0.08, peak value = 0.85, and number of inflection points = 3, and construct a time-domain feature subvector. .
[0082] The frequency domain feature extraction described above converts the time-domain signal into a frequency-domain signal using Fourier Transform (FFT) to extract features related to physiological rhythms. Specifically, frequency domain features include power spectral density (PSD): calculating the power proportion of different frequency ranges, such as the power proportion of the low-frequency band (0.04-0.15Hz) (reflecting sympathetic nerve activity) and the power proportion of the high-frequency band (0.15-0.4Hz) (reflecting parasympathetic nerve activity) of the heart rate signal; dominant frequency, dominant frequency power, and frequency band energy ratio. Example: After performing an FFT on the heart rate signal, extract the low-frequency power proportion = 0.65, the high-frequency power proportion = 0.35, and the dominant frequency = 0.08Hz to construct a frequency domain feature sub-vector. .
[0083] The time-domain feature vectors and frequency-domain feature vectors are concatenated to form a complete clinical physiological parameter feature vector. The formula is as follows:
[0084] In the example above, Here, the dimension is 7, and the specific dimension is adjusted according to the number of features extracted.
[0085] Feature extraction for the above-mentioned text medical record data includes that after processing the text medical record data, it becomes structured numerical data (such as the duration of the medical history) and unstructured text keywords (such as symptoms, diagnoses). A combined strategy of "direct extraction of structured features + extraction of semantic features of unstructured text" is adopted. The above-mentioned extraction of structured text features includes directly using the structured data (such as age, duration of the medical history, duration of medication, number of surgeries, etc.) after named entity recognition as basic features, and forming a structured feature sub-vector after sorting according to "clinical relevance". Example: Patient age = 55 years old (after standardization = 0.55), hypertension medical history = 0 years (after standardization = 0.2), duration of medication = 8 years (after standardization = 0.16), number of surgeries = 0, constructing . The above-mentioned extraction of semantic features of unstructured text includes using the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm to extract keyword semantic features for unstructured texts such as symptom descriptions, diagnosis results, medical orders, etc. Specifically as follows: Text preprocessing: Tokenize the structured text (using a special tokenization tool in the medical field), and remove stop words (removing meaningless words such as "de", "le", etc., and general medical words such as "patient", "visit"); Keyword screening: Based on the medical risk knowledge base, preset a keyword dictionary related to diseases (such as cardiovascular disease keywords: chest tightness, chest pain, palpitations, hypertension, coronary artery stenosis, etc.), and only retain the keywords in the dictionary to participate in feature calculation; TF-IDF calculation: Term frequency (TF): The number of occurrences of a keyword in the current medical record / the total number of words in the current medical record; Inverse document frequency (IDF): The total number of medical records / the number of medical records containing the keyword, which is used to reduce the weight of general keywords. Keyword weight: ; Construction of semantic feature sub-vector: Select the keywords with the top-N TF-IDF values (N is adjusted according to the length of the medical record text, usually N = 20), and form a semantic feature sub-vector after sorting according to the weight. Example: After processing the medical record text of a certain patient, the keywords "chest tightness" ( ), "hypertension" ( ), "aggravated after activity" ( ), constructing .
[0086] Finally, construct a text medical record feature vector, specifically including concatenating the structured feature sub-vector and the semantic feature sub-vector to form a complete text medical record feature vector. The formula is as follows: , Example: , The above dimension is 7-dimensional, and the specific dimension in this embodiment can be adjusted according to the number of structured features and the number of keywords.
[0087] Feature extraction is performed on the aforementioned image data, including manual feature extraction and feature extraction based on a CNN model. The manual feature extraction described above is a shallow feature extraction. In this embodiment, manual feature extraction specifically includes extracting morphological features, density / grayscale features, and texture features from the segmented lesion region. The morphological features include: lesion area (number of pixels), perimeter, roundness, major axis / minor axis ratio, and boundary roughness. The density / grayscale features include: average grayscale value of the lesion region, grayscale variance, grayscale entropy (reflecting the uniformity of grayscale distribution), and peak position of the grayscale histogram. The texture features include: extracting energy (reflecting texture uniformity), contrast (reflecting texture clarity), correlation (reflecting the degree of association between texture pixels), and entropy (reflecting texture complexity) using the Gray-Level Co-occurrence Matrix (GLCM). Example: For a lesion region in a patient's coronary artery CT scan, the extracted area = 0.72 (after standardization), roundness = 0.65, average grayscale value = 0.58, and texture energy = 0.81, constructing a manual feature subvector. The above feature extraction based on the CNN model includes using a pre-trained CNN model and fine-tuning it based on a medical image dataset to extract deep features. In this embodiment, the CNN model mentioned above may include ResNet50, VGG16, etc. Finally, the hand-crafted feature vectors are concatenated with the deep semantic feature vectors to form a complete image data feature vector. The formula is as follows:
[0088] In this embodiment, the above-mentioned combined strategy of "weighted summation fusion + feature selection optimization" is adopted to... , , The three types of feature vectors are fused into a unified fusion decision result, ensuring the simplicity and discriminative power of the fused features.
[0089] Specifically, based on the clinical importance and feature discriminative power of each modality feature, the weights of clinical physiological parameters are determined using a weighting method of "clinical experience weights + model adaptive weights". Text feature weights Image feature weights The above weights satisfy:
[0090] In this embodiment, an initial weight allocation can be performed on the above formula. For example, the initial weight can be set based on clinical diagnosis and treatment logic. For instance, in cardiovascular disease early warning, imaging features (such as the degree of vascular stenosis) have the highest weight. =0.5), followed by clinical physiological parameters ( =0.3), text features again ( =0.2); then, adaptive weight optimization is performed: using "classification accuracy of fused features" as the objective function, the weight values are adjusted through grid search or particle swarm optimization (PSO) algorithms. For example, if the contribution of "history of myocardial infarction" in the text features to risk assessment is higher than initially expected, the weights are automatically increased. To 0.25, correspondingly reduced Set the weight to 0.25 to ensure that the weight matches the actual value of the feature.
[0091] Finally, the feature data information is fused to satisfy:
[0092] in, To integrate decision-making results, For clinical physiological parameter feature vectors, The weights of the feature vector of clinical physiological parameters. For text feature vectors, For text feature vector weights, For image feature vectors, The weights are the image feature vectors. Finally, the aforementioned feature data is imported into a risk knowledge base to obtain potential risk factors. In this embodiment, the aforementioned potential risk factors refer to key factors directly related to the patient's condition that may lead to disease progression or poor prognosis, such as a history of hypertension, coronary artery stenosis, and chest tightness symptoms. These are obtained by matching the fusion features with the risk knowledge base. In this embodiment, the aforementioned risk knowledge base is a structured database containing association rules of "feature—risk factor—disease type—risk level". The fusion features are matched with the knowledge base, and the filtered fusion features are semantically mapped, converting the feature items into standard terms that the knowledge base can recognize (e.g., "heart rate variance" is mapped to "heart rate variability"). Then, matching is performed according to "modal priority": first, image features (lesion-related), then clinical physiological parameter features (physiological abnormality-related), and finally text features (medical history, symptom-related). The top N core potential risk factors are output for subsequent parameter adjustment of the dynamic risk assessment model.
[0093] The aforementioned evaluation output module 140 is used to import the fusion decision results into the constructed dynamic risk assessment model to obtain and output risk assessment results, wherein the dynamic risk assessment model satisfies:
[0094] In the above formula, Based on the risk assessment results, This refers to the vector information corresponding to multimodal data. As a potential risk factor, For bias terms, The weights of the aforementioned potential risk factors, It is the Sigmoid activation function.
[0095] Specifically, the above output is the filtered and fused decision result. With potential risk factors The model input matrix X is formed through structured integration; potential risk factors are encoded using one-hot encoding to convert discrete risk factors into binary vectors. In this embodiment, the above model adopts a three-stage architecture of "feature enhancement layer + adaptive classification layer + dynamic update layer" to adapt to the characteristics of high dimensionality and imbalanced sample classes in multimodal data.
[0096] In this embodiment, the above-mentioned standardized data information is used to extract features to obtain feature data information, and the feature data information is fused to obtain a fusion decision result. This includes: using the fused high-dimensional feature vector as sample features and the disease diagnosis and risk level results confirmed by clinical physician consultation as sample labels, and combining the characteristics of high dimensionality and imbalanced sample categories of multimodal data, a machine learning algorithm with both feature screening capability and classification accuracy is used to construct a dynamic risk assessment model.
[0097] In this embodiment, the step of outputting risk assessment results based on the dynamic risk assessment model includes classifying the risk assessment results to obtain risk level early warning information; and generating decision-making suggestion information based on the risk level early warning information. Specifically, the decision-making suggestion information includes: if the risk level early warning information is low-risk, the decision-making suggestion is routine monitoring; if the risk level early warning information is medium-risk, the decision-making suggestion is enhanced observation; if the risk level early warning information is high-risk, the decision-making suggestion is to recommend further examination for the patient; and if the risk level early warning information is extremely high-risk, the decision-making suggestion is to recommend emergency intervention.
[0098] The system disclosed in the above embodiments of this invention integrates three core medical data types: clinical physiological parameters, text medical records, and imaging data. It fully explores the complementary information of different types of data, avoids the one-sidedness of early warning caused by a single data source, and improves the comprehensiveness and accuracy of disease risk assessment. It uses an adaptive machine learning algorithm to build a dynamic risk assessment model, which can adjust the assessment parameters in real time according to the dynamic changes of the patient's disease data, adapt to the individual differences of different patients and the evolution of the disease, and improve the timeliness and dynamic adaptability of early warning. It has strong clinical applicability: it generates targeted decision-making suggestions through a risk grading mechanism, providing clinicians with clear and operable diagnostic and treatment references. The differentiated suggestions for low-risk routine monitoring and extremely high-risk emergency intervention help optimize the allocation of medical resources, improve the efficiency of diagnosis and treatment and the prognosis of patients.
[0099] A storage medium containing computer-executable instructions, which, when executed by a computer processor, implement the above-described disease early warning method based on multimodal data fusion.
[0100] The method disclosed in the above embodiments of the present invention integrates three core medical data types: clinical physiological parameters, text medical records, and imaging data. It fully explores the complementary information of different types of data, avoids the one-sidedness of early warning caused by a single data source, and improves the comprehensiveness and accuracy of disease risk assessment. It uses an adaptive machine learning algorithm to construct a dynamic risk assessment model, which can adjust the assessment parameters in real time according to the dynamic changes of the patient's disease data, adapt to the individual differences of different patients and the evolution of the disease, and improve the timeliness and dynamic adaptability of early warning. It has strong clinical applicability: it generates targeted decision-making suggestions through a risk grading mechanism, providing clinicians with clear and operable diagnostic and treatment references. The differentiated suggestions for low-risk routine monitoring and extremely high-risk emergency intervention help optimize the allocation of medical resources, improve the efficiency of diagnosis and treatment and the prognosis of patients.
[0101] It should be noted that the computer storage medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0102] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0103] The aforementioned computer storage medium may be included in the aforementioned electronic device; or it may exist independently and not be assembled into the electronic device.
[0104] The aforementioned computer storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0105] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A disease condition early warning method based on multi-modal data fusion, characterized in that, include: Acquire multimodal data information, including the patient's clinical physiological parameters, text medical record data, and image data; The multimodal data information is cleaned, denoised, and normalized to obtain standardized data information; Feature extraction is performed on the standardized data information to obtain feature data information, and the feature data information is fused to obtain the fused potential risk factors; The fusion decision results are imported into the constructed dynamic risk assessment model to obtain and output risk assessment results, wherein the dynamic risk assessment model satisfies: In the above formula, Based on the risk assessment results, This refers to the vector information corresponding to multimodal data. As a potential risk factor, For bias terms, The weights of the aforementioned potential risk factors, It is the Sigmoid activation function.
2. The disease early warning method based on multimodal data fusion according to claim 1, characterized in that, The multimodal data information is normalized, including: The multimodal data information is imported into a normalized data model, which satisfies the following: in, For the imported multimodal data information, The maximum value of the imported data. The minimum value of the imported data. This is data that has undergone normalization.
3. The disease early warning method based on multimodal data fusion according to claim 1, characterized in that, The process of extracting features from the standardized data information to obtain feature data information, and then fusing the feature data information to obtain the fused potential risk factors, includes: The mean, variance, and peak value time-domain features and power spectral density frequency-domain features of clinical physiological parameters are extracted to obtain clinical physiological parameter feature vectors. The TF-IDF method is used to extract features from text data to obtain text feature vectors. Convolutional neural networks are used to extract features from image data to obtain image feature vectors. The fusion based on the feature data information satisfies: in, To integrate decision-making results, For clinical physiological parameter feature vectors, The weights of the feature vector of clinical physiological parameters. For text feature vectors, For text feature vector weights, For image feature vectors, The weights are the image feature vectors. Import the aforementioned feature data into the risk knowledge base to obtain potential risk factors.
4. The disease early warning method based on multimodal data fusion according to claim 1, characterized in that, The acquisition of multimodal data information, including the patient's clinical physiological parameters, text medical record data, and image data, includes: The clinical physiological parameters are collected by establishing a communication connection with the electrocardiogram monitor, blood pressure monitor, and pulse oximeter via the HL7 protocol. The image data is acquired by interfacing with the PACS system via the DICOM protocol. The collected medical records are obtained from the HIS system through database queries or interface calls, and the unstructured text is converted into structured data using a named entity recognition algorithm based on natural language processing.
5. The disease early warning method based on multimodal data fusion according to claim 1, characterized in that, The denoising of the multimodal data information includes: The clinical physiological parameters are denoised using Kalman filtering or median filtering, the text medical record data is denoised using syntax checking and semantic analysis, and the image data is denoised using Gaussian filtering or mean filtering.
6. The disease early warning method based on multimodal data fusion according to claim 1, characterized in that, The step of extracting features from the standardized data information to obtain feature data information, and fusing the feature data information to obtain a fusion decision result, includes: Using the fused high-dimensional feature vectors as sample features and the disease diagnosis and risk level results confirmed by clinical physician consultation as sample labels, and taking into account the characteristics of high dimensionality and imbalanced sample categories of multimodal data, a dynamic risk assessment model is constructed using a machine learning algorithm that combines feature selection capability and classification accuracy.
7. The disease early warning method based on multimodal data fusion according to claim 1, characterized in that, The risk assessment results output based on the dynamic risk assessment model include: The risk assessment results are classified into risk levels to obtain risk level early warning information; Based on the aforementioned risk level warning information, decision-making recommendations are generated.
8. The disease early warning method based on multimodal data fusion according to claim 7, characterized in that, Based on the risk level warning information, the generator produces decision-making suggestion information, including: If the risk level warning information is a low-risk warning information, the decision suggestion information is routine monitoring; If the risk level warning information is a medium-risk warning information, the decision recommendation information is to strengthen observation; If the risk level warning information is a high-risk warning information, the decision suggestion information is to recommend that the patient undergo further examination; If the risk level warning information is an extremely high risk warning information, the decision suggestion information is to recommend emergency intervention.
9. A disease early warning system based on multimodal data fusion, characterized in that, The system includes: The data acquisition module is used to acquire multimodal data information, including the patient's clinical physiological parameters, text medical record data, and image data. The data processing module is used to clean, denoise, and normalize the multimodal data information to obtain standardized data information. The fusion decision module is used to extract features from the standardized data information to obtain feature data information, and to fuse the feature data information to obtain a fusion decision result. The evaluation output module is used to import the fused decision results into the constructed dynamic risk assessment model to obtain and output the risk assessment results, wherein the dynamic risk assessment model satisfies: In the above formula, Based on the risk assessment results, This refers to the vector information corresponding to multimodal data. As a potential risk factor, For bias terms, The weights of the aforementioned potential risk factors, It is the Sigmoid activation function.
10. A storage medium containing computer-executable instructions, which, when executed by a computer processor, implement the disease early warning method based on multimodal data fusion as described in any one of claims 1-8.