Cerebral arterial thrombosis prediction system based on big data

Through active learning and dynamic weight coefficient optimization data sets, combined with Gaussian process model and risk stratified loss function, the problems of patient type imbalance and dynamic risk changes in the existing technology are solved, and the accuracy and accuracy of ischemic stroke prediction are improved.

CN120260924AActive Publication Date: 2025-07-04FOURTH MILITARY MEDICAL UNIVERSITY
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510410225.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The existing ischemic stroke prediction system has imbalance in patient type, improper handling of noise and uncertainty in clinical data, resulting in poor prediction effect and failure to consider dynamic changes in patient risks, resulting in insufficient accuracy.

Method used

Active learning combined with weighted uncertainty is used to select samples, combine the prediction error of the Gaussian process model, and dynamically adjust the patient data set, and ischemic stroke prediction model is constructed through dynamic weight coefficients and risk stratified loss functions to achieve refined management of patients with different risk levels.

Benefits of technology

The accuracy and accuracy of ischemic stroke prediction is improved, especially the ability to identify high-risk samples, and the precise classification and prediction of patients with different risk levels is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260924A_ABST
    Figure CN120260924A_ABST
Patent Text Reader

Abstract

The invention discloses a cerebral arterial thrombosis prediction system based on big data. The cerebral arterial thrombosis prediction system comprises a data acquisition module, a data set optimization module, a cerebral arterial thrombosis prediction model construction module and a cerebral arterial thrombosis prediction module. The invention belongs to the field of data processing, and particularly relates to an ischemic cerebral apoplexy prediction system based on big data, according to the scheme, samples are selected through active learning in combination with weighted uncertainty, category distribution of the samples is considered through weighting coefficients, prediction errors of a Gaussian process model are combined, and on the basis of uncertainty and rare degree of each sample, a prediction result is obtained. Dynamically adjusting the patient data set; efficient selection of a patient data set is realized; by adopting a dynamic weight coefficient, according to the gradient density automatic weight of the sample, a high-risk sample patient of the cerebral arterial thrombosis is weighted more; constructing risk hierarchical loss to accurately classify patients of different risk levels; and thus, accurate prediction of cerebral arterial thrombosis is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and specifically refers to an ischemic stroke prediction system based on big data. Background Art

[0002] Ischemic stroke prediction systems mainly rely on a large amount of medical research and clinical data, and use statistical methods and machine learning algorithms to identify risk factors closely related to the occurrence of ischemic stroke. Then, the risk of an individual having an ischemic stroke is evaluated and predicted. However, general ischemic stroke prediction systems have problems such as unbalanced patient types, improper handling of the noise and uncertainties often present in clinical data, which leads to poor final ischemic stroke prediction effects; general ischemic stroke prediction systems have problems such as unbalanced patient risk categories, failure to consider the dynamic changes of patient risks, and lack of refined management for patients with different risk levels, which leads to poor accuracy of ischemic stroke prediction. Summary of the Invention

[0003] In view of the above situation, to overcome the defects of the prior art, the present invention provides an ischemic stroke prediction system based on big data. Aiming at the problems of unbalanced patient types, improper handling of the noise and uncertainties often present in clinical data, and resulting in poor final ischemic stroke prediction effects in general ischemic stroke prediction systems, this solution uses active learning combined with weighted uncertainty to select samples, the weighting coefficient considers the class distribution of samples, combines the prediction error of the Gaussian process model, and dynamically adjusts the patient dataset based on the uncertainty and rarity of each sample; realizes the efficient selection of the patient dataset; and then improves the final ischemic stroke prediction effect; aiming at the problems of unbalanced patient risk categories, failure to consider the dynamic changes of patient risks, and lack of refined management for patients with different risk levels, and resulting in poor accuracy of ischemic stroke prediction in general ischemic stroke prediction systems, this solution uses dynamic weight coefficients, automatically weights according to the gradient density of samples, and gives higher weights to high-risk sample patients with ischemic stroke; constructs a risk stratification loss to accurately classify patients with different risk levels; and then realizes the accurate prediction of ischemic stroke.

[0004] The technical solution adopted by the present invention is as follows: The ischemic stroke prediction system based on big data provided by the present invention includes a data acquisition module, a dataset optimization module, an ischemic stroke prediction model construction module, and an ischemic stroke prediction module;

[0005] The data acquisition module collects historical patient medical data and obtains an original dataset through feature engineering processing;

[0006] The dataset optimization module uses active learning technology to optimize the dataset based on weighted uncertainty;

[0007] The ischemic stroke prediction model construction module is based on a multi-layer perceptron, and designs a dynamic weight coefficient and a risk stratification loss function to construct an ischemic stroke prediction model;

[0008] The ischemic stroke prediction module realizes the prediction of ischemic stroke for real-time patient medical data based on the established ischemic stroke prediction model.

[0009] Furthermore, in the data acquisition module, the historical patient medical data includes clinical signs, laboratory test indicators, lifestyle, and ischemic stroke status; the ischemic stroke status is used as the data label; the ischemic stroke status includes non-occurrence of ischemic stroke and occurrence of ischemic stroke; the original data set is obtained through feature engineering processing.

[0010] Furthermore, the data set optimization module specifically includes the following:

[0011] Initialization unit; randomly select nl sample data from the original data set as the initial training set; the initial Gaussian process model for training is expressed as: ; assume that the noise follows ; the distribution of the true label is expressed as: ; where is the input feature of each sample in the initial training set; is the corresponding true label; is the implicit function corresponding to the sample assumed to follow a Gaussian process with a mean of and a covariance function of K; is the noise assumed to follow a normal distribution with a mean of 0 and a variance of ; is the probability distribution of the true label under the given input feature ; I is the identity matrix;

[0012] Active learning selection unit; predict the remaining data set, and calculate the uncertainty of each sample; expressed as: ; ; calculate the rarity , expressed as: ; construct the weighted uncertainty , expressed as: ; where is the prediction output of the Gaussian process model of the remaining data; is the covariance vector between the new sample and the initial training samples; T is the matrix transpose; is the uncertainty; is the number of samples corresponding to the i-th sample category; is the total number of samples in the dataset;

[0013] The dataset construction unit; sort the weighted uncertainties from smallest to largest, select the top z samples with lower weighted uncertainties and add them to the training set; repeat the active learning selection step until the maximum number of iterations is reached; obtain the final dataset.

[0014] Furthermore, in the ischemic stroke prediction model construction module, the construction of the ischemic stroke prediction model is realized based on the multi-layer perceptron model and the final dataset; specifically, it includes the following contents:

[0015] The feature extraction unit; uses a multi-layer perceptron to perform a non-linear transformation on the input features. For the input data X, after the first-layer linear transformation, it is expressed as: ; The second-layer fully connected transformation further linearly transforms the output of the first layer and applies the ReLU activation again, expressed as: ; Introduce a residual connection; when x and have matching dimensions, the output is expressed as: ; When x and do not have matching dimensions, the output is expressed as: ; Among them, and are the outputs of the first-layer transformation and the second-layer transformation respectively; and are the weight matrix and bias of the first layer; and are the weight matrix and bias of the second layer; ReLU(·) is the ReLU activation function; y is the feature extraction output; x is the input feature vector of the patient; and are the weights and biases used to adjust the input dimension matching;

[0016] The prediction output unit; uses a fully connected layer to output the final prediction result, expressed as: ; Among them, is the risk probability of ischemic stroke predicted by the model; sigmoid(·) is the sigmoid function;

[0017] The loss function design unit; specifically includes:

[0018] Construct the initial cross-entropy loss function, expressed as: ; Calculate the gradient, let the predicted probability ; The gradient is expressed as: ; Let ; Introduce the gradient density , expressed as: ; ; Define the dynamic weight coefficient , expressed as: ; where is the initial cross - entropy loss function; Y is the true label; is the gradient of the initial cross - entropy loss function; is the prediction error; and are the prediction errors of the i - th sample and the k - th sample respectively; n is the total number of samples; k is the sample index; is the local kernel function; u is the window width parameter;

[0019] Construct the risk - stratified loss; relabel the data of patients with ischemic stroke that has occurred, and the labels include extremely high risk, high risk, medium risk, and low risk; convert the probability output by the model for each patient into a risk score; expressed as: ; the risk - stratified loss is expressed as: ; where is the risk score of the i - th sample; is the predicted probability of the i - th sample; is the true label of the i - th sample; and are the balance factors for the intermediate risk intervals; output the data label of as extremely high risk; output the data label of as high risk; output the data label of as medium risk; output the data label of as high risk;

[0020] The loss function L of the final ischemic stroke prediction model is expressed as: .

[0021] Further, the ischemic stroke prediction module, based on the established ischemic stroke prediction model, real - time collects patient medical data, inputs it into the ischemic stroke prediction model after feature engineering processing, and takes the model output as the ischemic stroke prediction result.

[0022] The beneficial effects obtained by the present invention using the above - mentioned scheme are as follows:

[0023] (1) Aiming at the problem that the general ischemic stroke prediction system has an imbalance in patient types and fails to handle the noise and uncertainty often existing in clinical data properly, resulting in poor final ischemic stroke prediction effect. This scheme dynamically adjusts the patient data set by actively learning and combining weighted uncertainty to select samples, where the weighting coefficient considers the class distribution of samples and combines the prediction error of the Gaussian process model, based on the uncertainty and rarity of each sample; realizes the efficient selection of the patient data set; and further improves the final ischemic stroke prediction effect.

[0024] (2) Regarding the problem that the general ischemic stroke prediction system has an imbalance in patient risk categories, does not consider the dynamic changes in patient risks, and does not conduct refined management of patients with different risk levels, resulting in poor accuracy in ischemic stroke prediction. This solution uses dynamic weight coefficients, automatically weights according to the gradient density of the samples, and assigns higher weights to high-risk sample patients with ischemic stroke; constructs a risk stratification loss to accurately classify patients with different risk levels; and then realizes the accurate prediction of ischemic stroke. Description of the Drawings

[0025] Figure 1 It is a schematic flowchart of the ischemic stroke prediction system based on big data provided by the present invention;

[0026] Figure 2 It is a schematic flowchart of the ischemic stroke prediction model construction module.

[0027] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. Detailed Embodiments

[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0029] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention.

[0030] Example 1, refer to Figure 1 , the ischemic stroke prediction system based on big data provided by the present invention includes a data acquisition module, a data set optimization module, an ischemic stroke prediction model construction module, and an ischemic stroke prediction module;

[0031] The data acquisition module collects historical patient medical data and obtains the original data set through feature engineering processing;

[0032] The dataset optimization module uses active learning technology to optimize the dataset based on weighted uncertainty;

[0033] The ischemic stroke prediction model construction module designs a dynamic weight coefficient and a risk stratification loss function based on a multi-layer perceptron to construct an ischemic stroke prediction model;

[0034] The ischemic stroke prediction module realizes ischemic stroke prediction for real-time patient medical data based on the established ischemic stroke prediction model.

[0035] Example two, refer to Figure 1 , this example is based on the above example. In the data acquisition module, the historical patient medical data includes clinical signs, laboratory test indicators, lifestyle, and ischemic stroke status; the ischemic stroke status is used as a data label; the ischemic stroke status includes non-occurrence of ischemic stroke and occurrence of ischemic stroke; the clinical signs include blood pressure, heart rate, weight, height, BMI, waist circumference, and body fat percentage; the laboratory test indicators include blood glucose, glycated hemoglobin, blood lipids, inflammatory indicators, renal function indicators, and liver function indicators; the lifestyle includes smoking and drinking conditions, exercise habits, diet structure, and sleep quality; the original dataset is obtained through feature engineering processing.

[0036] Example three, refer to Figure 1 , this example is based on the above example. In the dataset optimization module, the proportions of patients with non-occurrence of ischemic stroke and occurrence of ischemic stroke in the overall data are different, and the active learning is used to optimize the dataset; the specific contents are as follows:

[0037] Initialization unit; randomly select nl sample data from the original dataset as the initial training set; the initial Gaussian process model for training is expressed as: ; assume that the noise follows ; the distribution of the true label is expressed as: ; where, is the input feature of each sample in the initial training set; is the corresponding true label; is the implicit function of the assumed sample that follows a Gaussian process with a mean of and a covariance function of K; is the noise assumed to follow a normal distribution with a mean of 0 and a variance of ; is the probability distribution of the true label under the given input feature ; I is the identity matrix;

[0038] Active learning selection unit; predicts the remaining dataset, calculates the uncertainty of each sample; expressed as: ; ; Calculate rarity , expressed as: ; Construct weighted uncertainty , expressed as: ; Wherein, is the predicted output of the Gaussian process model for the remaining data; is the covariance vector between the new sample and the initial training samples; T is the matrix transpose; is the uncertainty; is the number of samples corresponding to the i-th sample's category; is the total number of samples in the dataset;

[0039] Dataset construction unit; sorts according to the weighted uncertainty from small to large, selects the top z samples with lower weighted uncertainty and adds them to the training set; repeats the active learning selection step until the maximum number of iterations is reached; obtains the final dataset.

[0040] By performing the above operations, for the general ischemic stroke prediction system, there are problems such as unbalanced patient types, improper handling of the noise and uncertainty often existing in clinical data, and thus poor final ischemic stroke prediction effect. This solution selects samples through active learning combined with weighted uncertainty, the weighting coefficient considers the class distribution of the samples, combines the prediction error of the Gaussian process model, and dynamically adjusts the patient dataset based on the uncertainty and rarity of each sample; realizes the efficient selection of the patient dataset; and thus improves the final ischemic stroke prediction effect.

[0041] Example 4, refer to Figure 1 and Figure 2 , this example is based on the above example. In the ischemic stroke prediction model construction module, the prediction of ischemic stroke not only involves simple clinical information, but also complex non-linear relationships; uses a multi-layer perceptron model in deep learning to capture these non-linear features; designs a loss function based on dynamic weight coefficients and risk stratification loss; constructs an ischemic stroke prediction model based on the final dataset; specifically includes the following content:

[0042] Feature extraction unit; performs non-linear transformation on the input features using a multi-layer perceptron. For the input data X, after the first layer of linear transformation, it is expressed as: ; The second layer of fully connected transformation further linearly transforms the output of the first layer and applies the ReLU activation again, expressed as: ; Introduce residual connection; when x and have matching dimensions, the output is expressed as: ; When x and When the dimensions do not match, the output is expressed as: ; where and are the outputs of the first-layer transformation and the second-layer transformation respectively; and are the weight matrix and bias of the first layer; and are the weight matrix and bias of the second layer; ReLU(·) is the ReLU activation function; y is the output of feature extraction; x is the input feature vector of the patient; and are the weights and biases for adjusting the input dimension matching;

[0043] Prediction output unit; The final prediction result is output using a fully connected layer, expressed as: ; where is the risk probability of ischemic stroke predicted by the model; sigmoid(·) is the sigmoid function;

[0044] Loss function design unit; Since the data of ischemic stroke is usually highly imbalanced and the high-risk samples are often fewer; To solve the problem of sample imbalance, a dynamic weight coefficient is designed to make the model assign a greater weight to the high-risk samples during training; Specifically including:

[0045] Construct the initial cross-entropy loss function, expressed as: ; Calculate the gradient, let the predicted probability ; The gradient is expressed as: ; Let ; To measure the distribution of the sample gradient in the overall samples, the gradient density is introduced, expressed as: ; ; Define the dynamic weight coefficient , expressed as: ; where is the initial cross-entropy loss function; Y is the true label; is the gradient of the initial cross-entropy loss function; is the prediction error; and are the prediction errors of the i-th sample and the k-th sample respectively; n is the total number of samples; k is the sample index; is the local kernel function; u is the window width parameter;

[0046] Construct risk-stratified loss; relabel the data of patients with ischemic stroke that has occurred. The labels include extremely high risk, high risk, medium risk, and low risk. Since the risk of ischemic stroke shows a continuous distribution, in order to accurately manage different risk levels, convert the probability output by the model for each patient into a risk score, which is expressed as: ; Risk-stratified loss Expressed as: ; where is the risk score of the i-th sample; is the predicted probability of the i-th sample; is the true label of the i-th sample; and are the balance factors for the intermediate risk intervals; Output the data labels of as extremely high risk; Output the data labels of as high risk; Output the data labels of as medium risk; Output the data labels of as high risk;

[0047] The loss function L of the final ischemic stroke prediction model is expressed as: .

[0048] By performing the above operations, for the general ischemic stroke prediction system, there are problems such as the imbalance of patient risk categories, the failure to consider the dynamic changes of patient risks, and the lack of refined management of patients with different risk levels, which in turn lead to poor accuracy in ischemic stroke prediction. This solution uses dynamic weight coefficients to automatically weight according to the gradient density of samples, giving higher weights to high-risk sample patients with ischemic stroke; constructs risk-stratified loss to accurately classify patients with different risk levels; and then realizes the accurate prediction of ischemic stroke.

[0049] Example 5, refer to Figure 1 , based on the above example, the ischemic stroke prediction module, based on the established ischemic stroke prediction model, real-time collects patient medical data, and after feature engineering processing, inputs it into the ischemic stroke prediction model, and takes the model output as the ischemic stroke prediction result; for the model output of patients with ischemic stroke that has occurred and is not low risk, issue a warning to relevant personnel.

[0050] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.

[0051] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention.

[0052] The above describes the present invention and its implementation manners, and this description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. In general, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.

Claims

1. An ischemic stroke prediction system based on big data, characterized in that: The system includes a data acquisition module, a dataset optimization module, an ischemic stroke prediction model construction module, and an ischemic stroke prediction module; The data acquisition module collects historical patient medical data and obtains an original dataset through feature engineering processing; The dataset optimization module uses active learning technology to optimize the dataset based on weighted uncertainty; The ischemic stroke prediction model construction module designs a dynamic weight coefficient and a risk stratification loss function based on a multi-layer perceptron to construct an ischemic stroke prediction model; The ischemic stroke prediction module realizes ischemic stroke prediction for real-time patient medical data based on the established ischemic stroke prediction model.

2. The ischemic stroke prediction system based on big data according to claim 1, wherein: The dataset optimization module specifically includes the following: Initialization unit; randomly select nl sample data from the original dataset as the initial training set; used to train the initial Gaussian process model, expressed as: ; assume the noise follows ; the distribution of the true labels is expressed as: ; where, is the input feature of each sample in the initial training set; is the corresponding true label; is to assume that the implicit function corresponding to the sample follows a Gaussian process with a mean of and a covariance function of K; is the noise assumed to follow a normal distribution with a mean of 0 and a variance of ; is the probability distribution of the true label given the input feature ; I is the identity matrix; Active learning selection unit; make predictions on the remaining dataset and calculate the uncertainty of each sample; expressed as: ; ; Calculate the rarity , expressed as: ; Construct the weighted uncertainty , expressed as: ; where is the predicted output of the Gaussian process model for the remaining data; is the covariance vector between the new sample and the initial training samples; T is the matrix transpose; is the uncertainty; is the number of samples corresponding to the i-th sample's category; is the total number of samples in the dataset; A dataset construction unit; sorting according to weighted uncertainty from small to large, selecting the top z samples with lower weighted uncertainty and adding them to the training set; repeating the active learning selection step until the maximum number of iterations is reached; obtaining the final dataset.

3. The ischemic stroke prediction system based on big data according to claim 2, wherein: In the ischemic stroke prediction model construction module, the construction of the ischemic stroke prediction model is realized based on the multi-layer perceptron model and the final dataset; specifically Includes the following: Feature extraction unit; uses a multi-layer perceptron to perform non-linear transformation on the input features. For the input data X, after the first layer of linear transformation, it is represented as: ; The second layer of fully connected transformation further linearly transforms the output of the first layer and applies the ReLU activation again, represented as: ; Introduce residual connection; when x and have matching dimensions, the output is represented as: ; When x and do not have matching dimensions, the output is represented as: ; where, and are the outputs of the first layer transformation and the second layer transformation respectively; and are the weight matrix and bias of the first layer; and are the weight matrix and bias of the second layer; ReLU(·) is the ReLU activation function; y is the feature extraction output; x is the input feature vector of the patient; and are the weights and biases used to adjust the input dimension matching; Prediction output unit; the final prediction result is output using a fully connected layer, expressed as: ; where is the risk probability of ischemic stroke predicted by the model; sigmoid(·) is the sigmoid function; A loss function design unit; specifically including: Construct the initial cross - entropy loss function, denoted as: ; Calculate the gradient. Let the predicted probability be ; The gradient is denoted as: ; Let ; Introduce the gradient density , denoted as: ; ; Define the dynamic weight coefficient , denoted as: ; Among them, is the initial cross - entropy loss function; Y is the true label; is the gradient of the initial cross - entropy loss function; is the prediction error; and are the prediction errors of the i - th sample and the k - th sample respectively; n is the total number of samples; k is the sample index; is the local kernel function; u is the window width parameter; Construct risk-stratified loss; relabel the data of patients who have suffered ischemic stroke, and the labels include extremely high risk, high risk, medium risk, and low risk; convert the probability output by the model for each patient into a risk score; expressed as: ; Risk-stratified loss Expressed as: ; where is the risk score of the i-th sample; is the predicted probability of the i-th sample; is the true label of the i-th sample; and are the balance factors for the intermediate risk intervals; output the data label of as extremely high risk; output the data label of as high risk; output the data label of as medium risk; output the data label of as high risk; The loss function L of the final ischemic stroke prediction model is expressed as: .

4. The ischemic stroke prediction system based on big data according to claim 3, characterized in that: In the data acquisition module, the historical patient medical data includes clinical signs, laboratory test indicators, lifestyle, and ischemic stroke status; the ischemic stroke status is used as the data label; the ischemic stroke status includes not having had an ischemic stroke and having had an ischemic stroke; An original dataset is obtained through feature engineering processing.

5. The ischemic stroke prediction system based on big data according to claim 4, wherein: The ischemic stroke prediction module is based on the established ischemic stroke prediction model, collects patient medical data in real time, inputs it into the ischemic stroke prediction model after feature engineering processing, and uses the model output as the ischemic stroke prediction result.

Citation Information

Patent Citations

  • Cerebral apoplexy risk prediction method and device based on hybrid deep transfer learning

    CN111968746A

  • Neural network model training method and device for pathological image sample

    CN112232407A

  • Machine learning implementation for multi-analyte assay of biological samples

    CN112292697A

  • Pulmonary nodule risk assessment system

    CN112382392A

  • Active learning method and system based on adversarial training enhancement

    CN116187400A