A blood usage prediction method and prediction system for a blood center

By adopting new loss function and neural network structure in the blood volume prediction model, the implicit relationship between the patient's chain indicators and blood consumption is constructed, and the existing model's shortcomings in accuracy and universality are solved, and blood volume prediction with higher robustness and generalization ability is achieved.

CN115775618BActive Publication Date: 2025-05-16WESTLAKE UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211534022.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-01
Publication Date
2025-05-16
Estimated Expiration
2042-12-01

AI Technical Summary

Technical Problem

The existing blood volume prediction model has insufficient accuracy and universality, low data quality, and complex cooperative relationships between hospitals, resulting in low generalization of the model and unavailable to be widely used.

Method used

Using a new loss function based on neural network model, the cross-domain prediction capability of neural network models is improved by constructing the implicit relationship between patients' chain indicators and blood consumption. The specific steps include collecting the hospital's historical clinical data, normalizing and enhancing the data, building a latent space projection network and output network, and training using a weighted loss function of cross entropy loss and similarity loss.

Benefits of technology

The robustness and generalization ability of the blood volume prediction model are improved, the domain deviation problem is overcome, and more accurate and universal blood volume prediction is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115775618B_ABST
    Figure CN115775618B_ABST
Patent Text Reader

Abstract

The present invention discloses a blood usage prediction method and prediction system for a blood center. The method mainly constructs a blood usage prediction neural network, and designs a weighted loss of the cross entropy loss and the loss of similarity between the original data space and the latent space as the overall loss function in the blood usage prediction neural network training process. This can solve the problem that the self-supervisory signal of the hidden layer will alleviate overfitting, domain migration and other problems. The present invention can improve the cross-domain prediction ability of the neural network model, so that the prediction method of the present invention is more robust and the model generalization ability is stronger.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of blood bank scheduling, and in particular to a blood usage prediction method and prediction system for a blood center. Background Art

[0002] Predictive modeling of patient blood usage is a meaningful and challenging topic. Current blood usage prediction models are not perfect. Currently, most studies in the field of blood usage prediction are based on univariate time series methods (autoregressive models). In these studies, predictions are based only on previous demand values ​​without considering other characteristics that may affect demand.

[0003] For example, Wang et al. developed an early transfusion scoring system to predict the blood needs of patients with severe trauma in the pre-hospital or early emergency rescue. However, the above design is more suitable for triage in a single hospital rather than a blood center, because the system does not consider how to avoid the performance degradation caused by domain bias between different hospitals. Rebecca et al. summarized the existing research on how to predict whether trauma patients need massive transfusions, and listed the blood consumption scoring system (ABC) and the shock index scoring system (SI). The above systems use classical or ML-based methods to predict the blood consumption during patient treatment. Unfortunately, although the progress is significant, it is not satisfactory in terms of accuracy and universality, so the blood demand prediction model cannot be widely used. We summarize these problems as follows. The data quality is not high. The cooperative relationship between blood centers and hospitals makes it very difficult to establish a strict patient information feedback system. In addition, the bias and data missing caused by differences in hospital equipment pose challenges to the constant blood usage prediction model. The generalization of the model is low. The blood usage prediction model is often established for a specific hospital or city, without considering expansion to a wider range of applications, so the generalization performance is poor. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention proposes a blood usage prediction method and system for a blood center. The method constructs a new loss function based on a neural network model to construct an implicit relationship between the patient's linkage indicators and blood consumption, thereby improving the cross-domain prediction capability of the neural network model, making the prediction method of the present invention more robust.

[0005] The purpose of the present invention is achieved through the following technical solutions:

[0006] A method for predicting blood usage in a blood center, the method comprising the following steps:

[0007] Step 1: Collect the historical clinical data of patients from various hospitals, and normalize the actual blood usage of the patients so that the minimum blood usage is 0 and the maximum blood usage is 1, and use the normalized blood usage as the label of data x;

[0008] Step 2: Select k neighboring data of hospitals other than the hospital where data x is located, perform linear interpolation, and obtain enhanced data of data x; Enhance data x ′ The labels are the same as the corresponding original data; the historical clinical data and its enhanced data together constitute the training data set;

[0009] Step 3: construct a blood usage prediction neural network, which includes a latent space projection network and an output network. Both the latent space projection network and the output network are multi-layer perception networks, each including an input layer, multiple hidden layers and an output layer, and the output of the latent space projection network is used as the input of the output network; the latent space projection network is used for preprocessing to combat the adverse effects of noise and missing values ​​of input data; the output network is used for dimensionality reduction;

[0010] Step 4: constructing a weighted loss consisting of a cross entropy loss and a loss of similarity between the original data space and the latent space as an overall loss function in the training process of the blood usage prediction neural network, and using the training data set obtained in step 2 to train the blood usage prediction neural network;

[0011] Step 5: Input the clinical data of the patient to be predicted into the trained blood usage prediction neural network to obtain the patient's predicted blood usage.

[0012] Furthermore, the overall loss function is:

[0013] L=L D +βL CE

[0014] Among them, L D is the loss of similarity between the original data space and the latent space, L CE is the cross entropy loss, β is a trade-off L D and L CE Hyperparameters of L D The calculation formula is as follows:

[0015]

[0016]

[0017] Among them, L(x i , x j ) is the data x sampled from the training data set i and xj The similarity loss between j =τ(x i ) is to judge x j Is it x i The enhanced data; B is the number of samples selected for one training; v y , v z are the degrees of freedom of two spaces, i.e., spatial dimensions; Calculated by the following formula:

[0018] D(p, q)=p log q+(1-p)log(1-q)

[0019]

[0020] Where D(p, q) is the fuzzy inter-set cross entropy loss, p∈[0,1]; Gamma(·) is the Gamma function; v is the degree of freedom parameter that controls the shape of the distribution; d is the Euclidean distance between any two data point pairs; κ(d, v) is the manifold similarity between any two data point pairs;

[0021] L CE is the cross entropy loss:

[0022]

[0023] Among them, l i is the label of data point i, σ(z i ) is the output of the output network.

[0024] Furthermore, in step 2, the calculation formula for obtaining the enhanced data x′ from the original data x is as follows:

[0025] x′=T(x)={τ(x i,1 ), ..., τ(x i,f ), τ(x i,n )}

[0026]

[0027] Among them, the various characteristics of x'τ(x i,f ) is the original feature x collected from various hospitals i,f Obtained through linear interpolation; r is the linear combination coefficient, which is sampled from the uniform distribution U(0, p), and p is a hyperparameter; is a set of neighbor nodes sampled from hospitals different from x.

[0028] Furthermore, the latent space projection network and the output network each include 3 hidden layers.

[0029] A blood usage prediction system for implementing a blood usage prediction method, the system comprising the following modules:

[0030] Database module: used to store training data sets;

[0031] Feature enhancement module: used to use the linear difference method to select k neighboring data of hospitals other than the hospital where the data x is located, perform linear interpolation, and output enhanced data of data x;

[0032] Blood usage normalization module: used to normalize the patient's historical blood usage so that the minimum blood usage is 0 and the maximum blood usage is 1, to facilitate the training of cross entropy loss;

[0033] Loss function module: used to calculate the loss of similarity between the original data space and the latent space, as well as the cross entropy loss.

[0034] Blood usage prediction module: It has a built-in blood usage prediction neural network, receives input of historical clinical data and its enhanced data, and outputs a prediction of the patient's blood usage.

[0035] The beneficial effects of the present invention are as follows:

[0036] The present invention is based on a new loss function of the neural network model to construct an implicit relationship between the patient's linkage indicators and blood consumption, thereby improving the cross-domain prediction ability of the neural network model, making the prediction method of the present invention more robust and the model more generalizable. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 The present invention is a flowchart of a method for predicting blood usage in a blood center according to an embodiment of the present invention.

[0038] Figure 2 FIG. 1 is a schematic diagram of a blood usage prediction neural network according to an embodiment of the present invention.

[0039] Figure 3 Schematic diagram of the domain bias problem.

[0040] Figure 4 It is the ROC curve diagram. DETAILED DESCRIPTION

[0041] The present invention will be described in detail below based on the accompanying drawings and preferred embodiments, and the purpose and effects of the present invention will become more clear. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] like Figure 1 As shown, the blood consumption prediction method for a blood center of the present invention comprises the following steps:

[0043] Step 1: Collect the historical clinical data of patients from various hospitals, and normalize the actual blood usage of the patients so that the minimum blood usage is 0 and the maximum blood usage is 1, and use the normalized blood usage as the label of data x;

[0044] Step 2: Select k neighboring data of hospitals other than the hospital where data x is located, perform linear interpolation, and obtain enhanced data of data x; the label of enhanced data x′ is the same as its corresponding original data; the historical clinical data and its enhanced data together constitute a training data set;

[0045] The data enhancement here uses the MIXUP enhancement method. The MIXUP enhancement method is widely used in the training process of neural networks. Most data enhancement methods fill the original sample space based on some artificially summarized data distribution characteristics, which can effectively reduce the overfitting of the model and improve the generalization performance of the model. Secondly, the data enhancement method essentially strengthens the basic assumption of DR, that is, the local connectivity of the neighborhood. It learns the detailed data distribution by generating more data in the manifold based on the sampling points. The data enhancement of the present invention generates new data points x′=T(x) from the original data x under a unified framework.

[0046] X′=T(x)={τ(x i,1 ), ..., τ(x i,f ), τ(x i,n )}

[0047]

[0048] in, That is x a The features of are the original features x collected from various hospitals i,f Obtained by linear interpolation; r is the linear combination coefficient, which is sampled from the uniform distribution U(0, p), and p is a hyperparameter; is sampled from the neighborhood N(x) of x, is the k neighborhood of data point i in h hospitals, H is the hospital set, H / h i This means deleting the hospital that contains the neighborhood of data point i.

[0049] Cross-hospital data augmentation generates new data by combining data point i with its neighborhood in different hospitals. It reduces the impact of missing data and increases the divergence of training data. Therefore, the model learns an accurate data distribution, which can improve the performance of the method. In addition, it works together with the loss function to align the data offsets from different hospitals, thereby overcoming domain bias.

[0050] Step 3: construct a blood usage prediction neural network, which includes a latent space projection network and an output network. Both the latent space projection network and the output network are multi-layer perception networks, each including an input layer, multiple hidden layers and an output layer, and the output of the latent space projection network is used as the input of the output network; the latent space projection network is used for preprocessing to combat the adverse effects of noise and missing values ​​of input data; the output network is used for dimensionality reduction;

[0051] The general neural network model will directly use the network output to calculate the supervised loss function. This may lead to some undesirable consequences (such as overfitting). The invention proposes a manifold learning regularizer to use the information of data in the latent space (i.e., the low-dimensional feature space, the manifold hypothesis assumes that high-dimensional data actually has a low-dimensional representation that can describe its characteristics) to suppress overfitting and other problems. The invention converts the overall network F(·, w i , w j ) is divided into a latent space projection network f i (·, w i ) and the output network f j (·, w j ) two parts, the specific forward propagation process is as follows:

[0052] (1) The enhanced data x′=T(x) is projected through the latent space network to obtain the output y=f of the latent space. i (x′,w i ).

[0053] (2) After y passes through the output network, the output z = f j (y,w j ).

[0054] All experiments use a fixed MLP network structure, defined as a 10-layer structure. The first five layers are latent space projection networks and the last five layers are output networks. The number of epochs is 400. The batch size is 2048.

[0055] Step 4: Construct a weighted loss consisting of the cross entropy loss and the loss of similarity between the original data space and the latent space as the overall loss function in the training process of the blood usage prediction neural network, and use the training data set obtained in step 2 to train the blood usage prediction neural network.

[0056] Among them, the cross entropy loss utilizes the label information, and the loss of similarity between the original data space and the latent space is called the "Manifold Regularize (Regularizer in the original text here)" loss.

[0057] The following focuses on the "manifold regularization" loss. The "manifold regularization" loss handles the domain bias of different hospitals during the training phase and provides manifold constraints to prevent overfitting. The inconsistency of medical equipment in different hospitals and the subjective diagnosis of doctors lead to domain bias in the data between hospitals. The "manifold regularization" loss guides the mapping of the neural network to be insensitive to hospitals, thereby overcoming the domain bias, such as Figure 2 As shown. Therefore, we first define the pairwise similarity between nodes. Considering the inconsistent dimensions of the latent space, we use the t distribution with variable degrees of freedom as the kernel function to measure the point-to-point similarity of the data. First, we define the manifold similarity between two pairs of data points:

[0058]

[0059] Where Gamma(.) is the Gamma function, v is the degree of freedom parameter that controls the shape of the distribution, and d is the Euclidean distance between any two pairs of data points. Based on the similarity of data points in this unified paradigm in different dimensions, we also need to introduce the "fuzzy inter-set cross entropy loss".

[0060] D(p, q)=p logq+(1-p)log(1-q)

[0061] Where p∈[0,1]. It is not difficult to find that the above D(p,q) is actually a continuous version of the cross entropy.

[0062] Now define the loss of similarity between the original data space and the latent space. This loss will guide the network to narrow the structural distance between the two spaces (that is, the similarity between all pairs of data points):

[0063]

[0064]

[0065] Where B is the batch size. j =τ(x i ) is to judge x j Is it x i If x j is x i For the enhanced data, the loss function will shorten their distance in the latent space; if it is not a loss function, it will try to keep their distance in the original space to protect the manifold structure information of the data from being destroyed. y , v z These are the degrees of freedom (that is, dimensions) of two spaces.

[0066] Figure 3It shows how the "manifold regularization" loss works. Without the "manifold regularization" loss, data from different hospitals will be clustered into one category in the latent space (due to a specific domain bias). However, after the "manifold regularization" loss, the model will intentionally ignore the domain bias and focus more on the structural information of the data.

[0067] Next, we will introduce another loss of the model: cross entropy loss. Cross entropy loss uses the label information of the data.

[0068]

[0069] Among them l i is the label of data point i, σ(z i ) is the output of the entire network. This loss supervises the alignment of the output of the data after passing through the network with its corresponding label.

[0070] The overall loss function is:

[0071] L=L D +βL CE

[0072] Where β is a trade-off L D and L CE Hyperparameters of .

[0073] Step 5: Input the clinical data of the patient to be predicted into the trained blood usage prediction neural network to obtain the patient's predicted blood usage.

[0074] According to another aspect of the present invention, a blood usage prediction system for a blood center is provided, the system comprising the following modules:

[0075] Database module: used to store training data sets;

[0076] Feature enhancement module: used to use the linear difference method to select k neighboring data of hospitals other than the hospital where the data x is located, perform linear interpolation, and output enhanced data of data x;

[0077] Blood usage normalization module: used to normalize the patient's historical blood usage so that the minimum blood usage is 0 and the maximum blood usage is 1, to facilitate the training of cross entropy loss;

[0078] Loss function module: used to calculate the loss of similarity between the original data space and the latent space, as well as the cross entropy loss. Blood usage prediction module: built-in blood usage prediction neural network, receiving input historical clinical data and its enhanced data, and outputting a prediction of the patient's blood usage.

[0079] A specific application example of the method and system of the present invention is given below.

[0080] 1. Collect original data and collect data in actual conditions. The data in the database of a provincial blood center and 12 hospitals include data of 2025 patients, of which 1970 data are from emergency trauma patients in 10 hospitals in a province from 2018 to 2020. Each data set mainly includes the following parts. (1) General information of patients, including case number, consultation time, pre-test classification, injury time, gender, age, weight, diagnosis, penetrating injury, heart rate, diastolic blood pressure, systolic blood pressure, body temperature, shock index, and Glasgow Coma Scale. (2) Injury status, including pleural effusion, peritoneal effusion, limb injury status, thoracic and abdominal injury status, and pelvic injury status. (3) Laboratory tests: hemoglobin, hematocrit, albumin, hemoglobin value 24 hours after transfusion therapy, base residual, pH value (acidity), base deficiency, oxygen saturation, PT (prothrombin time), and APTT (partial thromboplastin time). (4) Burn conditions, including burn area and burn depth. (5) Patient recovery status, including whether blood was used: whether suspended red blood cells were transfused, blood transfusion status and amount of blood transfusion on the first day of hospitalization.

[0081] 2. The original dataset generates an enhanced dataset through the data augmentation method introduced above.

[0082] 3. Initialize the network: Use Kaiming initializer to initialize other network parameters. This example uses AdamW optimizer, learning rate is 0.02, weight decay is 0.5. All experiments use a fixed MLP network structure, defined as a 10-layer structure. The first five layers are the latent space projection network and the last five layers are the output network. The number of epochs is 400. The batch size is 2048.

[0083] 4. Training Network: Latent Space Projection Network i (·,w i ) will project the enhanced dataset into the latent space to obtain a feature dataset. At this time, the "manifold regularization" loss function D A structural loss is calculated using the feature dataset and the enhanced dataset to guide the network to overcome the impact of domain shift.

[0084] 5. Use the network: input test data, output network i (·,w i ) will project the feature dataset into the output space to obtain the final prediction value. At this time, the cross entropy loss function will use the output and label information to calculate a prediction error loss to guide the network to align input and output.

[0085] Although the blood data of each patient is collected for the accurate regression subtask, it is equally important to predict whether the patient needs a blood transfusion. Especially in emergency situations, it is more important to indicate whether the patient needs a blood transfusion in the short term than to predict the amount of blood used in the entire treatment cycle. Therefore, the performance of the BUPNN of the present invention on the classification subtask is first evaluated. Then, two evaluation indicators, namely, classification accuracy (ACC) and area under the ROC curve (AUC), are used to compare the performance of the classifier from multiple perspectives.

[0086] Figure 4 For ROC curve comparison, missing values ​​are filled with median values. The closer the curve is to the upper left corner, the better the performance of the model. The symmetry of the curve along the line from (0,1) to (1,0) indicates the balanced performance of the model. It can be seen that the performance of the method curve of the present invention is higher than all existing methods.

[0087] Table 1 shows the comparison of Midiean classification accuracy with baseline methods for all data from different hospitals, with the best result shown in blood. The second result is underlined. The brackets on the right end show how much BUPNN exceeds the best indicator among the other methods.

[0088] Table 1 shows the comparison of Midiean classification accuracy and baseline method for all data from different hospitals

[0089]

[0090] Table 2 shows the MSE comparison of different methods on all data sets

[0091]

[0092] Table 2 is a comparison of the MSE of different methods under all data sets, with the best result in bold and the second-ranked result underlined, and the brackets on the right side show how much BUPNN exceeds. The data in Tables 1 and 2 both show that the prediction effect of the BUPNN model of the present invention is much better than other existing methods.

[0093] Those skilled in the art can understand that the above are only preferred examples of the invention and are not intended to limit the invention. Although the invention is described in detail with reference to the above examples, those skilled in the art can still modify the technical solutions recorded in the above examples or replace some of the technical features therein with equivalents. Any modification, equivalent replacement, etc. made within the spirit and principle of the invention shall be included in the protection scope of the invention.

Claims

1. A method for predicting blood usage in a blood center, characterized in that: The method comprises the following steps: Step 1: Collect the historical clinical data of patients from various hospitals, and normalize the actual blood usage of the patients so that the minimum blood usage is 0 and the maximum blood usage is 1, and use the normalized blood usage as the label of data x; Step 2: Select k neighboring data of hospitals other than the hospital where data x is located, perform linear interpolation, and obtain enhanced data of data x; the label of enhanced data x′ is the same as its corresponding original data; the historical clinical data and its enhanced data together constitute a training data set; Step 3: construct a blood usage prediction neural network, which includes a latent space projection network and an output network. Both the latent space projection network and the output network are multi-layer perception networks, each including an input layer, multiple hidden layers and an output layer, and the output of the latent space projection network is used as the input of the output network; the latent space projection network is used for preprocessing to combat the adverse effects of noise and missing values ​​of input data; the output network is used for dimensionality reduction; Step 4: constructing a weighted loss consisting of a cross entropy loss and a loss of similarity between the original data space and the latent space as an overall loss function in the training process of the blood usage prediction neural network, and using the training data set obtained in step 2 to train the blood usage prediction neural network; Step 5: Input the clinical data of the patient to be predicted into the trained blood usage prediction neural network to obtain the predicted blood usage of the patient; The overall loss function is: L=L D +βL CE Among them, L D is the loss of similarity between the original data space and the latent space, L CE is the cross entropy loss, β is a trade-off L D and L CE Hyperparameters of L D The calculation formula is as follows: Among them, L(x i ,x j ) is the data x sampled from the training data set i and x j The similarity loss between j =τ(x i ) is to judge x j Is it x i The enhanced data; B is the number of samples selected for one training; v y ,v z are the degrees of freedom of two spaces, i.e., spatial dimensions; Calculated by the following formula: D(p,q)=plogq+(1-p)log(1-q) Where D(p,q) is the fuzzy inter-set cross entropy loss, p∈[0,1]; Gamma(·) is the Gamma function; v is the degree of freedom parameter that controls the shape of the distribution; d is the Euclidean distance between any two data point pairs; κ(d,v) is the manifold similarity between any two data point pairs; L CE is the cross entropy loss: Among them, l i is the label of data point i, σ(z i ) is the output of the output network.

2. The blood consumption prediction method for a blood center according to claim 1, characterized in that: In step 2, the calculation formula for obtaining the enhanced data x′ from the original data x is as follows: x′=T(x)={τ(x i,1 ),…,τ(x i,f ),τ(x i,n )} Among them, x ‘ The various characteristics of τ(x i,f ) is the original feature x collected from various hospitals i,f Obtained through linear interpolation; r is the linear combination coefficient, which is sampled from the uniform distribution U(0,p), and p is a hyperparameter; is a set of neighbor nodes sampled from hospitals different from x.

3. The blood usage prediction method for a blood center according to claim 1, characterized in that: The latent space projection network and the output network both include 3 hidden layers.

4. A blood usage prediction system for implementing the blood usage prediction method according to claim 1, characterized in that: The system includes the following modules: Database module: used to store training data sets; Feature enhancement module: used to use the linear difference method to select k neighboring data of hospitals other than the hospital where the data x is located, perform linear interpolation, and output enhanced data of data x; Blood usage normalization module: used to normalize the patient's historical blood usage so that the minimum blood usage is 0 and the maximum blood usage is 1, to facilitate the training of cross entropy loss; Loss function module: used to calculate the loss of similarity between the original data space and the latent space, as well as the cross entropy loss. Blood usage prediction module: It has a built-in blood usage prediction neural network, receives input of historical clinical data and its enhanced data, and outputs a prediction of the patient's blood usage.

Citation Information

Patent Citations

  • Arrival angle estimation method based on adversarial regularization deep neural network

    CN111767791A

  • Deep learning-based blood inventory early warning method and system

    CN114171173A