Hepatitis B treatment time prediction method and device based on neural network model
Through a neural network model-based method, the recurrent neural network-long and short-term memory network is used to process data of chronic hepatitis B patients, combined with dynamic risk adjustment factors, the problem of insufficient accuracy in the treatment time of traditional linear regression prediction methods is solved, and more accurate prediction of treatment time and judgment of cure possibility is achieved.
Patent Information
- Application Number
- CN202510423229.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-11
AI Technical Summary
The existing linear regression prediction methods cannot accurately predict the treatment effect of patients with chronic hepatitis B, especially because the data relationship is complex and the nonlinear relationship is difficult to fit, resulting in insufficient accuracy in the prediction of hepatitis B treatment time.
Using a neural network model-based method, the target model is generated by obtaining multi-faceted information and related sample sets of target users, and the initial recurrent neural network-long and short-term memory network model is used to process the target model. Combined with dynamic risk adjustment factors, the treatment time of hepatitis B is predicted, and the SA-Golden standard is used to assist in judging the possibility of cure.
It improves the accuracy of prediction of hepatitis B treatment time, can understand the possibility of cure after medication, helps patients and doctors make better treatment decisions, arrange the R&D process reasonably, and reduce costs.
Smart Images

Figure CN120299731A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a method and device for predicting the hepatitis B treatment time based on a neural network model. Background Art
[0002] Hepatitis B virus (HBV) infection is a severe global public health problem. A large number of these infected individuals urgently need effective drug treatment. The clinical treatment goals of chronic hepatitis B (CHB) include long-term inhibition of HBV replication, reduction of hepatic cell inflammatory necrosis and hepatic fibrosis hyperplasia, delay and reduction of the occurrence of complications such as liver failure and decompensated cirrhosis, so as to extend the survival time of patients. Based on the disease characteristics of CHB and the action mechanism of new drugs, the determination of the drug trial period needs to comprehensively consider efficacy and safety. For new drugs for long-term treatment to inhibit the virus, the confirmatory clinical trial of efficacy requires continuous administration for at least 48 weeks, and it is recommended to observe for a longer time. For treatment drugs with a limited treatment course, a longer trial period is required.
[0003] During the process of drug research and development and clinical treatment, pharmaceutical researchers hope to predict the clinical trial results in advance, which helps to reasonably arrange the research and development process and reduce costs. For patients, understanding in advance whether they can be clinically cured after taking medicine for a certain period of time is of great significance for their treatment decisions and life planning. However, there is currently a lack of effective methods to meet this need. In recent years, although there have been many attempts to combine HBV-related research with machine learning, the traditional linear regression prediction method has obvious deficiencies. Traditional linear regression uses regression analysis in mathematical statistics to determine the quantitative relationship between variables. Although it has a fast modeling speed and strong interpretability, its assumption that the relationship between variables is linear cannot well fit non-linear data. When dealing with data related to the treatment of chronic hepatitis B patients, due to the complex relationship between the data, which is not a simple linear relationship, the traditional linear regression model is difficult to accurately predict the treatment effect of patients.
[0004] It should be noted that the information disclosed in the above background art section is only used to strengthen the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The purpose of this application is to provide a prediction method and device for the hepatitis B treatment time based on a neural network model, which at least overcome the problems existing in the prior art to a certain extent. By obtaining multi-faceted information of the target user, the relevant sample set, and the initial model, and through processing the initial recurrent neural network-long short-term memory network model to generate the target model, steps such as data feature sampling are involved. Then, the data of the target user is processed to obtain the target access value and the dynamic risk adjustment factor, and both are input into the target model to obtain the access prediction value. Then, the HBsAg value and the Alt value are extracted from the prediction value, and the attribute classification information is generated through preset threshold processing to judge the cure type, and the cure time node is determined in combination with the preset HBsAg threshold. The SA-Golden standard (HBsAg value / Alt value) is used to assist in judging the cure possibility of the patient.
[0006] Other features and advantages of this application will become apparent through the following detailed description, or will be learned in part through the practice of the present invention.
[0007] According to one aspect of this application, a prediction method for the hepatitis B treatment time based on a neural network model is provided, including: obtaining the clinical parameter information of the target user, the living habit information of the target user, the different frequency access values of the target user within a preset time period, the initial recurrent neural network-long short-term memory network model, the training sample set, and the validation sample set, where the training sample set is used to represent users who have received different treatment cycles of PEG-IFNα-2b; processing the initial recurrent neural network-long short-term memory network model based on the training sample set and the validation sample set to generate the target recurrent neural network-long short-term memory network model; performing data cleaning and missing value supplementation processing on the clinical parameter information of the target user and the different frequency access values of the target user within a preset time period to generate the target access value of the target user within a preset time period, where the target access value of the target user within a preset time period includes at least two of HBsAg, HBeAg, HBVDNA, HBsAb, HBeAb, and Alt; processing the living habit information of the target user to generate a dynamic risk adjustment factor; processing the target access value of the target user within a preset time period and the dynamic risk adjustment factor based on the target recurrent neural network-long short-term memory network model to generate the access prediction value of the target user within the target time period; processing the access prediction value of the target user within the target time period to generate the drug cure prediction time information of the target user.
[0008] Another aspect of the present application is a prediction device for the treatment time of hepatitis B based on a neural network model, which is characterized by including: an acquisition module, configured to acquire the clinical parameter information of the target user, the living habit information of the target user, the different frequency access values of the target user within a preset time period, an initial recurrent neural network-long short-term memory network model, a training sample set, and a verification sample set, wherein the training sample set is used to represent users who have received different treatment cycles of PEG-IFNα-2b; a processing module, configured to process the initial recurrent neural network-long short-term memory network model based on the training sample set and the verification sample set to generate a target recurrent neural network-long short-term memory network model; perform data cleaning and missing value supplementation processing on the clinical parameter information of the target user and the different frequency access values of the target user within a preset time period to generate the target access values of the target user within the preset time period, wherein the target access values of the target user within the preset time period include at least two of HBsAg, HBeAg, HBVDNA, HBsAb, HBeAb, and Alt; process the living habit information of the target user to generate a dynamic risk adjustment factor; process the target access values and the dynamic risk adjustment factor of the target user within the preset time period based on the target recurrent neural network-long short-term memory network model to generate the access prediction values of the target user within the target time period; process the access prediction values of the target user within the target time period to generate the drug cure prediction time information of the target user.
[0009] According to yet another aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a second processor, it implements the above-mentioned prediction method for the treatment time of hepatitis B based on a neural network model.
[0010] The present application provides a prediction method and device for the treatment time of hepatitis B based on a neural network model. The present application introduces a construction and application process of a neural network model for predicting the efficacy of PEG-IFNα-2b in the treatment of chronic hepatitis B. First, multi-faceted information of the target user, relevant sample sets, and an initial model are obtained. By processing the initial recurrent neural network-long short-term memory network model, a target model is generated, which involves steps such as data feature sampling. Then, the target user data is processed to obtain the target access values and the dynamic risk adjustment factor, and the two are input into the target model to obtain the access prediction values. Then, the HBsAg value and the Alt value are extracted from the prediction values, and after being processed by a preset threshold, the attribute classification information is generated to judge the cure type, and the cure time node is determined by combining the preset HBsAg threshold. The SA-Golden standard (HBsAg value / Alt value) is used to assist in judging the cure possibility of the patient.
[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings
[0012] Figure 1 The flowchart shows a prediction method for the hepatitis B treatment time provided by an embodiment of the present application based on a neural network model; Figure 2 The structural schematic diagram shows a prediction device for the hepatitis B treatment time provided by an embodiment of the present application based on a neural network model. Detailed Embodiments
[0013] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0014] The following combines with Figure 1 to describe the prediction method for the hepatitis B treatment time based on a neural network model according to an exemplary embodiment of the present application. In one embodiment, the present application also proposes a prediction method and device for the hepatitis B treatment time based on a neural network model. Figure 1 The flowchart schematically shows a prediction method for the hepatitis B treatment time based on a neural network model according to an embodiment of the present application. As Figure 1 shown, this method is applied to a server and includes: S101, obtaining the clinical parameter information of the target user, the living habit information of the target user, the different frequency access values of the target user within a preset time period, an initial recurrent neural network-long short-term memory network model, a training sample set, and a validation sample set.
[0015] In one implementation, clinical parameter information covers multiple aspects and is crucial for understanding the user's condition and physical status. The user's age and gender are recorded because age and gender affect drug metabolism and treatment response. Research has found that there are differences in drug tolerance and treatment effects among users of different age groups, and gender differences also lead to different immune responses, thus affecting the treatment outcome. At the same time, six important indicator parameters (HBsAg, HBeAg, HBVDNA, HBsAb, HBeAb, Alt) during the visits of chronic hepatitis B users are key points for inclusion. These indicators directly reflect the hepatitis B virus infection status and liver function and are the key basis for evaluating the treatment effect. Taking age as an example, for older users, the liver's metabolism and repair functions are relatively weak, the drug metabolism rate is slower, and it takes longer to achieve the ideal treatment effect during the treatment process. In terms of gender, during special periods such as the physiological cycle and pregnancy, changes in hormone levels in the female body affect the immune system, resulting in different responses to drugs compared to men. When including the indicator parameters, HBsAg is an important marker of hepatitis B virus infection, and changes in its level directly reflect the active degree of virus replication in the body; HBeAg is closely related to the virus's infectivity; HBVDNA quantification can accurately measure the virus content in the blood; HBsAb is a protective antibody, and an increase in its level means that the body has a certain resistance to the hepatitis B virus; the appearance of HBeAb usually indicates that virus replication is inhibited; Alt is a sensitive indicator reflecting liver cell damage, and an increase in its value indicates inflammation or necrosis of liver cells. These indicators are interrelated, and comprehensive analysis can provide a more comprehensive understanding of the user's condition and treatment progress.
[0016] Lifestyle information has a potential impact on the user's treatment effect. In actual research, the lifestyle information of the target users is collected and analyzed. Through questionnaires or interviews, the dietary intake characteristics of the users are obtained to understand whether the users prefer high-fat and high-sugar foods and whether they have the habit of drinking alcohol, etc. Because long-term high-fat and high-sugar diets increase the liver burden and affect the treatment effect, and alcohol consumption directly damages the liver. The exercise habit characteristics of the users are inquired, including exercise frequency, exercise intensity, etc. Moderate exercise helps improve the body's immunity, promote liver blood circulation, and is beneficial for drug treatment. The work and rest characteristics of the users are understood, such as whether they have regular work and rest times and daily sleep times. Regular work and rest are beneficial for the liver's self-repair. The smoking habit characteristics of the users are understood. Smoking increases the liver metabolism burden and affects the treatment effect.
[0017] In a prediction model, the different frequency access values of the target user within a preset time period are important data sources. For example, when collecting data, the visit time nodes are set as the 0th, 4th, 8th, 12th, 24th, 36th, 48th, 60th, 72nd, 84th, 96th, 108th, 120th, 132nd, and 144th weeks. At these time nodes, the users are visited to obtain the relevant index values. Taking the HBsAg index as an example, the numerical changes are recorded during each visit. By analyzing the HBsAg values at different time nodes, the dynamic changes in the hepatitis B virus surface antigen level can be observed, and the effect of drug treatment on virus inhibition can be understood. At the same time, by combining the values of other indicators such as Alt at each visit node, the changes in the user's liver function over time can be comprehensively evaluated, providing a basis for predicting the treatment effect. At the 0th week visit, the initial HBsAg and Alt values of the user are obtained as the basis for subsequent comparison. As the treatment progresses, these indicators are continuously monitored during the 4th week, 8th week, etc. visits. If the HBsAg value gradually decreases during the treatment process, it indicates that the inhibitory effect of the drug on the virus is gradually emerging; if the Alt value gradually returns to normal after increasing initially, it means that the liver cells are gradually repairing after experiencing an inflammatory reaction. By continuously observing the changes in these indicators at different visit nodes, the treatment effect and the development trend of the disease can be judged more accurately. In addition, the changes in other indicators, such as the seroconversion of HBeAg and the changes in HBVDNA quantification, can also be combined to comprehensively evaluate the progress and effect of the treatment.
[0018] The initial Recurrent Neural Network-Long Short-Term Memory Network (RNN-LSTM) model is one of the core tools for prediction. In the initial stage of the research, an RNN-LSTM model with a general structure was constructed as the initial model. This model has the characteristic of highlighting the role of time series and is suitable for dealing with the time series nature of user visit data. The model is mainly composed of a convolutional layer, a pooling layer, and an LSTM layer. Before entering the fully connected layer, the feature dimension is changed through the LSTM layer. In the initial stage, the parameters of the model such as Input_size, Hidden_size, Batch_size, Learningrate, Seq_length, Epoch, etc. are given initial values. These initial values can refer to relevant research or be set according to experience. Subsequently, the model is trained and optimized through the training sample set, and the model parameters are adjusted to make it more suitable for the prediction task. When constructing the initial model, the Input_size parameter determines the number of features of the model input data. According to the previous analysis of user data, its value is determined to ensure that the model can fully process the input data. The Hidden_size parameter controls the number of neurons in the hidden layer and affects the learning ability and complexity of the model. The Batch_size parameter determines the number of data samples input into the model during each training. A suitable Batch_size can balance the training speed and model performance. The Learningrate parameter controls the step size of parameter update during model training. An overly large step size causes unstable model training, while an overly small step size makes the training speed too slow. The Seq_length parameter specifies the length of the time series of the input data and is set according to the characteristics of user visit data. The Epoch parameter represents the number of rounds of model training. Through multiple trainings, the model gradually learns the rules in the data. The setting of these initial parameters lays the foundation for the training of the model, and subsequent continuous optimization through training improves the prediction accuracy of the model.
[0019] The training sample set and the validation sample set are crucial for the training and evaluation of the model. The training sample set is used to represent users who receive different treatment cycles of PEG-IFNα-2b, and a large amount of user data receiving this drug treatment is collected as the training sample set. For example, eligible user data are screened out from three projects, namely TB1901IFN, TB1211IFN, TB1007IFN, and the Everest project. These data include various aspects such as users' clinical parameter information, lifestyle information, and different frequency access values. These data are sorted and labeled to form a training sample set for training the initial RNN-LSTM model, enabling the model to learn the features and patterns in the data. The validation sample set also comes from user data receiving PEG-IFNα-2b treatment but is independent of the training sample set. After using the training sample set to train the model, the validation sample set is used to validate the trained model. Through the validation sample set, the performance of the model can be evaluated, such as indicators like prediction accuracy and mean squared error, to determine whether the model is overfitting or underfitting, and then the model is adjusted and optimized to ensure that the model has good generalization ability on different data. When collecting the training sample set, the data are strictly screened to exclude samples with severe data missing or not meeting the research standards to ensure the quality of the training data. When labeling the data, according to the treatment results of users, such as whether the clinical cure standard (HBsAg < 0.05 IU / ml) is achieved, the data are labeled into corresponding categories for the model to learn. During the use of the validation sample set, the trained model is applied to the validation sample set, and indicators such as prediction accuracy and mean squared error are calculated. If the prediction accuracy is low and the mean squared error is large, it indicates that the model has an underfitting problem and the model complexity needs to be increased or the training parameters need to be adjusted; if it performs well on the training sample set but the accuracy drops significantly and the mean squared error increases on the validation sample set, then an overfitting phenomenon occurs, and measures such as reducing the model complexity and increasing the regularization term need to be taken for optimization to improve the generalization ability and prediction accuracy of the model.
[0020] S102, process the initial recurrent neural network-long short-term memory network model based on the training sample set and the validation sample set to generate a target recurrent neural network-long short-term memory network model.
[0021] In one implementation, any number of data features in the training sample set are obtained. From the training sample sets collected from the three projects of TB1901IFN, TB1211IFN, and TB1007IFN and the Everest project, data features such as the age, gender, medication duration, HBsAg, HBeAg, HBVDNA, HBsAb, HBeAb, and Alt of the patients are selected. For example, age reflects the differences in the physical functions of patients. Younger patients have a faster metabolism and respond more quickly to drugs. Gender differences can affect the immune response and thus the treatment effect. Indicators such as HBsAg and Alt directly reflect the condition of hepatitis B and liver function. These features reflect the situation of patients from different perspectives and provide a basis for subsequent analysis and model training.
[0022] Based on the number of each data feature in the training sample set, a sampling ratio is generated. Based on the sampling ratio, the training sample set is sampled to generate a preset number of sampling features. The number of occurrences of each data feature in the training sample set is counted. Suppose there are 1000 cases of patient data in the training sample set, 950 cases with complete age features recorded, 1000 cases with complete gender features recorded, 980 cases with complete HBsAg data, etc. According to these numbers, the sampling ratio is calculated. To ensure a reasonable proportion of each feature in the sampling, relative ratio calculation can be used. Taking the age feature as an example, its sampling ratio is ; the sampling ratio of the gender feature is . The sampling ratio can ensure that in subsequent sampling, each data feature participates in model training according to its importance and integrity in the overall population.
[0023] It is set to generate 500 sampling features. According to the sampling ratio calculated above, the data in the training sample set is sampled. For example, data is selected from the age feature at a sampling ratio of 0.95, and approximately age data are selected; 500 data are selected from the gender feature. During the sampling process, methods such as random sampling or stratified sampling can be used to ensure the randomness and representativeness of the sampling. Random sampling can avoid human bias, and stratified sampling can ensure that reasonable proportions of different categories of data (such as different age groups, genders, etc.) are selected. The 500 sampling features obtained in this way contain representative samples of each data feature and provide effective data for subsequent grouping and model training.
[0024] Based on any data feature and each sampling feature, response groups and poor response groups are generated. Each group of data sets contains a preset number of data samples, and at least one data sample includes identification information. Missing value imputation processing is performed on the response group and the poor response group respectively to generate missing value imputation prediction information. Taking the HBsAg data feature as an example, grouping is carried out in combination with other sampling features. According to clinical criteria, HBsAg < 0.05 IU / ml is the response group (Sustained response, SR), indicating a better treatment effect; HBsAg >= 0.05 IU / ml is the poor response group (Nonresponse, NR). When grouping, other sampling features such as age, gender, and medication duration are comprehensively considered. For a patient aged 35, male, with 12 weeks of medication and an HBsAg value of 0.03 IU / ml, since the HBsAg value meets the criteria of the response group, the patient is classified into the response group; while another patient aged 42, female, with 8 weeks of medication and an HBsAg value of 0.1 IU / ml is assigned to the poor response group. Each group is set to contain 250 data samples, and each data sample records the detailed information of the patient as identification information, such as patient number, basic information, and numerical values of various indicators, for convenient subsequent analysis.
[0025] In the response group and the poor response group, some data samples have missing values. For example, in the response group, some patients are missing the Alt value at a certain visit. Methods such as linear prediction are used for missing value imputation. Linear prediction constructs a prediction function to find the parameter vector of the model, minimizing the sum of the squares of the differences between the predicted values and the true values in the training set. Suppose a patient is missing the Alt value at the 8th week. Based on the Alt values at other visit times of this patient and the change trends of relevant indicators, the predicted value of the missing value is calculated using linear prediction. The same treatment is also carried out for the poor response group to generate the missing value imputation prediction information for both groups, ensuring data integrity and improving the model training effect.
[0026] Train a preset recurrent neural network - long short - term memory network model based on the data samples in the responder group and the non - responder group to generate a trained recurrent neural network - long short - term memory network model. Input the processed data samples of the responder group and the non - responder group into the preset RNN - LSTM model for training. The RNN - LSTM model consists of a convolutional layer, a pooling layer, and an LSTM layer and can process time - series data. During training, use the relevant index data (such as HBsAg, Alt, etc.) of the patient's first 5 visits as input, and the real data of the next visit as the label to let the model learn the relationship between the data. For example, input the HBsAg and Alt values of a patient's first 5 visits, and the model learns to predict the value of the 6th visit. After multiple rounds of training, adjust the model parameters (such as Input_size, Hidden_size, Batch_size, Learningrate, etc.) to make the model better fit the data and generate the trained RNN - LSTM model.
[0027] Process the trained recurrent neural network - long short - term memory network model based on the validation sample set to generate a validation result. If the data samples containing identification information in the validation result are risk factors that affect chronic hepatitis B users, then regard the trained recurrent neural network - long short - term memory network model as the target recurrent neural network - long short - term memory network model. Select data from an independent validation sample set and input it into the trained RNN - LSTM model. The validation sample set also contains various data characteristics and visit index data of patients. The model makes predictions on the data in the validation sample set and outputs the prediction results. For example, predict the HBsAg value and cure status of patients in the validation sample set. Compare these prediction results with the real data in the validation sample set to generate a validation result for evaluating the model performance. In the validation result, check the data samples containing identification information. If certain features (such as high HBsAg value, high ALT value, specific age range, etc.) in these data samples are proven to be risk factors that affect the treatment effect of chronic hepatitis B users, it indicates that the model can effectively identify these factors and has a certain reliability and effectiveness. At this time, determine the trained RNN - LSTM model as the target recurrent neural network - long short - term memory network model for subsequent prediction of the treatment effect of target users; if the model fails to accurately identify risk factors, or the prediction results have a large difference from the actual situation, then it is necessary to readjust the model parameters, improve the model structure, or re - process and train the data until the model achieves a satisfactory validation effect.
[0028] In another implementation, the missing sample information of the response group and the poor response group is obtained respectively. When processing the data of chronic hepatitis B patients, we have divided the patients into a response group (HBsAg < 0.05 IU / ml) and a poor response group (HBsAg >= 0.05 IU / ml) according to the HBsAg level. Taking the actual collected mixed dataset and the Everest project data as examples, in these two groups, there are data missing due to reasons such as patient loss to follow-up. For example, in the response group, the visit data of a patient at the 8th week and the 12th week is missing, and the missing indicators include HBsAg, Alt, etc.; in the poor response group, there are also some patients with different indicators missing at different time points. We organize the data to clarify the sample numbers where these missing values are located, the missing visit times, and the corresponding variable names, so as to obtain the missing sample information of the response group and the poor response group respectively.
[0029] The response group and the poor response group are processed respectively based on the missing sample information of the response group and the poor response group to generate predicted values for the missing samples. For the obtained missing sample information, a suitable method is used to generate predicted values for the missing samples. The article mentions using the method of linear prediction, and its principle is to calculate the sum of the squares of the differences between the predicted values and the true values in the training set , where, represents the predicted value, represents the true value corresponding to this predicted value, T represents the number of samples) and make it the smallest, so as to construct a prediction function to map the linear relationship between the input feature matrix and the label value. For example, for a patient in the response group whose Alt value at the 8th week is missing, we collect the Alt values and related indicators (such as HBsAg value, HBeAg value, etc.) of this patient at other visit times as the feature matrix, and use the linear prediction algorithm to calculate the predicted value of the missing Alt value at the 8th week. Similarly, for the samples with missing values in the poor response group, they are also processed according to this method to obtain the predicted values of the corresponding missing samples.
[0030] Obtain the true values corresponding to the missing samples, process the predicted values of the missing samples and the true values corresponding to the missing samples, and generate missing value imputation prediction information. For samples with missing values, if there are other reliable ways to obtain the true values corresponding to these missing samples, then record them. Suppose that during the subsequent research process, through supplementary detection or other channels, the true Alt values of the patients missing the 8th week Alt value in the above response group, as well as the true values corresponding to other missing samples in the poor response group, are obtained, and these true values will be used for comparative analysis with the predicted values. After obtaining the predicted values and true values of the missing samples, compare and analyze the two. By calculating the difference between the predicted value and the true value, a formula is used to measure this difference. For example, calculate the sum of the squares of the differences between the predicted value and the true value of each missing sample in the response group and the poor response group, and these calculation results constitute a part of the missing value imputation prediction information. Through this processing of a large number of missing samples, we can comprehensively understand the deviation between the predicted value and the true value, and then evaluate the accuracy of the linear prediction method, providing an important reference for subsequent data processing and model training. In practical applications, T is the number of samples with missing values. By performing such calculations on the predicted values and true values of all missing samples, we can obtain a comprehensive evaluation index for measuring the accuracy of the prediction and the performance of the model.
[0031] S103. Perform data cleaning and missing value supplementation processing on the clinical parameter information of the target user and the different frequency access values of the target user within the preset time period, and generate the target access values of the target user within the preset time period.
[0032] In one implementation, outlier detection is performed on the clinical parameter information of the target user and the different frequency access values of the target user within a preset time period to generate initial target user information, where the initial target user information is used to represent that there is no data with an error value exceeding a preset threshold in the clinical parameter information of the target user and the different frequency access values of the target user within the preset time period. Suppose there is a new target user coming for detection now. We have collected his clinical parameter information within a preset time period (such as from week 0 to week 144, including multiple visit nodes), covering age, gender, medication group, etc., and different frequency access values, such as the values of indicators like HBsAg, HBeAg, HBVDNA, HBsAb, HBeAb, and Alt. In actual operation, the outlier judgment criteria are set based on medical knowledge and past research results. For example, for the HBsAg indicator, the normal range is set at 0 - 1000 IU / ml (this is an example range, and the actual value should be based on professional medical standards). If it exceeds this range, it is considered an outlier; for age, considering the common age range of chronic hepatitis B patients, the outlier range is set at less than 10 years old or greater than 80 years old (also an assumed range). When viewing the data of this target user, it is found that at the 12th week visit, the HBsAg value reached 3000 IU / ml, far exceeding the normal range; at the same time, the age record is 90 years old, also exceeding the set normal age range. At this time, these abnormal data are marked or excluded. The data set obtained after such processing is the initial target user information. There is no data with an error value exceeding the preset threshold in the clinical parameter information and different frequency access values in this information, providing a reliable data basis for subsequent accurate analysis.
[0033] Obtain the missing value thresholds for each clinical parameter information and the missing value thresholds for different frequencies of visit values. Process the initial target user information based on the missing value thresholds for each clinical parameter information and the missing value thresholds for different frequencies of visit values to generate missing sample information, where the missing sample information is used to characterize the samples with missing values and the positions and variable names where the missing values are located. Determine the missing value thresholds for each clinical parameter information and different frequencies of visit values. For example, it is stipulated that if the data of the HBsAg index is missing for 3 consecutive visits, it is considered that there is a serious missing problem in the sample for the HBsAg index; for the visit time node, if the missing exceeds 30% of the total number of visits, it is also considered that there is a missing problem. Check the data of this target user and find that the data for the 8th, 12th, and 16th weeks are all missing for the HBsAg index. According to the set HBsAg missing value threshold, mark this sample as a sample with a missing value, and record the positions where the missing values are located (the 8th, 12th, and 16th weeks) and the variable name (HBsAg). If the total number of visits of this target user is 15 times and 6 of them have missing data, exceeding the 30% threshold, it is also marked as a sample with a missing value, and the relevant detailed information is recorded. Through comprehensive inspection, generate complete missing sample information, which details the situation of each sample with a missing value and provides a key basis for subsequent processing.
[0034] Process the missing sample information to generate missing value imputation prediction information, and process the missing value imputation prediction information to generate the target visit values of the target user within a preset time period. For the missing sample information, adopt appropriate methods to generate missing value imputation prediction information. Use the method of linear prediction, the principle of which is to calculate the sum of the squares of the differences between the predicted values and the true values in the training set and minimize it, so as to construct a prediction function to map the linear relationship between the input feature matrix and the label value. Taking the missing HBsAg values of this target user as an example, collect the HBsAg values and related indicators (such as HBeAg, HBVDNA, Alt, etc.) at other visit times as the feature matrix, and use the linear prediction algorithm to calculate the predicted values of the missing HBsAg values for the 8th, 12th, and 16th weeks. Suppose that after calculation, the predicted value of the HBsAg value for the 8th week of this target user is 150 IU / ml, the 12th week is 120 IU / ml, and the 16th week is 90 IU / ml. The samples with missing values in the group with poor response are also processed in the same way to obtain the predicted values of each missing sample, and then generate missing value imputation prediction information, which provides a reference for filling in the missing data.
[0035] The imputed prediction information of missing values is further processed to obtain the target access value of the target user within a preset time period. The calculated predicted values of missing values are inserted into the original dataset according to the corresponding samples and visit times. For this target user, the predicted HBsAg values of 150 IU / ml, 120 IU / ml, and 90 IU / ml at the 8th, 12th, and 16th weeks are inserted into their corresponding positions respectively, thus filling in the missing data of the HBsAg index of this target user. After such processing is performed on all samples with missing values, the target access values containing complete or imputed clinical parameter information and different frequency access values are obtained. These data will serve as important inputs for subsequent model training and prediction, used to construct a more accurate prediction model and improve the accuracy of predicting the treatment effect of this target user.
[0036] S104, process the lifestyle information of the target user to generate a dynamic risk adjustment factor.
[0037] In one implementation, feature extraction processing is performed on the lifestyle information of the target user to generate dietary intake features, exercise habit features, rest schedule features, and smoking habit features. Suppose there is a chronic hepatitis B patient as the target user, and researchers collect his lifestyle information through methods such as questionnaires, interviews, and wearable device monitoring. In terms of dietary intake, it is learned that in this user's daily diet, the intake frequency of fried foods is relatively high, reaching 4 times a week, and the daily vegetable intake is less than 200 grams, and he likes high-sugar beverages, drinking at least 1 bottle every day. Based on this information, the extracted dietary intake features are: high intake of fried foods, insufficient vegetable intake, and excessive consumption of high-sugar beverages. In terms of exercise habits, this user exercises only 1 - 2 times a week, with each exercise duration not exceeding 30 minutes, and the exercise intensity is relatively low, mostly walking. Thus, the generated exercise habit features are: low exercise frequency, insufficient exercise duration, and low exercise intensity. In terms of the rest schedule, this user often stays up late, going to bed at 1 - 2 am and getting up at 9 - 10 am, and has poor sleep quality, with frequent dreams and easy waking. Based on this, the rest schedule features are obtained as: late bedtime, irregular sleep time, and poor sleep quality. In terms of smoking habits, this user smokes 10 - 15 cigarettes a day and has a smoking history of 10 years. This smoking habit feature is: large smoking amount and long smoking history.
[0038] Process the dietary intake characteristics, exercise habit characteristics, rest characteristics, and smoking habit characteristics to generate a dietary risk score, an exercise risk score, a rest quality score, a smoking risk score, and a dynamic feature weight value. Among them, the dynamic feature weight value is generated based on the lifestyle information of different target users. According to the extracted lifestyle characteristics, researchers formulate corresponding scoring rules to generate risk scores. For dietary intake characteristics, bad eating habits such as high intake of fried foods, insufficient intake of vegetables, and excessive consumption of sugary drinks each increase the burden on the liver and affect the treatment effect, so a certain risk score is given respectively. For example, a high intake of fried foods is scored 3 points, insufficient intake of vegetables is scored 2 points, and excessive consumption of sugary drinks is scored 3 points. The comprehensive dietary risk score of this user is 8 points. In terms of exercise habits, a low exercise frequency is scored 3 points, insufficient exercise duration is scored 2 points, and low exercise intensity is scored 2 points, and the exercise risk score is 7 points. In the rest characteristics, going to bed late is scored 3 points, irregular sleep time is scored 2 points, and poor sleep quality is scored 3 points, and the rest quality score is 8 points. In terms of smoking habits, a large smoking amount is scored 4 points, and a long smoking history is scored 4 points, and the smoking risk score is 8 points. When generating the dynamic feature weight value, considering that different lifestyles have different degrees of influence on the treatment effects of different patients, the weights are determined based on the analysis of a large amount of sample data. Suppose it is found through data analysis that for the specific patient group where the target user is located, diet has a relatively large impact on the treatment effect, followed by exercise and rest, and smoking has a relatively small impact. Then the generated dynamic feature weight values are: the weight of dietary intake characteristics is 0.4, the weight of exercise habit characteristics is 0.2, the weight of rest characteristics is 0.25, and the weight of smoking habit characteristics is 0.15.
[0039] Process the dietary risk score, exercise risk score, rest quality score, smoking risk score, and dynamic feature weight value to generate a dynamic risk adjustment factor. After obtaining the dietary risk score, exercise risk score, rest quality score, smoking risk score, and dynamic feature weight value, a dynamic risk adjustment factor is generated through a specific calculation method. The weighted summation method can be used, that is, the dynamic risk adjustment factor = dietary risk score × weight of dietary intake characteristics + exercise risk score × weight of exercise habit characteristics + rest quality score × weight of rest characteristics + smoking risk score × weight of smoking habit characteristics. Substitute the data in the above example into the formula, and the dynamic risk adjustment factor of this target user . This dynamic risk adjustment factor comprehensively reflects the potential impact degree of the lifestyle of this target user on the efficacy of PEG-IFNα-2b in the treatment of chronic hepatitis B. The higher the value, the greater the negative impact of the lifestyle on the treatment effect. In the subsequent prediction model, this factor can be used to more accurately evaluate the treatment situation of patients and provide a reference basis for personalized treatment.
[0040] S105. Process the target access value and the dynamic risk adjustment factor of the target user within a preset time period based on the target recurrent neural network-long short-term memory network model to generate the access prediction value of the target user within the target time period.
[0041] In one implementation, process the target access value and the dynamic risk adjustment factor of the target user within a preset time period based on the target recurrent neural network-long short-term memory network model to generate the access value of the first frequency and the moment state vector of the second frequency. Suppose the visit data of a certain chronic hepatitis B patient (target user) is as follows (unit: HBsAg is IU / ml, Alt is U / L): The visit time is the 0th week, with HBsAg of 250, Alt of 80, and the dynamic risk adjustment factor (DRF) of 7.8; the visit time is the 4th week, with HBsAg of 180, Alt of 75, and the dynamic risk adjustment factor (DRF) of 7.8; the visit time is the 8th week, with HBsAg of 150, Alt of 70, and the dynamic risk adjustment factor (DRF) of 7.8; the visit time is the 0th week, with HBsAg of 120, Alt of 65, and the dynamic risk adjustment factor (DRF) of 7.8.
[0042] Obtain the parameter matrix, and process the parameter matrix, the access value of the first frequency, and the moment state vector of the second frequency to generate the moment state vector of the first frequency. The method includes the calculation formula for obtaining the moment state vector of the first frequency, and the calculation formula is: ; where represents the moment state vector of the first frequency, A represents the parameter matrix, represents the moment state vector of the second frequency, represents the bias vector of the hidden state. Take the first 5 visit data (in this example, the first 4 supplemented data + DRF) as the input feature matrix, and the target is the predicted value of the 5th visit.
[0043] Initial state vector Randomly generated, for example =[0.1, -0.2, 0.3, 0.05].
[0044] Parameter matrix A and bias vector Obtained through training: .
[0045] Gradually calculate the state vector: The 1st visit (t = 1): Input feature =[250, 80, 7.8], combined with Calculate: .
[0046] Second visit (t = 2): Input features = [180, 75, 7.8], based on Update: .
[0047] Third visit (t = 3): Input after completion = [150, 70, 7.8], update . After 4 iterations, the final state vector is obtained , which is used to predict the HBsAg value at the 5th visit.
[0048] Generate the access prediction value of the target user within the target time period based on the moment state vector of the first frequency and the moment state vector of the second frequency. The model will, according to the information in these state vectors, combine the previously learned patterns to predict the relevant index values of the target user at each future visit node. For example, predict the predicted values of indicators such as HBsAg and HBeAg of the target user at visit nodes such as the 16th week and the 24th week. These predicted values comprehensively consider the current state of the target user (target access value), the impact of living habits (dynamic risk adjustment factor), as well as the time series features and patterns learned by the model, providing an important reference basis for doctors to judge the treatment effect of patients and formulate subsequent treatment plans. If the predicted HBsAg value gradually decreases in subsequent visits, it means that the treatment effect is good; on the contrary, if the predicted value does not change significantly or increases, the treatment plan needs to be adjusted. For example, to predict the HBsAg value at the 5th visit (the 16th week), based on and the output of the fully connected layer, the predicted value is 85 IU / ml. Repeatedly apply the state update formula to predict the HBsAg values at the 6th - 13th visits. When the predicted value is lower than 0.05 IU / ml (such as at the 10th visit), it is marked as the cure time node.
[0049] S106, Process the access prediction value of the target user within the target time period to generate the drug cure prediction time information of the target user.
[0050] In one implementation, the access prediction values of the target user within the target time period are processed to generate HBsAg values and Alt values with different frequencies within the target time period. The access prediction values output by the model include the prediction results of multiple indicators, from which the HBsAg values and Alt values are extracted. Suppose the predicted HBsAg values of the target user at visit nodes such as the 16th week, 24th week, and 36th week are 80 IU / ml, 60 IU / ml, and 30 IU / ml respectively, and the corresponding Alt values are 50 U / L, 45 U / L, and 40 U / L respectively. These prediction values reflect the changing trends of the hepatitis B surface antigen level and the hepatocyte damage index in the patient's body as the treatment time progresses, and are an important basis for subsequent analysis.
[0051] Based on a preset detection threshold, the HBsAg values and Alt values with different frequencies within the target time period are processed to generate the attribute classification information of the target user, where the attribute classification information of the target user is used to characterize the drug cure type of the target user. The preset detection threshold is set according to medical knowledge and clinical experience. For example, the critical value of HBsAg is set to 0.05 IU / ml (a key indicator for judging clinical cure), and at the same time, the normal range threshold of Alt is set to 0 - 40 U / L. For the predicted HBsAg value, if it is greater than 0.05 IU / ml, it means that the virus is not effectively controlled, and if it is less than or equal to 0.05 IU / ml, it indicates that the clinical cure standard may be achieved. For the Alt value, if it exceeds the normal range, it means that there is inflammation or damage to the hepatocytes, and if it is within the normal range, it indicates that the hepatocyte state is relatively good. According to these thresholds, the above-predicted HBsAg values and Alt values are processed. At the 16th week, the HBsAg value of 80 IU / ml is greater than 0.05 IU / ml, and the Alt value of 50 U / L exceeds the normal range. It is comprehensively judged that the attribute classification of the target user at this time is "not cured and there is inflammation in the hepatocytes"; at the 24th week, the HBsAg value of 60 IU / ml is greater than 0.05 IU / ml, but the Alt value of 45 U / L is close to the normal range, and the attribute classification is "not cured but the hepatocyte inflammation is reduced"; at the 36th week, the HBsAg value of 30 IU / ml is greater than 0.05 IU / ml, and the Alt value of 40 U / L is within the normal range, and the attribute classification is "not cured but the hepatocyte state is good". These attribute classification information clearly show the treatment status and liver condition of the patient at different time points, providing a more detailed basis for predicting the cure time in the future.
[0052] Process the HBsAg values with different frequencies within the target time period based on the preset HBsAg threshold and the attribute classification information of the target user to generate the drug cure prediction time information of the target user, where the drug cure prediction time information of the target user is used to characterize the cure time node information of the target user. Based on the preset HBsAg threshold (0.05 IU / ml) and the above-generated attribute classification information, further process the HBsAg values with different frequencies within the target time period. As the treatment continues, if it is predicted that the HBsAg value at a certain visit node is less than or equal to 0.05 IU / ml for the first time, and this value continues to remain at this level or lower in subsequent visits, and the hepatocyte status is judged to be stable or good in combination with the attribute classification information, then this time point is recognized as the predicted cure time node. Suppose that in subsequent predictions, at the 72nd week, the HBsAg value is 0.04 IU / ml, and the Alt value is 35 U / L within the normal range, and the attribute classification is "close to cure and good hepatocyte status". In the visit predictions after the 72nd week, the HBsAg values also remain below 0.05 IU / ml. At this time, the drug cure prediction time information of the target user can be determined to be the 72nd week, that is, it is predicted that the target user will reach the clinical cure state around the 72nd week. This information is of great guiding significance for doctors to formulate treatment plans, evaluate treatment effects, and patients to understand the development of their own conditions. For example, doctors can reasonably adjust the drug dosage or decide whether to continue the current treatment plan according to the predicted cure time; patients can also have a clearer expectation of the treatment process, better cooperate with the treatment and arrange their lives.
[0053] Obtain various information of the target user by the server, including clinical parameters, living habits, numerical values of different frequency visits, as well as the initial model, training sample set and validation sample set. Use the training and validation sample sets to process the initial recurrent neural network-long short-term memory network model to generate the target model, where there are steps such as data feature sampling, sample grouping, and missing value processing. Clean the data of the target user and supplement the missing values to obtain the target visit numerical values, process the living habit information to generate the dynamic risk adjustment factor, and combine the two to input into the target model to obtain the visit prediction value. Further process the visit prediction value to generate the drug cure prediction time information. First, extract the HBsAg value and Alt value from the prediction value, generate the attribute classification information based on the preset threshold, and use this to judge the drug cure type. Then, combine the preset HBsAg threshold and the attribute classification information to determine the cure time node. The SA-Golden standard (HBsAg value / Alt value) is used to assist in judging the cure possibility of the patient.
[0054] In one implementation, as Figure 2 shown, the present application also provides a prediction device for the hepatitis B treatment time based on a neural network model, including: An acquisition module 201, configured to acquire clinical parameter information of a target user, lifestyle information of the target user, different frequency access values of the target user within a preset time period, an initial recurrent neural network-long short-term memory network model, a training sample set, and a validation sample set, where the training sample set is used to represent users who have received different treatment cycles of PEG-IFNα-2b; A processing module 202, configured to process the initial recurrent neural network-long short-term memory network model based on the training sample set and the validation sample set to generate a target recurrent neural network-long short-term memory network model; perform data cleaning and missing value supplementation processing on the clinical parameter information of the target user and the different frequency access values of the target user within the preset time period to generate target access values of the target user within the preset time period, where the target access values of the target user within the preset time period include at least two of HBsAg, HBeAg, HBVDNA, HBsAb, HBeAb, and Alt; process the lifestyle information of the target user to generate a dynamic risk adjustment factor; process the target access values of the target user within the preset time period and the dynamic risk adjustment factor based on the target recurrent neural network-long short-term memory network model to generate access prediction values of the target user within a target time period; process the access prediction values of the target user within the target time period to generate drug cure prediction time information of the target user.
[0055] In another implementation manner of the present application, the processing module 202 is configured to process the initial recurrent neural network-long short-term memory network model based on the training sample set and the validation sample set to generate a target recurrent neural network-long short-term memory network model, including: Acquire any quantity of data features in the training sample set; Generate a sampling ratio based on the quantity of each data feature in the training sample set; Sample the training sample set based on the sampling ratio to generate a preset quantity of sampling features; Process any data feature and each sampling feature to generate a response group and a non-response group, where each data group contains a preset quantity of data samples, and at least one data sample includes identification information; Perform missing value imputation processing on the response group and the non-response group respectively to generate missing value imputation prediction information; Train a preset recurrent neural network-long short-term memory network model based on the data samples in the response group and the non-response group to generate a trained recurrent neural network-long short-term memory network model; Process the trained recurrent neural network-long short-term memory network model based on the validation sample set to generate a validation result; If the data samples containing identification information in the verification result are risk factors characterizing the impact on chronic hepatitis B users, the trained recurrent neural network-long short-term memory network model is used as the target recurrent neural network-long short-term memory network model.
[0056] In another implementation manner of the present application, the processing module 202 is configured to process the initial recurrent neural network-long short-term memory network model based on the training sample set and the verification sample set to generate a target recurrent neural network-long short-term memory network model, and further includes: Obtain the missing sample information of the responder group and the non-responder group respectively; Process the responder group and the non-responder group respectively based on the missing sample information of the responder group and the non-responder group to generate predicted values of the missing samples; Obtain the true values corresponding to the missing samples, and process the predicted values of the missing samples and the true values corresponding to the missing samples to generate missing value imputation prediction information; The method further includes a calculation formula for obtaining the missing value imputation prediction information, and the calculation formula is: ; Wherein, represents the predicted value, represents the true value corresponding to the predicted value, and T represents the number of samples.
[0057] In another implementation manner of the present application, the processing module 202 is configured to perform data cleaning and missing value supplementation processing on the clinical parameter information of the target user and the different frequency access values of the target user within a preset time period to generate the target access values of the target user within the preset time period, including: Perform outlier detection processing on the clinical parameter information of the target user and the different frequency access values of the target user within a preset time period to generate initial target user information, where the initial target user information is used to characterize that there is no data with an error value exceeding a preset threshold in the clinical parameter information of the target user and the different frequency access values of the target user within a preset time period; Obtain the missing value threshold of each clinical parameter information and the missing value threshold of the different frequency access values; Process the initial target user information based on the missing value threshold of each clinical parameter information and the missing value threshold of the different frequency access values to generate missing sample information, where the missing sample information is used to characterize the samples with missing values and the positions and variable names where the missing values are located; Process the missing sample information to generate missing value imputation prediction information; Process the missing value imputation prediction information to generate the target access values of the target user within a preset time period.
[0058] In another embodiment of the present application, the processing module 202 is configured to process the lifestyle information of the target user to generate a dynamic risk adjustment factor, including: Performing feature extraction processing on the lifestyle information of the target user to generate dietary intake features, exercise habit features, work and rest features, and smoking habit features; Processing the dietary intake features, exercise habit features, work and rest features, and smoking habit features to generate a dietary risk score, an exercise risk score, a work and rest quality score, a smoking risk score, and a dynamic feature weight value, where the dynamic feature weight value is generated based on the lifestyle information of different target users; Processing the dietary risk score, the exercise risk score, the work and rest quality score, the smoking risk score, and the dynamic feature weight value to generate a dynamic risk adjustment factor.
[0059] In another embodiment of the present application, the processing module 202 is configured to process the target access value and the dynamic risk adjustment factor of the target user within a preset time period based on the target recurrent neural network-long short-term memory network model to generate an access prediction value of the target user within the target time period, including: Processing the target access value and the dynamic risk adjustment factor of the target user within a preset time period based on the target recurrent neural network-long short-term memory network model to generate an access value of the first frequency and a moment state vector of the second frequency; Obtaining a parameter matrix; Processing the parameter matrix, the access value of the first frequency, and the moment state vector of the second frequency to generate a moment state vector of the first frequency; Generating an access prediction value of the target user within the target time period based on the moment state vector of the first frequency and the moment state vector of the second frequency; The method includes a calculation formula for obtaining the moment state vector of the first frequency, and the calculation formula is: ; where represents the moment state vector of the first frequency, A represents the parameter matrix, represents the moment state vector of the second frequency, represents the bias vector of the hidden state.
[0060] In another embodiment of the present application, the processing module 202 is configured to process the access prediction value of the target user within the target time period to generate drug cure prediction time information of the target user, including: Processing the access prediction value of the target user within the target time period to generate HBsAg values and Alt values of different frequencies within the target time period; Process the HBsAg values and Alt values with different frequencies within a target time period based on a preset detection threshold to generate attribute classification information of a target user, where the attribute classification information of the target user is used to characterize the drug cure type of the target user; Process the HBsAg values with different frequencies within a target time period based on a preset HBsAg threshold and the attribute classification information of the target user to generate drug cure prediction time information of the target user, where the drug cure prediction time information of the target user is used to characterize the cure time node information of the target user.
[0061] Each embodiment in this application is described in a related manner. For the same and similar parts among the embodiments, reference can be made to each other. The key points of each embodiment are the differences from other embodiments. In particular, for the embodiments of the prediction method, electronic device, electronic equipment, and readable storage medium for evaluating the hepatitis B treatment time based on a neural network model, since they are basically similar to the embodiment of the prediction method for the hepatitis B treatment time based on a neural network model described above, the description is relatively simple, and reference can be made to the partial description of the embodiment of the prediction method for the hepatitis B treatment time based on a neural network model for the related parts.
Claims
1. A prediction method for the treatment time of hepatitis B based on a neural network model, characterized in that, Including: Obtaining the clinical parameter information of the target user, the lifestyle information of the target user, the different frequency access values of the target user within a preset time period, an initial recurrent neural network-long short-term memory network model, a training sample set, and a validation sample set, where the training sample set is used to represent users who have received different treatment cycles of PEG-IFNα-2b; Processing the initial recurrent neural network-long short-term memory network model based on the training sample set and the validation sample set to generate a target recurrent neural network-long short-term memory network model; Performing data cleaning and missing value supplementation processing on the clinical parameter information of the target user and the different frequency access values of the target user within a preset time period to generate the target access values of the target user within the preset time period, where the target access values of the target user within the preset time period include at least two of HBsAg, HBeAg, HBVDNA, HBsAb, HBeAb, and Alt; Processing the lifestyle information of the target user to generate a dynamic risk adjustment factor; Processing the target access values of the target user within a preset time period and the dynamic risk adjustment factor based on the target recurrent neural network-long short-term memory network model to generate the access prediction value of the target user within the target time period; Processing the access prediction value of the target user within the target time period to generate the drug cure prediction time information of the target user.
2. The method according to claim 1, wherein Processing the initial recurrent neural network-long short-term memory network model based on the training sample set and the validation sample set to generate a target recurrent neural network-long short-term memory network model, including: Obtaining any quantity of data features in the training sample set; Generating a sampling ratio based on the quantity of each data feature in the training sample set; Sampling the training sample set based on the sampling ratio to generate a preset quantity of sampling features; Processing any data feature with each sampling feature to generate a response group and a non-responding group, where each data group contains a preset quantity of data samples, and at least one data sample includes identification information; Performing missing value imputation processing on the response group and the non-responding group respectively to generate missing value imputation prediction information; Training a preset recurrent neural network-long short-term memory network model based on the data samples in the response group and the non-responding group to generate a trained recurrent neural network-long short-term memory network model; Processing the trained recurrent neural network-long short-term memory network model based on the validation sample set to generate a validation result; If the data sample containing identification information in the validation result represents a risk factor affecting chronic hepatitis B users, then taking the trained recurrent neural network-long short-term memory network model as the target recurrent neural network-long short-term memory network model.
3. The method according to claim 2, characterized in that, Processing the initial recurrent neural network-long short-term memory network model based on the training sample set and the validation sample set to generate a target recurrent neural network-long short-term memory network model, further including: Respectively obtaining the missing sample information of the response group and the non-responding group; Processing the response group and the non-responding group respectively based on the missing sample information of the response group and the non-responding group to generate predicted values of the missing samples; Obtain the true value corresponding to the missing sample; Process the predicted value of the missing sample and the true value corresponding to the missing sample to generate missing value imputation prediction information; The method further includes a calculation formula for obtaining the missing value imputation prediction information, and the calculation formula is: ; Among them, represents the predicted value, represents the true value corresponding to the predicted value, and T represents the number of samples.
4. The method according to claim 1, wherein Perform data cleaning and missing value supplementation processing on the clinical parameter information of the target user and the different frequency access values of the target user within a preset time period to generate the target access values of the target user within the preset time period, including: Perform outlier detection processing on the clinical parameter information of the target user and the different frequency access values of the target user within a preset time period to generate initial target user information, where the initial target user information is used to represent that there is no data with an error value exceeding a preset threshold in the clinical parameter information of the target user and the different frequency access values of the target user within the preset time period; Obtain the missing value threshold of each clinical parameter information and the missing value threshold of different frequency access values; Process the initial target user information based on the missing value threshold of each clinical parameter information and the missing value threshold of different frequency access values to generate missing sample information, where the missing sample information is used to represent the samples with missing values and the positions and variable names where the missing values are located; Process the missing sample information to generate missing value imputation prediction information; Process the missing value imputation prediction information to generate the target access values of the target user within the preset time period.
5. The method according to claim 1, characterized in that Process the lifestyle information of the target user to generate a dynamic risk adjustment factor, including: Perform feature extraction processing on the lifestyle information of the target user to generate dietary intake features, exercise habit features, rest schedule features, and smoking habit features; Process the dietary intake features, exercise habit features, rest schedule features, and smoking habit features to generate a dietary risk score, an exercise risk score, a rest quality score, a smoking risk score, and a dynamic feature weight value, where the dynamic feature weight value is generated based on the lifestyle information of different target users; Process the dietary risk score, the exercise risk score, the rest quality score, the smoking risk score, and the dynamic feature weight value to generate a dynamic risk adjustment factor.
6. The method according to claim 5, characterized in that, Based on the target recurrent neural network-long short-term memory network model, process the target access values of the target user within a preset time period and the dynamic risk adjustment factor to generate the access prediction value of the target user within the target time period, including: Based on the target recurrent neural network-long short-term memory network model, process the target access values of the target user within a preset time period and the dynamic risk adjustment factor to generate the access value of the first frequency and the moment state vector of the second frequency; Obtain the parameter matrix; Process the parameter matrix, the access value of the first frequency, and the moment state vector of the second frequency to generate the moment state vector of the first frequency; Generate the access prediction value of the target user within the target time period based on the moment state vector of the first frequency and the moment state vector of the second frequency; The method includes a calculation formula for obtaining the moment state vector of the first frequency, and the calculation formula is: ; Among them, represents the moment state vector of the first frequency, A represents the parameter matrix, represents the moment state vector of the second frequency, represents the bias vector of the hidden state.
7. The method according to claim 6, wherein Process the access prediction value of the target user within the target time period to generate the drug cure prediction time information of the target user, including: Process the access prediction value of the target user within the target time period to generate HBsAg values and Alt values with different frequencies within the target time period; Process the HBsAg values and Alt values with different frequencies within the target time period based on a preset detection threshold to generate the attribute classification information of the target user, where the attribute classification information of the target user is used to characterize the drug cure type of the target user; Process the HBsAg values with different frequencies within the target time period based on a preset HBsAg threshold and the attribute classification information of the target user to generate the drug cure prediction time information of the target user, where the drug cure prediction time information of the target user is used to characterize the cure time node information of the target user.
8. A prediction device for the treatment time of hepatitis B based on a neural network model, characterized in that, The device includes: An acquisition module, configured to acquire the clinical parameter information of the target user, the living habit information of the target user, the access numerical values with different frequencies of the target user within a preset time period, an initial recurrent neural network-long short-term memory network model, a training sample set, and a validation sample set, where the training sample set is used to characterize users who have received different treatment cycles of PEG-IFNα-2b; A processing module, configured to process the initial recurrent neural network-long short-term memory network model based on the training sample set and the validation sample set to generate a target recurrent neural network-long short-term memory network model; perform data cleaning and missing value supplementation processing on the clinical parameter information of the target user and the access numerical values with different frequencies of the target user within a preset time period to generate the target access numerical values of the target user within the preset time period, where the target access numerical values of the target user within the preset time period include at least two of HBsAg, HBeAg, HBVDNA, HBsAb, HBeAb, and Alt; process the living habit information of the target user to generate a dynamic risk adjustment factor; process the target access numerical values of the target user within the preset time period and the dynamic risk adjustment factor based on the target recurrent neural network-long short-term memory network model to generate the access prediction value of the target user within the target time period; process the access prediction value of the target user within the target time period to generate the drug cure prediction time information of the target user.
9. An electronic device, characterized in that, Includes: A first processor; And a memory, configured to store the executable instructions of the first processor; Wherein, the first processor is configured to execute the prediction method for the hepatitis B treatment time based on the neural network model according to any one of claims 1 to 7 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a second processor, it implements the prediction method for the hepatitis B treatment time based on the neural network model according to any one of claims 1 to 7.