Model training method and device, risk determination method and device, equipment and medium

CN120112903APending Publication Date: 2025-06-06BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380011007.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In deep learning technology, especially in prognostic risk assessment in the medical field, it is difficult for the prior art to effectively utilize scarce labeled data, resulting in insufficient diversity of training samples and overfitting of models.

Method used

By obtaining a first data sample corresponding to a plurality of sample objects and simulating the data factors affecting the target risk based on the noise data, a plurality of second data samples are generated. Then, the preset model is trained based on the first data sample and the second data sample to obtain a risk determination model.

Benefits of technology

The training samples are expanded through the second data sample, which improves the diversity and richness of the training samples, avoids the problem of overfitting the model and insufficient training samples, and improves the generalization ability of the risk determination model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112903A_ABST
    Figure CN120112903A_ABST
Patent Text Reader

Abstract

The model training method comprises the steps that first data samples corresponding to multiple sample objects are acquired, each first data sample comprises at least one data factor, and the data factors influence target risks of the sample objects in the target direction; simulating at least one data factor influencing the target risk of the sample object based on the noise data to obtain a plurality of second data samples; training a preset model based on the first data sample and the second data sample to obtain a risk determination model; wherein the risk determination model is used for determining a first risk category corresponding to the target direction of the sample object. The invention further provides a risk determination method and device, equipment and a medium, and aims to avoid model overfitting and improve the risk determination accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Model training methods, risk determination methods, devices, equipment and media Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a model training method, risk determination method, device, equipment and medium. Background Art

[0002] With the development of artificial intelligence, deep learning technology has been applied to all aspects of human life. Deep learning requires a large amount of labeled data as training samples, and labeled data is crucial to the effectiveness of deep learning. For example, in the medical field, the acquisition of clinical data on gene or site mutations within genes is crucial to the effectiveness of prognostic risk assessment.

[0003] Overview

[0004] The present disclosure provides a model training method, wherein the method comprises:

[0005] Acquire a first data sample corresponding to each of a plurality of sample objects, wherein the first data sample includes at least one data factor, and the data factor affects a target risk of the sample object in a target direction;

[0006] Simulating at least one data factor affecting the target risk of the sample object based on the noise data to obtain a plurality of second data samples;

[0007] A preset model is trained based on the first data sample and the second data sample to obtain a risk determination model; wherein the risk determination model is used to determine a first risk category corresponding to the target direction of the sample object.

[0008] Optionally, a first number of the first data samples is greater than a second number of the second data samples, and a ratio of the first number to the second number is greater than or equal to 1.5 and less than or equal to 3.

[0009] Optionally, the first data sample and the second data sample both carry their own corresponding first labels, and the first labels are used to indicate the first risk category corresponding to the sample object and to supervise the training.

[0010] Optionally, the noise data includes at least one data group, and different data groups are used to simulate second data samples corresponding to different sample objects; and simulating at least one data factor affecting the target risk of the sample object based on the noise data to obtain multiple second data samples includes:

[0011] Acquire category information corresponding to each of the data groups, the category information being used to indicate a second risk category of the data group in the target direction;

[0012] Based on at least one of the data groups and the corresponding category information, at least one second data sample is generated that meets the risk category indicated by the category information.

[0013] Optionally, the first data sample and the second data sample both carry respective corresponding first labels, where the first labels are used to indicate the first risk category corresponding to the sample object;

[0014] The second risk category is used to indicate the risk level of the sample object in the target direction, and the second risk category is the same as the first risk category;

[0015] Alternatively, the second risk category is used to indicate the level of risk of the sample object in the target direction, and the second risk category is used to indicate the controllable range after the risk occurs to the sample object in the target direction.

[0016] Optionally, the noise data includes at least one data group, and different data groups are used to simulate characteristic data corresponding to different objects; the noise data is generated by the following steps:

[0017] Generate basic data that conforms to normal distribution based on target commands;

[0018] Performing at least one round of sampling on the basic data to obtain data groups corresponding to the at least one round of sampling; wherein, in each round of sampling, data corresponding to at least one description dimension is sampled from the basic data;

[0019] Different description dimensions correspond to different influencing factors of the object related to the target direction.

[0020] Optionally, simulating at least one data factor affecting the target risk of the sample object based on the noise data to obtain a plurality of second data samples includes:

[0021] Inputting the noise data into a generator to obtain a plurality of second data samples output by the generator; wherein the generator is obtained by training a generative adversarial network using the plurality of third data samples and the noise data samples;

[0022] In which, the third data sample carries a second label representing whether the data is real data, the generator in the generative adversarial network is used to generate predicted data based on the noise data sample, and the discriminator in the generative adversarial network is used to determine whether the predicted data is real data.

[0023] Optionally, the noise data includes at least one data group, and different data groups are used to simulate characteristic data corresponding to different objects; inputting the noise data into the generator to obtain a plurality of second data samples output by the generator includes:

[0024] Inputting at least one of the data groups and corresponding category information into the generator, and obtaining second data samples output by the generator that correspond to the at least one of the data groups;

[0025] The category information is used to indicate a second risk category of the data group in the target direction.

[0026] Optionally, the step of training the generative adversarial network includes:

[0027] Using the plurality of the third data samples, training a discriminator in the generative adversarial network;

[0028] The noise data samples are used to train the generator in the generative adversarial network.

[0029] Optionally, one third data sample corresponds to one sample object, and the third data sample further carries a third label representing a second risk category of the sample object in the target direction; and using the plurality of third data samples to train the discriminator in the generative adversarial network includes:

[0030] Inputting a plurality of the third data samples into the discriminator to obtain a first prediction result and a second prediction result output by the discriminator; wherein the first prediction result is used to indicate whether the third data sample is real data, and the second prediction result is used to indicate a second risk category corresponding to the third data sample;

[0031] Based on the first prediction result, the second label, the second prediction result and the third label, the parameters of the discriminator are updated.

[0032] Optionally, before training the generator in the generative adversarial network using the noise data sample, the method further includes:

[0033] inputting the noise data sample and the category information corresponding to the noise data sample into the generator;

[0034] Inputting the predicted data corresponding to the noise data sample output by the generator into the discriminator;

[0035] Obtaining a first prediction result and a second prediction result output by the discriminator; wherein the first prediction result is used to indicate whether the predicted data is real data, and the second prediction result is used to indicate a second risk category corresponding to the predicted data;

[0036] The parameters of the discriminator are updated based on the first prediction result.

[0037] Optionally, the plurality of third data samples include real data samples and synthetic data samples, the second label carried by the real data sample indicates that the real data sample is real data, and the second label carried by the synthetic data sample indicates that the real data sample is fake data; and updating the loss of the discriminator based on the first prediction result, the second label, the second prediction result, and the third label includes:

[0038] Updating the loss of the discriminator based on the first prediction result, the second prediction result, the second label, the second prediction result, and the third label corresponding to the real data sample;

[0039] Based on the first prediction result and the second label corresponding to the synthetic data sample, the parameters of the discriminator are updated.

[0040] Optionally, the training a generator in the generative adversarial network using the noise data sample includes:

[0041] inputting the noise data sample and the category information corresponding to the noise data sample into the generator;

[0042] Inputting the predicted data corresponding to the noise data sample output by the generator into the discriminator;

[0043] Obtaining a first prediction result and a second prediction result output by the discriminator; wherein the first prediction result is used to indicate whether the predicted data is real data, and the second prediction result is used to indicate a second risk category corresponding to the predicted data;

[0044] The parameters of the generator are updated based on the first prediction result, or the parameters of the generator are updated based on the first prediction result, the second prediction result and the category information.

[0045] Optionally, multiple sample objects originate from the same individual, and multiple third data samples include at least one real data sample. The target probability corresponding to the sample object to which the real data sample belongs is less than or equal to 0.5, and the target probability represents the statistical significance of the sample object in causing the target risk to the individual in the target direction.

[0046] Optionally, the generative adversarial network includes multiple network layers, each of which includes at least one neuron. During the training of the generative adversarial network, the method further includes:

[0047] In at least one training session, some target neurons among the plurality of neurons are inactivated.

[0048] Optionally, in the at least one training session, the inactivation of some of the neurons includes at least one of the following:

[0049] In N consecutive training sessions, the target neurons subjected to the inactivation treatment are different;

[0050] In N consecutive training sessions, the target neurons subjected to the inactivation treatment are partially repeated;

[0051] In N consecutive trainings, the target neurons subjected to the inactivation process are located in different network layers;

[0052] In N consecutive trainings, the network layer where the target neurons subjected to the inactivation processing are located is repeated.

[0053] Optionally, the deactivation process includes: setting the weight parameter corresponding to the target neuron to 0, or fixing the weight parameter of the target neuron at that time.

[0054] Optionally, inputting the noise data into a generator to obtain a plurality of second data samples output by the generator includes:

[0055] Inputting the noise data into the trained generative adversarial network to obtain a plurality of candidate data output by the generator, and a prediction probability corresponding to each candidate data, wherein the prediction probability is a probability representing that the candidate data is true data;

[0056] Based on the predicted probability, a plurality of second data samples are screened out from a plurality of the candidate data.

[0057] The present disclosure also provides a risk determination method, wherein the method comprises:

[0058] Acquiring characteristic distribution data of a target object to be classified; wherein the characteristic distribution data includes at least one data factor affecting the target risk of the sample object in the target direction;

[0059] Inputting the characteristic distribution data into a risk determination model to obtain a first risk category corresponding to the target risk of the target object;

[0060] Wherein, the risk determination model is obtained based on the training method of the model.

[0061] Optionally, the target direction includes gene mutation direction, and the target risk includes prognostic risk; or, the target direction includes site mutation direction, and the target risk includes prognostic risk.

[0062] By adopting the training method of the model of the embodiment of the present disclosure, a first data sample corresponding to each of multiple sample objects can be obtained, and based on the noise data, at least one data factor affecting the target risk of the sample object is simulated to obtain multiple second data samples; then, the preset model is trained based on the first data sample and the second data sample to obtain a risk determination model; wherein, the first data sample includes at least one data factor, and the data factor affects the target risk of the sample object in the target direction. The first data sample and the second data sample both carry their own corresponding first labels, and the first label is used to indicate the first risk category corresponding to the sample object, and is used to supervise the training. The risk determination model is used to determine the first risk category corresponding to the risk occurring in the target direction of the sample object.

[0063] Since the training samples used in the process of training the risk determination model include the first data sample and the second data sample, wherein the second data sample is simulated based on the noise data, the second data sample can be used to expand the first data sample, so that the sample objects with few data samples can also be expanded through the second data sample, thereby improving the diversity and richness of the training samples, avoiding the problem of poor model training effect caused by insufficient training samples, and also avoiding the problem of strong tendency of training samples, for example, the sample objects targeted by the first data sample all tend to a certain feature, which leads to the problem of model overfitting.

[0064] The present disclosure also discloses an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the model training method or risk determination method as described above when executing the computer program.

[0065] The present disclosure also discloses a computer-readable storage medium, which stores a computer program that enables a processor to execute the model training method or risk determination method described in the present disclosure.

[0066] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the specific implementation methods of the present disclosure are listed below.

[0067] BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following is a brief introduction to the drawings required for the description of the embodiments or related technologies. Obviously, the drawings described below are some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. It should be noted that the scales in the drawings are for illustration only and do not represent the actual scale.

[0069] FIG1 is a schematic diagram showing a step flow of a model training method according to an embodiment of the present disclosure;

[0070] FIG2 is a schematic diagram showing the overall principle of a model training method according to an embodiment of the present disclosure;

[0071] FIG3 shows a schematic diagram of ROC of testing the trained preset model under the first quantity ratio of the second data sample and the first data sample;

[0072] FIG4 shows a schematic diagram of ROC of testing the trained preset model under the second data sample and the first data sample of the second quantity ratio;

[0073] FIG5 shows a schematic diagram of ROC for testing the trained preset model under the third quantity ratio of the second data sample and the first data sample;

[0074] FIG6 is a schematic diagram showing the overall process of generating a second data sample by using a generative adversarial network in an embodiment of the present disclosure;

[0075] FIG7 is a schematic diagram showing a process of training a discriminator in an embodiment of the present disclosure;

[0076] FIG8 shows another schematic diagram of a process for training a discriminator in an embodiment of the present disclosure;

[0077] FIG9 is a schematic diagram showing a process of training a generator in an embodiment of the present disclosure;

[0078] FIG10a is a schematic diagram showing the ROC results of a risk determination model trained using the second data sample generated by the first generative adversarial network in a test set in an embodiment of the present disclosure;

[0079] FIG10 b is a schematic diagram showing a process of training a first type of generative adversarial network in an embodiment of the present disclosure;

[0080] FIG11a is a schematic diagram showing ROC results of a risk determination model trained using a second data sample generated by a second generative adversarial network in a test set in an embodiment of the present disclosure;

[0081] FIG11b is a schematic diagram showing a process of training a second generative adversarial network according to an embodiment of the present disclosure;

[0082] FIG12 shows a schematic diagram of a network structure of a generative adversarial network according to an embodiment of the present disclosure;

[0083] FIG13 shows a schematic diagram of the process of applying the model training method in the medical field, taking the TP53 gene as an example;

[0084] FIG14 shows the distribution of statistical mutation sites of TP53 along the entire amino acid chain;

[0085] FIG15 shows the characteristic data distribution of statistical mutation sites of TP53;

[0086] FIG16 shows a schematic diagram of the process of applying the model training method to another medical field, taking the TP53 gene as an example;

[0087] FIG17 shows a schematic flow chart of steps of a risk determination method according to an embodiment of the present disclosure;

[0088] FIG18 shows a schematic diagram of the framework structure of a training device for a model according to an embodiment of the present disclosure;

[0089] FIG19 shows a schematic diagram of the framework structure of a risk determination device in an embodiment of the present disclosure.

[0090] Detailed description

[0091] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0092] In related technologies, in some fields that use deep learning technology, it is difficult to obtain labeled data. For example, in the prognostic risk assessment in the medical field, it is necessary to determine the extent to which mutations in some genes or sites in genes affect the prognostic risk. If deep learning technology is used for prediction, it is difficult to ensure the diversity and richness of training samples because some genes or sites have more clinical data, while other genes or sites have less clinical data.

[0093] For example, taking the high-frequency mutation gene TPS3 as an example, by statistically analyzing the distribution of TPS3 mutation sites on the entire amino acid chain, it was found that TP53 missense mutations are mainly concentrated in the site ranges of 100-281 and 326-335, while the remaining sites are low-frequency mutation sites. There is a large amount of clinical data for high-frequency mutation sites, while the clinical data for low-frequency mutation sites are very scarce.

[0094] Taking risk prediction as an example, risk prediction in the medical field includes disease prognosis prediction, which includes both survival prediction and prognostic risk prediction. Therefore, if deep learning is used to predict the prognostic risk of various TPS3 mutation sites for certain diseases, the clinical data is primarily concentrated on a few high-frequency mutation sites. This can lead to overfitting during model training, resulting in poor generalization performance, such as poor performance on small-sample data classification problems.

[0095] Of course, the risk prediction described in this application is not limited to disease prognosis prediction. It can also be applied to other scenarios requiring risk prediction, such as predicting the risks of surgery in the medical field to reduce the stress of surgical patients. In this scenario, the surgical risks of patients can be predicted based on data factors such as the patient's age, gender, and medical history. For example, in natural disaster risk assessment, it is necessary to predict the risk of natural disasters occurring in a region, such as predicting the risk of mudslides in region H during heavy rain.

[0096] From this we can see that in the field of risk prediction, there are problems such as less sample data and low sample diversity.

[0097] In light of this, the present disclosure aims to increase the diversity of training samples and address the problem of model overfitting caused by the high concentration of training samples in deep learning technologies. Its core concept is to simulate false data samples based on real data samples, construct a training set including false and real data samples, and train a preset model to obtain a risk determination model. In this way, the false data samples can enhance the diversity of training samples and avoid model overfitting during training.

[0098] 1 and 2 , FIG1 shows a schematic diagram of a step flow of a model training method, and FIG2 shows a schematic diagram of the overall principle of a model training method. As shown in FIG1 and FIG2 , the following steps may be specifically included:

[0099] Step S101: Acquire a first data sample corresponding to each of a plurality of sample objects, wherein the first data sample includes at least one data factor, and the data factor affects a target risk of the sample object in a target direction.

[0100] In this embodiment, the sample objects can be determined based on the application scenario. For example, in predicting the prognosis risk of a disease, the sample objects can be genes or sites on genes. In predicting the prognosis and survival period of a disease, the sample objects can be genes, sites on genes, patients, etc. In predicting the risk of natural disasters, the sample objects can be regions, geological structures, etc.

[0101] In this embodiment, the first data sample may include text data, or may include image data, or may include both text data and image data. More specifically, if the first data sample is text data, the text data may be data describing the risk of the sample object in the target direction. For example, in the medical field, the sample object may be a patient, and the first data sample may include the patient's clinical data. If the first data sample is image data, the image data may be data obtained by collecting images of the sample object. For example, in the medical field, the sample object may be a patient, and the first data sample may be data obtained by collecting images of the patient's lesion site, such as CT images, B-ultrasound images, and magnetic resonance images. For another example, in the field of natural disaster risk prediction, the sample object may be a certain region, and the first data sample may be image data collected in region H before, during, and after a rainstorm, such as thermal imaging data and high-definition imaging data.

[0102] In this case, a first data sample corresponding to a sample object can be obtained from a database. The first data sample is a real data sample of the sample object. For example, in predicting the prognosis risk of a disease, taking the sample object as a gene as an example, missense mutation data and clinical data of the TP53 gene can be obtained from the ICGC and MSK databases. Based on the obtained missense mutation data and clinical data, the first data sample of the TP53 gene can be obtained. For example, in predicting the prognosis survival period of a disease, taking the sample object as a site in a gene as an example, missense mutation data and clinical data of the mutation site in the TP53 gene can be obtained from the ICGC and MSK databases. Based on the obtained missense mutation data and clinical data, the first data sample of the mutation site in the TP53 gene can be obtained. For another example, in predicting the risk of natural disasters, taking the sample object as region H as an example, meteorological record data and geological change data of region H can be obtained from historical disaster records and geological data sources of region H. Based on the meteorological record data and geological change data, the first data sample of region H can be obtained.

[0103] Specifically, the first data sample includes at least one data factor, and different data factors can describe the relationship between the sample subject and the target risk from different perspectives. This can also be understood as: at least one data factor can influence the probability of the sample subject experiencing the target risk. Taking the prediction of disease prognosis risk as an example, the data factors in the first data sample may include: the impact score of the site mutation on protein structure / function, the impact level of the missense mutation at the site, and the location of the site on the gene. This information can reflect the probability of the sample subject experiencing the target risk from different perspectives. By combining these data factors, the prognostic risk of the site after the mutation can be determined. Taking the prediction of wind for natural disasters as an example, the data factors in the first data sample may include: the looseness of the soil, the composition of the soil, the slope gradient, the rock layer structure, etc. This information can reflect the probability of a debris flow disaster in region H from different perspectives. By combining these data factors, the risk of debris flows in region H during the rainy season can be determined.

[0104] The target risk of a sample object refers to the risk occurring in the target direction. The target direction can represent the direction of the risk. For example, in disease prognosis prediction, the target direction can be understood as the direction in which the sample object mutates, such as the prognostic risk after a site mutation. For example, in disease prognosis and survival, the target direction can be the direction in which the sample object receives treatment, such as the prognostic survival of a patient after liver cancer surgery. For another example, in natural disaster warning, the target direction can be the direction of heavy rain in region H, such as the risk of mudslides in region H after a heavy rain, or the risk of landslides in region H after a heavy rain.

[0105] In this embodiment, a sample object can correspond to a first data sample, and multiple sample objects can correspond to multiple first data samples. Among them, the sample object can be an object that obtains data through transactions in the application scenario. For example, in the prognosis prediction of a disease, the prognostic risk caused by gene mutations and site mutations is worthy of attention. However, for some genes and sites, the probability of mutation is not high. For objects with a low mutation probability, the amount of data in their first data samples is small. For other genes and sites, the probability of mutation is high, such as the TP53 gene. The probability of mutation is high. For objects with a high mutation probability, the amount of data in their first data samples is large. Therefore, the sample object in this embodiment can be an object that obtains data samples through transactions in the application scenario.

[0106] Step S102: Based on the noise data, simulate at least one data factor that affects the target risk of the sample object to obtain a plurality of second data samples.

[0107] In this embodiment, the noise data can be understood as the data required to generate a false second data sample. The noise data can be a set of random numbers, and the data distribution of the noise data can match the data distribution of the data factors contained in the sample object. For example, if the data factors contained in the sample object are distributed in the impact score of the site mutation on the protein structure / function, the impact level of the missense mutation of the site, the location information of the site on the gene, etc., then the noise data also has a data distribution similar to the impact score of the site mutation on the protein structure / function, the impact level of the missense mutation of the site, the location information of the site on the gene, etc.

[0108] At least one data factor of the sample object can be simulated based on the noise data, thereby obtaining multiple second data samples. In practice, the noise data can be fitted based on an analysis of each data factor in the actual first data sample, thereby obtaining the second data samples. In some examples, the noise data may include multiple data groups, and the data distribution of each data group matches the data distribution of the data factors contained in the sample object. In this way, a fitting can be performed on the data in each data group to obtain a second data sample, and multiple data groups can be fitted to obtain multiple second data samples.

[0109] Among them, since the second data sample is data generated based on noise data, the data type of the second data sample can be the same as that of the first data sample. For example, if the first data sample is text data, the second data sample is also text data, and the data dimensions included are consistent, such as both include clinical data; if the first data sample is image data, the second data sample is also image data, and the image type is consistent. For example, if the first data sample is a magnetic resonance image, the second data sample is also a magnetic resonance image. For another example, if the first data sample is an ultrasound image, the second data sample can also be an ultrasound image; if the first data sample includes text data and image data, the second data sample can also include text data and image data.

[0110] In this embodiment, since the second data sample is simulated from noise data, the corresponding sample object can be a virtual object. For example, in disease prognosis prediction, the second data sample can be a data sample simulating a site mutation, representing a virtual site. In practice, since this second data sample object is obtained by fitting the noise data based on an analysis of various data factors in the first data sample, and represents a virtual object, it can be used to expand the real sample object. For example, if the real sample object is a gene or site with a high mutation probability, the simulated virtual second data sample can be data corresponding to a gene or site with a low mutation probability, thereby increasing the size of the training data.

[0111] It should be noted that the data factors included in the second data sample and the data factors included in the first data sample may have the same distribution in terms of quantity and type.

[0112] Step S103: training a preset model based on the first data sample and the second data sample to obtain a risk determination model; wherein the risk determination model is used to determine a first risk category corresponding to the target direction of the sample object.

[0113] In this embodiment, the first data sample and the second data sample can be used as training sets to train the preset model. The preset model can be used to determine the first risk category corresponding to the target direction of the sample object based on the input data sample. The first risk category can be the risk category corresponding to the target risk occurring in the target direction. For example, in the prognosis prediction of a disease, the first risk category can be the high or low risk of the prognosis risk after a site mutation; or, the first risk category can be the risk category corresponding to other risks occurring in the target direction. For example, in the prognosis survival prediction of a disease, the first risk category can be the category of the prognosis survival of the disease after a site mutation.

[0114] Among them, in the case where the first risk category can be the risk category corresponding to other risks occurring in the target direction, the data factors included in the first data sample affect the target risk occurring in the sample object, and it is associated with other risks, that is, the data factors have an influence relationship on both the target risk and other risks.

[0115] For example, in the prediction of the prognosis survival period of a disease, the first risk category can be the category of the prognosis survival period of the disease after a site mutation occurs. The data factors included in the first data sample and the second data sample are both related to the prognosis risk after the site mutation occurs, and the level of the prognosis risk is directly correlated with the prognosis survival period, which is manifested in that the higher the prognosis risk, the shorter the prognosis survival period may be, and the lower the prognosis risk, the longer the prognosis survival period may be. Therefore, during the training process, the preset model can determine the category of the prognosis survival period of the sample object by analyzing and judging the data factors included in the first data sample and the second data sample.

[0116] As another example, in the risk prediction of natural disasters, the first risk category can be the probability of a mudslide disaster. The data factors included in the first data sample and the second data sample are both related to the risk of landslides during heavy rain attacks, and the level of landslide risk is correlated with mudslide disasters. The higher the risk of landslides, the higher the probability of a mudslide disaster, and the lower the risk of landslides, the lower the probability of a mudslide disaster. Therefore, during the training process, the preset model can determine the probability of a mudslide disaster occurring in the sample object by analyzing and judging the data factors included in the first data sample and the second data sample.

[0117] When training the preset model using the first and second data samples, supervised or unsupervised training methods can be used. In supervised training, a category label can be assigned to each first and second data sample to indicate its category. This category label can be used as a gradient to update the model parameters of the preset model. In unsupervised training, data classification without prior (known) classification criteria can be performed based on the differences in category characteristics in the feature space between the first and second data samples (not only between the first and second data samples, but also between the first and second data samples). Specifically, decision rules are established based on the statistical characteristics of the data samples to be classified, without requiring prior knowledge of the category characteristics. The spatial distribution of each data sample is segmented or merged into clusters based on their similarity. The feature category represented by each cluster is determined through field surveys or comparisons with known features. In this case, common algorithms for the preset model include regression analysis, trend analysis, equal mixture distance method, cluster analysis, principal component analysis, and pattern recognition.

[0118] Among them, since the second data sample is a data sample simulated based on the noise data, in practice, when generating the second data sample based on the noise data, it can be instructed to generate the second data sample according to the set first label, so that the data factors included in the generated second data sample meet the characteristics of the first label. In this way, the first label of the second data sample can be the set first label and participate in the training supervision of the preset model.

[0119] By adopting the technical solution of the embodiment of the present disclosure, since the training samples used in the process of training the risk determination model include a first data sample and a second data sample, wherein the second data sample is simulated based on noise data, the purpose of expanding the first data sample can be achieved through the second data sample, so that sample objects with few data samples can also be expanded through the second data sample, thereby improving the diversity and richness of the training samples, avoiding the problem of poor model training effect caused by insufficient training samples, and also avoiding the problem of strong tendency of training samples, for example, the sample objects targeted by the first data sample all tend to a certain feature, which causes the model to overfit.

[0120] In some embodiments, the first data sample and the second data sample both carry their own corresponding first labels, where the first labels are used to indicate a first risk category corresponding to the sample object and to supervise the training.

[0121] In this embodiment, the first data sample and the second data sample both carry their respective corresponding first labels, and the first label is used to indicate the first risk category corresponding to the sample object. For example, when the first risk category includes categories of high and low prognostic risks, if the first label is 0, it can represent a low prognostic risk, and if the first label is 1, it can represent a high prognostic risk. For another example, when the first risk category includes categories of prognostic survival, the first label can include four levels, such as 0 / 1 / 2 / 3, with different levels corresponding to different prognostic survival ranges, such as 0 representing a survival period of 0-6 months, 1 representing a survival period of 6 months to 1.5 years, 2 representing a survival period of 1.5 years to 3 years, and 3 representing a survival period of 3 years to 6 years. Of course, this is only an exemplary description. In the division of prognostic survival, reference can be made to relevant fields and no special limitation is made here. For another example, when the first risk category includes categories of high and low risk of debris flow disasters, if the first label is 0, it can represent a low risk of debris flow disasters, and if the first label is 1, it can represent a high risk of debris flow disasters.

[0122] Among them, the first label can be used as supervision in the process of training the preset model. Specifically, the prediction result corresponding to the first data sample output by the preset model can be obtained, and the loss function can be constructed according to the first label and the prediction result corresponding to the first data sample to obtain the loss value, and the parameters of the preset model can be updated according to the loss value; and the prediction result corresponding to the second data sample output by the preset model can be obtained, and the loss function can be constructed according to the first label and the prediction result corresponding to the second data sample to obtain the loss value, and the parameters of the preset model can be updated according to the loss value.

[0123] By adopting the technical solution of this embodiment, supervised training is adopted when training the preset model. The preset model can converge as quickly as possible through the first label, thereby improving the efficiency of model training.

[0124] In some embodiments, the number of first data samples and the number of second data samples may be different, such as the first number of first data samples may be greater than the second number of second data samples, and the ratio of the first number to the second number is greater than or equal to 1.5 and less than or equal to 3.

[0125] In some embodiments, the ratio of the first number to the second number may be 2, ie, the number of first data samples may be twice the number of second data samples.

[0126] For example, referring to Figures 3 to 5, ROC diagrams of testing the trained preset model under different ratios of second data samples and first data samples are shown. As shown in Figures 3 to 5, 213 first data samples are included; among them, as shown in Figure 3, 50 second data samples are combined with 213 first data samples, training set ACC = 0.997, test set ACC = 0.864, and test set AUC = 0.90; as shown in Figure 4, 100 second data samples are combined with 213 first data samples, training set ACC = 0.9868, test set ACC = 0.955, and test set AUC = 0.975; as shown in Figure 5, 200 second data samples are combined with 213 first data samples, training set ACC = 0.9838, test set ACC = 0.864, and test set AUC = 0.850.

[0127] Therefore, it can be seen that by comparing the ACC and AUC of the test set under the three quantity ratios, it can be found that when the ratio between the first data sample and the second data sample is close to 2:1, the performance of the risk determination model on the test set is best.

[0128] The following describes how to obtain the second data sample based on the noise data:

[0129] In some embodiments A, the noise data may be a random number conforming to a normal distribution. Based on the noise data, data factors of multiple sample objects may be simulated, thereby generating multiple second data samples at once based on the noise data. The noise data may include at least one data group, with different data groups being used to simulate different second data samples. In other words, different data groups may correspond to different virtual sample objects; thus, one data group may be used to generate one second data sample.

[0130] In which, the distribution of each data group can be consistent with the distribution of data factors in the first data sample. For example, if the first data sample includes M data factors, the data group can also include M-dimensional data, and the noise data can be N*M size data, where N represents the number of data groups. Based on the N*M size noise data, N second data samples can be generated, and each second data sample includes M data factors.

[0131] Each data group in the noise data can be assigned category information, which can be used to indicate a second risk category for the data group in the target direction. When generating the second data sample, a data sample that meets the second risk category can be generated. In a specific implementation, category information corresponding to each data group can be obtained, which can be used to indicate the second risk category of the data group in the target direction. Based on at least one data group and the corresponding category information, at least one second data sample that meets the risk category indicated by the category information can be generated.

[0132] The second risk category is the same as or different from the first risk category. Specifically, when the second risk category is different from the first risk category, the second risk category and the first risk category are associated. Furthermore, the second risk category to which the sample object belongs affects the first risk category to which the sample object belongs. This allows for the generation of a second data sample related to the characteristics identified by the second risk category, thereby improving the authenticity of the second data sample.

[0133] Of course, when the second risk category is the same as the first risk category, the second data sample generated by the noise data can directly adopt the first label indicated by the category information to be used as supervision when the preset model is trained, thereby improving the efficiency of collecting training samples.

[0134] In one example, in the medical field, the first data sample may include clinical data and mutation genomic data, and the first risk category includes at least one of the prognostic risk category, the gene mutation category, and the prognostic survival category; the second risk category includes at least one of the prognostic risk category, the gene mutation category, and the prognostic survival category.

[0135] In this example, the first risk category can be any one of the prognostic risk category, gene mutation category, and prognostic survival category, or multiple thereof. Of course, in the case of multiple categories, the preset model can implement multi-category classification tasks, and the second risk category can also be any one of the prognostic risk category, gene mutation category, and prognostic survival category, or multiple thereof. Specifically, there can be an association relationship between the first risk category and the second risk category, and the association relationship can mean that the two are the same or that the two have a certain causal relationship in the occurrence of risks in the target direction.

[0136] In this case, the first risk category and the second risk category may be the same or different.

[0137] In some examples, the first risk category may be the same as the second risk category. In this case, both the first risk category and the second risk category may be used to indicate the risk level of the sample object in the target direction. For example, indicating the risk level of the sample object in the target direction. For example, in the medical field, it may refer to the prognostic risk of a disease, or the prognostic survival category or the mutation risk of a gene locus (gene mutation category); for another example, in natural disaster risk assessment, the first risk category may be the risk level of a debris flow disaster or a landslide in a certain area.

[0138] Alternatively, in some other examples, the first risk category may be different from the second risk category. In this case, the second risk category may be used to indicate the level of risk of the sample object in the target direction, and the first risk category may be used to indicate the controllable range of the sample object after a risk occurs in the target direction. In other words, the second risk category to which the sample object belongs affects the first risk category to which the sample object belongs, or there is an association between the second risk category and the first risk category. The controllable range refers to: when a risk occurs to the sample object, a series of operations are performed on the sample object to eliminate the risk, and the range in which the risk is eliminated is also referred to as the controllable range of the risk of the sample object. Within this controllable range, the sample object is safe.

[0139] In this embodiment, when the first risk category and the second risk category are the same, the simulated second data sample is simulated in the direction indicated by the first risk category, so that the data structure of the second data sample adapts to the first risk category, thereby improving the authenticity of the data and reducing the difference in data structure between it and the first data sample. Therefore, when training the preset model, not only the data samples of the preset model are enriched, but the consistency between the data samples is also improved, thereby improving the accuracy of the risk determination model.

[0140] When the first risk category and the second risk category are different, the simulated second data sample is simulated in the direction indicated by the second risk category, so that the second data sample becomes a data sample associated with the first risk category, and can be close to the first data sample in data structure. For example, the generated second data sample is a data sample with a higher prognostic risk, and its data structure can contain implicit features that are useful for predicting prognosis survival. Thus, the preset model can learn to predict risks based on implicit features related to the first risk category through the second data sample during training, thereby optimizing the performance of the risk determination model, so that when applying reasoning, it can not be limited to a specific data structure (the structure of the first data sample) and can be adapted to a variety of data.

[0141] For example, in the medical field, the first risk category can be a prognostic survival category, and the second risk category can be a category of high or low prognostic risk. For example, the first risk category corresponding to the site mutation is a prognostic survival category, and the second risk category has a range that affects the prognostic survival. For example, the higher the prognostic risk, the lower the level of the first risk category, the shorter the survival period, and the smaller the controllable range. For another example, in the field of natural disaster assessment, the first risk category can be the degree of debris flow damage (which indirectly reflects the size of its controllable range), and the second risk category can be the high or low risk of debris flow occurrence. The high or low risk of debris flow occurrence has a certain degree of high or low with the degree of debris flow destructiveness. The higher the risk, the higher the degree of destructiveness may be, and the lower the risk, the lower the degree of destructiveness may be.

[0142] In this case, the first label carried by the generated second data sample can be determined based on the set second risk category. In practice, a correspondence between the second risk category and the first risk category can be established, and the first label can be set based on the correspondence. In this case, the generated second data sample is data that matches the prognostic risk feature, and the first label it carries is used to indicate the prognostic survival period corresponding to the sample object, so that the risk determination model can determine the prognostic survival period of the sample object in the target direction based on the relationship between the input data samples (first data sample and second data sample) and the prognostic risk. In this way, the richness of the sample can be increased, thereby improving the generalization ability of the risk determination model.

[0143] In one example, in order to improve the authenticity of the second data sample, the second risk category and the first risk category can be divided more finely. For example, if the second risk category is a prognostic risk category and the first risk category is a prognostic survival category, the prognostic risk category can be divided into multiple levels that are not much different from the prognostic survival category. For example, if the first risk category includes four levels, such as 0 / 1 / 2 / 3, the second risk category can also include four levels, such as very low, low, medium, relatively high, and very high. The first risk category and the second risk category are thereby aligned to obtain a second data sample that meets the first risk category based on noise data that meets the second risk category.

[0144] For example, since noise data is a random number that conforms to a normal distribution, specifically, a normal distribution with a mean of 0 and a variance of 1, it is randomly sampled from this normal distribution to serve as noise data. In a specific implementation, basic data that conforms to a normal distribution can be generated based on the target command; and at least one round of sampling is performed on the basic data to obtain data groups corresponding to the at least one round of sampling. In each round of sampling, data corresponding to at least one descriptive dimension is sampled from the basic data, and different descriptive dimensions correspond to different influencing factors of the object related to the target direction.

[0145] Specifically, the target command can be a random command. Based on the random command, basic data that obeys a normal distribution with a mean of 0 and a variance of 1 can be obtained. Then, according to the distribution of the data factors in the first data sample, multiple rounds of sampling are performed from the basic data, and a data group is sampled in each round, that is, a random number of size 1*M that conforms to the normal distribution is sampled to generate a corresponding second data sample. Among them, the sampled data group needs to conform to the distribution of the data factors in the first data sample, and different data factors are used to describe the relationship between the sample object and the target risk from different angles. Then, the sampled data group includes data corresponding to at least one descriptive dimension, and one descriptive dimension corresponds to one data factor, that is, data corresponding to M descriptive dimensions, so that the data distribution in the sampled data group is consistent with the distribution of the data factors in the first data sample. If it contains M data factors, then the data group also contains M data.

[0146] After N rounds of sampling, random numbers of N*M size that conform to the normal distribution will be obtained, thereby obtaining the noise data. When adopting this implementation method, since the noise data conforms to the normal distribution with a mean of 0 and a variance of 1, it is consistent with the numerical range of the data factor in the first data sample, thereby enhancing the simulation of the second data sample to the first data sample and improving the authenticity of the second data sample.

[0147] The N here can be greater than or equal to 1, which is different from the meaning of N (N training times) later.

[0148] In some embodiments B, the noise data can be a random number that conforms to the normal distribution. When generating multiple second data samples based on the noise data, a generative adversarial network can be used to obtain the second data samples. Specifically, the noise data can be input into the generator to obtain multiple second data samples output by the generator; wherein the generator is obtained by training the generative adversarial network using multiple third data samples and noise data samples, and the third data sample carries a second label that characterizes whether the data is real data. The generator in the generative adversarial network is used to generate predicted data based on the noise data sample, and the discriminator in the generative adversarial network is used to determine whether the predicted data is real data.

[0149] Referring to Figure 6, a schematic diagram of the overall process of generating a second data sample using a generative adversarial network is shown. As shown in Figure 6, the generative adversarial network can be trained using multiple third data samples and noise data samples. After the training is completed, the noise data can be input into the generator in the generative adversarial network to obtain the second data sample.

[0150] Among them, when using multiple third data samples and noise data samples to train the generative adversarial network, the generator and discriminator in the generative adversarial network can be trained in stages, such as first using multiple third data samples to train the discriminator so that the discriminator can distinguish the true or false of the data, and then using noise data to train the generator so that the predicted data generated by the generator cannot be identified as true or false by the discriminator. Alternatively, the generator and discriminator in the generative adversarial network can also be trained simultaneously, such as inputting multiple third data samples into the discriminator in the generative adversarial network, inputting noise data samples into the generator in the generative adversarial network, and inputting the predicted data output by the generator into the discriminator at the same time, so that the discriminator judges the true or false of the input predicted data and the third data samples, and updates the parameters of the discriminator based on the true or false judgment results of the two, and updates the parameters of the generator based on the true or false results of the predicted data output by the discriminator.

[0151] When this embodiment is adopted, since the second data sample can be generated based on the pre-trained generative adversarial network, the authenticity of the second data sample is improved, thereby improving the training effect of the preset model.

[0152] As described in implementation A above, category information can be set for each data group in the noise data to characterize the second risk category of the data group in the target direction, thereby generating a second data sample that meets the characteristics of the second risk category. For example, when the noise data is input into the generator, as shown in Figure 6, at least one of the data groups and the corresponding category information can also be input into the generator to obtain a second data sample output by the generator corresponding to at least one data group.

[0153] Accordingly, during the training process of the generative adversarial network, each data group in the noise data sample can also carry category information, allowing the generator to generate predicted data that meets the second risk category, thereby improving the authenticity of the predicted data. As a result, after multiple training sessions, the second data samples generated by the generator meet the second risk category, further improving the authenticity of the second data samples.

[0154] Below, we explain how to train a generative adversarial network:

[0155] For example, a generative adversarial network includes a generator and a discriminator connected to the generator, wherein the predicted data generated by the generator will be input into the discriminator to judge whether the data is true or false. When the predicted data output by the generator is judged by the discriminator to be false data, it means that the degree of simulation of the data factors of the sample object by the generator is insufficient, and the parameters of the generator need to be adjusted; when the predicted data output by the generator is judged by the discriminator to be true data, it means that the degree of simulation of the data factors of the sample object by the generator is sufficient, and the training of the generative adversarial network can be terminated.

[0156] As described in the previous example, the generator and discriminator can be trained in stages. Specifically, the discriminator can be trained first to acquire the ability to distinguish true from false data, and then the generator can be trained to generate predicted data that the discriminator cannot identify as false. In a specific implementation, the discriminator in the generative adversarial network can be trained using multiple third data samples. After the discriminator is trained, the generative adversarial network can be trained using noise data samples.

[0157] In this example, the multiple third data samples may include real third data samples (real data samples) and false third data samples (synthetic data samples). Specifically, each third data sample may carry a second label, which characterizes whether the third data sample is a real data sample. The multiple third data samples may be input into the discriminator, and a loss function may be constructed based on the discrimination result output by the discriminator (hereinafter referred to as the first prediction result) and the second label corresponding to the third data sample. The parameters of the discriminator are updated, and when the training end condition of the discriminator is met, the training of the discriminator is stopped and the training of the generator is started.

[0158] When training the generator, the noise data samples and the category information carried by each data group in the noise data samples (hereinafter referred to as the third label) can be input into the generator, so that the generator generates predicted data, and the predicted data is input into the trained discriminator. The discriminator outputs a first prediction result. Since the predicted data is false data, and the purpose of training is to make the predicted data generated by the generator recognized by the discriminator as real data, the parameters of the generator can be updated according to the difference between the first prediction result output by the discriminator and the discrimination result of the real data (represented as the second label of the real data). When the training end condition of the generator is met, the training of the generator is stopped, thereby obtaining a trained generative adversarial network.

[0159] Among them, the training termination condition of the discriminator may refer to: the loss value is less than the preset value, or the number of training times reaches the preset number; the training termination condition of the generator may refer to: the difference between the first prediction result and the discrimination result of the real data is less than the preset difference, or the number of training times reaches the preset number.

[0160] As described above, in some examples, the second data sample may be a sample that meets the second risk category, and the first data sample may also be a sample that meets the second risk category, wherein the discriminator in the generative adversarial network can also judge the second risk category of the input data (the third data sample and the predicted data), so that the discriminator can not only judge whether the data is true or false but also judge the second risk category of the data. In this way, when the predicted data output by the generator is input into the discriminator, the discriminator can judge the second risk category of the generated predicted data, so that the predicted data generated by the generator is closer to the real data and is more in line with the second risk category.

[0161] Accordingly, referring to Figure 7, a schematic diagram of the process of training the discriminator is shown. As shown in Figure 7, the third data sample can carry a third label, and the third label represents the second risk category of the third data sample in the target direction. Then, multiple third data samples can be input into the discriminator to obtain the first prediction result and the second prediction result output by the discriminator, and the parameters of the discriminator are updated based on the first prediction result, the second label, the second prediction result and the third label.

[0162] The first prediction result is used to indicate whether the third data sample is real data, and the second prediction result is used to indicate the second risk category corresponding to the third data sample.

[0163] In this embodiment, when the second risk category is the same as the first risk category, the generative adversarial network can not only be used to generate a real second data sample, but also can identify the first risk category to which the data sample belongs, so that when the generator generates the second data sample, it is closer to the risk characteristics represented by the first label, thereby improving its authenticity.

[0164] Specifically, as shown in Figure 7, multiple third data samples can be input into the discriminator, and the discriminator judges the truth or falsehood of the third data samples to obtain a first prediction result and the second risk category to which the third data samples belong, and obtains a second prediction result. Then, the first loss errorD_real for the discriminator to judge the truth or falsehood can be determined based on the first prediction result and the second label, and the second loss errorD_real_risk for the discriminator to judge the second risk category can be determined based on the second prediction result and the third label. The parameters of the discriminator are updated based on the first loss and the second loss.

[0165] Exemplarily, in the process of training the discriminator, the multiple third data samples input to the discriminator may include real data samples and false data samples, wherein the real data samples can be called real data samples, and the false data samples can be called synthetic data samples. The second label carried by the real data sample represents that the real data sample is real data, and the second label carried by the synthetic data sample represents that the real data sample is false data. Accordingly, since the discriminator not only needs to distinguish the true or false of the data, but also needs to distinguish the second risk category to which the data belongs, in practice, if it is a synthetic data sample, since the third label is not the second risk category in the true sense, it is the second risk category corresponding to the virtual sample object. In order to avoid increasing errors when updating the parameters of the discriminator, when the data input to the discriminator is a synthetic data sample, the error returned to the discriminator may not include the third label and the second prediction result corresponding to the third label.

[0166] Accordingly, when the loss of the discriminator is updated based on the first prediction result, the second label, the second prediction result and the third label, the loss of the discriminator can be updated based on the first prediction result, the second prediction result, the second label, the second prediction result and the third label corresponding to the real data sample; and the parameters of the discriminator can be updated based on the first prediction result and the second label corresponding to the synthetic data sample.

[0167] In this example, when the input to the discriminator is a real data sample, the loss of the discriminator can be updated according to the first prediction result, the second prediction result, the second label, the second prediction result and the third label output by the discriminator. Specifically, the first loss errorD_real for the discriminator to judge true or false can be determined according to the first prediction result and the second label, and the second loss errorD_real_risk for the discriminator to judge the second risk category can be determined according to the second prediction result and the third label. The parameters of the discriminator are updated according to the first loss and the second loss.

[0168] When the input to the discriminator is a synthetic data sample, in order to avoid a large error in the return, the loss of the discriminator can be updated according to the first prediction result and the second label output by the discriminator. Specifically, the first loss errorD_real for the discriminator to judge the truth or falsehood can be determined according to the first prediction result and the second label, and the parameters of the discriminator can be updated according to the first loss.

[0169] Exemplarily, after the discriminator training is completed, the generator needs to be trained. When training the generator, the output results of the discriminator need to be used to update the parameters of the generator. Among them, the noise data sample can include at least one data group, and different data groups are used to simulate the second data samples corresponding to different objects. Each data group can be set with category information. When the noise data sample is input into the generator in the generative adversarial network, the category information can also be input into the generator in the generative adversarial network. The category information can be used as the third label of the predicted data output by the generator, so that the discriminator can predict the second risk category of the data and judge the truth or falsehood of the predicted data. As described above, after judging the second risk category and the truth or falsehood of the predicted data, the discriminator can further determine the loss based on the second label and the third label to update the generator. Of course, since the predicted data input to the discriminator is false data, when updating, it is also possible not to refer to the third label and the second risk category (second prediction result) corresponding to the predicted data, thereby avoiding increasing the error.

[0170] Accordingly, referring to Figure 8, another flow chart for training the discriminator is shown. As shown in Figure 8, when the generative adversarial network is trained using noise data samples, the noise data samples and the category information corresponding to the noise data samples can be input into the generator, and the predicted data corresponding to the noise data samples output by the generator can be input into the discriminator; and the first prediction result and the second prediction result output by the discriminator are obtained; then, the parameters of the discriminator can be updated based on the first prediction result.

[0171] The first prediction result is used to indicate whether the predicted data is real data, and the second prediction result is used to indicate the second risk category corresponding to the predicted data.

[0172] Referring to FIG9 , a schematic diagram of the process of training the generator is shown. As shown in FIG9 , a noise data sample can be input into the generator, and the generator outputs prediction data corresponding to the noise data sample. Then, the prediction data is input into the discriminator, and the discriminator predicts the truth or falsity of the prediction data to obtain a first prediction result, and predicts the second risk list to which the prediction data belongs, a second prediction result. Among them, the third loss errorD_fake can be determined based on the distance between the first prediction result and the second label representing the real data, and the fourth loss errorD_fake_risk can be determined based on the second prediction result and the third label carried by the noise data sample; then, the parameters of the generator can be updated according to the third loss errorD_fake, or the parameters of the generator can be updated according to the third loss and the fourth loss, as shown in FIG9 . The update path shown by the dotted line in the figure is an optional strategy.

[0173] Among them, when the parameters of the generator are updated, the parameters of the discriminator can be fixed.

[0174] As another example, when the generator and the discriminator in the generative adversarial network are trained simultaneously using multiple third data samples and noise data samples, the training process can also be described as follows:

[0175] First, multiple third data samples are input into the discriminator in the generative adversarial network, noise data samples are input into the generator in the generative adversarial network, and prediction data output by the generator is input into the discriminator;

[0176] Next, the discriminator determines whether the input prediction data and the third data sample are true or false (a first prediction result), and determines the second risk category to which the prediction data belongs and the second risk category to which the third data sample belongs (a second prediction result);

[0177] Afterwards, the discriminator and the generator are updated separately. When updating the parameters of the discriminator, a first loss can be determined based on the first prediction result and the second label corresponding to the real data sample in the third data sample, and a second loss can be determined based on the second prediction result and the third label corresponding to the real data sample, and the parameters of the discriminator are updated based on the first loss and the second loss; and, a first loss can be determined based on the first prediction result and the second label corresponding to the synthetic data sample, and the parameters of the discriminator are updated based on the first loss; and, a third loss can be determined based on the first prediction result corresponding to the predicted data, and the parameters of the discriminator are updated based on the third loss.

[0178] When updating the parameters of the generator, a third loss may be determined based on the first prediction result corresponding to the prediction data, and the parameters of the generator may be updated based on the third loss.

[0179] Among them, when updating the parameters of the discriminator, the parameters of the generator are fixed, and when updating the parameters of the generator, the parameters of the discriminator can be fixed.

[0180] 10a-11b , schematic diagrams are shown of the results of the risk determination model by the discriminator in the generative adversarial network in two situations, namely: the discriminator not only distinguishes true from false but also distinguishes the second risk category, and the discriminator only distinguishes true from false.

[0181] As shown in Figures 10a and 10b, Figure 10a shows a schematic diagram of the ROC results of the risk determination model trained using the second data samples generated by the first generative adversarial network on the test set, and Figure 10b shows a schematic diagram of the training process of the first generative adversarial network. The discriminator in the first generative adversarial network only distinguishes true from false. Taking the first risk category as the high and low prognostic risk categories as an example, the first 100 second data samples generated by the generator are combined with 223 first data samples to construct the test set and training set. The test set includes a total of 21 data samples, including 19 Class 1 (high risk) data and 2 Class 0 (low risk) data. The risk determination model is trained using the training set and tested using the test set. As shown in Figure 10a, the results are: training set ACC = 0.9868, test set ACC = 0.955, test set AUC = 0.975, and TP_count = 20, TN_count = 1, FP_count = 1, and FN_count = 0.

[0182] As shown in Figures 11a and 11b, Figure 11a shows a schematic diagram of the ROC results of the risk determination model trained using the second data sample generated by the second generative adversarial network under the test set, and Figure 11b shows a schematic diagram of the training process using the second generative adversarial network. The discriminator in the second generative adversarial network not only distinguishes true from false but also distinguishes the second risk category. Taking the first risk category as the high and low prognostic risk categories as an example, the first 100 second data samples generated by the generator are combined with 223 first data samples to construct a test set and a training set. The test set includes a total of 21 data samples, including 19 Class 1 (high risk) data and 2 Class 0 (low risk) data. The risk determination model is obtained by training with the training set, as shown in Figure 11a, and the results are: training set ACC = 0.9912, test set ACC = 0.955, test set AUC = 0.975, and TP_count = 19, TN_count = 2, FP_count = 0, and FN_count = 1.

[0183] Figures 10a and 11a show that compared to a discriminator that only distinguishes true from false, if it also distinguishes the second risk category, it can effectively improve the recognition accuracy of the small sample class. As shown in Figures 10b and 11b, when the discriminator not only distinguishes true from false but also distinguishes high and low risk, the discriminator in the generative adversarial network is more oscillatory.

[0184] In this embodiment B, during the training of the generative adversarial network, some neurons in the generative adversarial network may be randomly inactivated. Specifically, some neurons may be randomly inactivated during both the generator and discriminator training stages, or some neurons may be randomly inactivated during the generator training stage while not randomly inactivated during the discriminator training stage.

[0185] Accordingly, the generative adversarial network includes multiple network layers, each network layer includes at least one neuron. During the training of the generative adversarial network, some target neurons among the multiple neurons can be inactivated according to a preset ratio in at least one training session.

[0186] Among them, the architecture of the generative adversarial network model is mainly composed of a bidirectional long short-term memory network (Bi_LSTM) and a linear fully connected network (Linear fully connection), and layer normalization (LayerNorm) is also added. Referring to Figure 12, a schematic diagram of the network structure of a generative adversarial network is shown. As shown in Figure 12, the discriminator and the generator both include a bidirectional long short-term memory network (Bi_LSTM), wherein the bidirectional long short-term memory network in the discriminator is 64-dimensional, and the bidirectional long short-term memory network of the generator is 128-dimensional.

[0187] Among them, the discriminator also includes layer normalization (LayerNorm) connected to the bidirectional long short-term memory network, a View layer connected to layer normalization (LayerNorm), a 128-dimensional linear layer connected to the View layer, an activation function layer leaky_relu connected to the linear layer, a random dropout layer dropout connected to the activation function layer leaky_relu, and a sigmoid layer connected to the random dropout layer.

[0188] Among them, the generator also includes a 256-dimensional layer normalization (LayerNorm) connected to the bidirectional long short-term memory network, a View layer connected to the layer normalization (LayerNorm), a 256-dimensional linear layer connected to the View layer, a 256-dimensional layer normalization (LayerNorm) connected to the linear layer, an activation function layer leaky_relu connected to the 256-dimensional layer normalization (LayerNorm), a 256-dimensional linear layer connected to the activation function layer leaky_relu, and a 72-dimensional output layer connected to the 256-dimensional linear layer.

[0189] In this example, layer normalization (LayerNorm) and random packet loss (dropout) in the discriminator can prevent the model from over-learning. In addition, the activation function layer in the discriminator and generator uses leaky_relu as the activation function, which is shown in the following formula (1):

[0190] The advantage of the LeakyReLU method is that during the back-propagation process, the gradient can also be calculated for the part of the LeakyReLU activation function input that is less than zero, thus avoiding the gradient vanishing problem.

[0191] The generative adversarial network includes multiple network layers, each of which contains at least one neuron. A neuron can be understood as an operator in the network layer. During the training process, at least one operator can be inactivated by random packet loss, for example, so that the operator does not participate in the calculation.

[0192] Part of the neurons may be inactivated during at least one training session. Specifically, part of the neurons may be inactivated during each training session. Alternatively, neurons may not be inactivated at the initial stage of training, but may be inactivated during the middle stage of training and not inactivated during the late stage of training. Alternatively, neurons may be inactivated at the initial stage of training and not inactivated during the middle and late stages of training. Alternatively, neurons may not be inactivated during the initial stage of training and not inactivated during the middle stage of training. In specific implementation, neurons may not be inactivated during the initial stage of training and may be inactivated during the middle stage of training and late stage of training.

[0193] Specifically, in each neuron inactivation process, some target neurons among multiple neurons can be inactivated according to a preset ratio; for example, in each training, some target neurons can be inactivated according to the same preset ratio, such as inactivating some target neurons at a ratio of 0.2, that is, 20% of the neurons in the network are inactivated.

[0194] For example, in different training sessions, some target neurons can be inactivated at different preset ratios. In the previous training session, some target neurons were inactivated at a ratio of 0.2, while in the next training session, some target neurons were inactivated at a ratio of 21%. In this example, the ratio of inactivated target neurons can be within a preset range, such as 15% to 30%. During training, the ratio of inactivated target neurons fluctuates within this range.

[0195] For example, in at least one training session, when performing inactivation processing on some neurons among the plurality of neurons according to a preset ratio, at least one of the following methods may be used:

[0196] Method 1: In N consecutive training sessions, the neurons that are inactivated are different;

[0197] Mode 2: In N consecutive training sessions, the neurons that are inactivated are partially repeated;

[0198] Mode 3: In N consecutive trainings, the neurons subjected to the inactivation treatment are in different network layers;

[0199] Mode 4: In N consecutive trainings, the network layer where the neurons subjected to the inactivation processing are located is repeated.

[0200] Wherein, N is an integer greater than or equal to 2.

[0201] In the first approach, different target neurons are inactivated during N consecutive training sessions. For example, different target neurons are inactivated during two consecutive training sessions.

[0202] In the second approach, during N consecutive training sessions, the target neurons that are inactivated are partially repeated. For example, during N consecutive training sessions, the same target neuron may be inactivated at least twice.

[0203] In approach 3, the target neurons to be inactivated during N consecutive training cycles can be located in different network layers. For example, during N consecutive training cycles, target neurons in different network layers are inactivated between two consecutive training cycles. As shown in Figure 12, during the nth training cycle, target neurons in the linear layer and the view layer can be inactivated, while during the n+1th training cycle, target neurons in the activation function layer and the 256-dimensional linear layer can be inactivated.

[0204] In method 4, during N consecutive trainings, neurons in the same network layer may be inactivated in at least two trainings. As shown in FIG12 , during the nth training and the n+1th training, all neurons in the 256-dimensional linear layer may be inactivated.

[0205] Among them, method 3 and method 1 can be combined. In this way, during N consecutive trainings, the target neurons to be inactivated can be located in different network layers, and the target neurons to be inactivated are also different.

[0206] Among them, method 2 and method 4 can be combined, so that during N consecutive trainings, the same neuron in the same network layer can be inactivated at least twice.

[0207] Specifically, any combination of methods 1 to 4 can be used to inactivate different target neurons in the same network layer in different training sessions, or the same target neuron in the same network layer inactivated in different training sessions, or target neurons in different network layers inactivated in different training sessions.

[0208] Exemplarily, the deactivation process performed on the target neuron may include setting the weight parameter corresponding to the target neuron to 0, or fixing the weight parameter of the target neuron at that time.

[0209] In this example, when the weight parameter corresponding to the target neuron is set to 0, the target neuron can be prevented from participating in the current calculation, so that the target neuron does not participate in the current training process; fixing the weight parameter of the target neuron in the previous training of the current training can mean: when the network parameters are updated in the current training, the weight parameter of the target neuron can be not updated in the current training, so that the weight parameter of the target neuron after this training is still the weight parameter updated in the previous training, so that the target neuron does not transmit the training error of this time.

[0210] As described above, the second number of second data samples is smaller than the first number of first data samples. In this embodiment B, noise data needs to be input into the generative adversarial network to obtain multiple second data samples. For example, the generative adversarial network can output a large number of candidate data for screening based on the noise data. After being discriminated by the discriminator, these candidate data have a predicted probability indicating that the candidate data is real data, thereby screening out multiple second data samples from the multiple candidate data based on the predicted probability.

[0211] Accordingly, the noise data is input into the trained generative adversarial network to obtain multiple candidate data output by the generator, as well as the prediction probability corresponding to each candidate data, where the prediction probability is the probability that the candidate data is the true data; and based on the prediction probability, multiple second data samples are screened out from the multiple candidate data.

[0212] In this embodiment, the predicted probability can be similar to the first prediction result, representing the truth or falsehood of the input data, wherein the higher the predicted probability, the closer the candidate data is to the real data, and the lower the predicted probability, the farther the candidate data is from the real data. Thus, multiple candidate data can be sorted according to the predicted probability, and multiple second data samples can be screened out according to the sorting results. For example, the first data samples are sorted in descending order according to the predicted probability, thereby screening out a preset number of second data samples with a higher ranking, or the second data samples are sorted in descending order according to the predicted probability. The preset number can be determined according to the ratio between the number of the first data samples and the number of the first data samples as described in the above embodiment. For example, if there are 200 first data samples and the ratio of the number of the first data samples to the number of the second data samples is 2, then 100 second data samples need to be screened out.

[0213] In one example of this embodiment B, the risk determination model can be applied in the medical field to predict the prognostic risk of a disease after a gene or a site on a gene mutates. When the risk determination model is used to predict the prognostic risk of a disease at a site on a gene, the third data sample used to train the generative adversarial network may include a clinical data sample obtained for a high-frequency mutation site on the gene, i.e., a real data sample. Accordingly, multiple sample objects may originate from the same individual, and the multiple third data samples include at least one real data sample. The target probability corresponding to the sample object to which the real data sample belongs is less than or equal to 0.5, and the target probability represents the statistical significance of the sample object in the target direction in causing the target risk to the individual.

[0214] In this example, if applied to prognostic risk prediction, the individual can refer to the gene, and the sample object can be the site in the gene; if applied to natural disaster risk prediction, the individual can be region H, and the sample object can be each monitoring location in region H; if applied to prognostic survival prediction, the individual can be the disease, and the sample object can be the gene related to the disease.

[0215] Among them, the real data sample is an objectively existing data sample, which corresponds to a real sample object. For example, in the prediction of prognostic risk, the target direction refers to the mutation direction of the site, and the target risk is the prognostic risk affected by the site mutation. Then, if the target probability corresponding to the sample object to which the real data sample belongs is less than or equal to 0.5, it means that the data sample of the sample object is statistically significant for evaluating the prognostic risk after the site mutation. Similarly, it applies to prognostic survival. In prognostic survival, the individual can be a disease, and the sample object can be a gene related to the disease. Then, if the target probability corresponding to the sample object to which the real data sample belongs is less than or equal to 0.5, it means that the data sample of the sample object (gene) is statistically significant for evaluating the effect of gene mutation on the prognostic survival of the disease.

[0216] For example, in the risk prediction of natural disasters, the target direction is heavy rain attack, the target risk is the risk of mudslides caused by heavy rain attack, the individual can be region H, and the sample object can be each monitoring location in region H. Then, the target probability corresponding to the sample object to which the real data sample belongs is less than or equal to 0.5, indicating that the data sample of the sample object is statistically significant for the risk of mudslides in region H.

[0217] Taking prognostic risk as an example, if individuals represent genes and samples represent loci within the genes, then the real data can be data on high-frequency mutation sites within the gene. In practice, missense mutation data and clinical data for the gene can be obtained from the publicly available ICGC and MSK databases to form the mutation group data set, while data on the gene without mutations serve as the control group data set. Then, using the Cox regression method, the prognostic risk of each amino acid mutation site in the gene can be determined. This is categorized based on the statistical P value, which measures whether the observed phenomenon in a statistical hypothesis test is a random event. A P value ≤ 0.05 indicates that the locus has a statistically significant effect on the prognostic risk, while a P value > 0.05 indicates that the locus has no statistically significant effect on the prognostic risk (a random event). Next, at least one data factor from the statistically significant locus can be used as the real data sample.

[0218] In some embodiments, the risk determination model can be applied to the medical field, where the target direction can include gene mutation direction, and the target risk can include prognostic risk; alternatively, the target direction can include site mutation direction, and the target risk can include prognostic risk. Alternatively, the target direction can include gene mutation direction, and the target risk can include prognostic survival; alternatively, the target direction can include site mutation direction, and the target risk can include prognostic survival.

[0219] For example, in predicting the prognostic risk after a gene or a site mutation on a gene, the prognostic risk can be predicted for the site mutation in the gene, and the sample object may include the site in the gene. The multiple first data samples may include: data samples of the first category of sample objects and data samples of the second category of objects; wherein the mutation frequency corresponding to the first category of sample objects is higher than the mutation frequency of the second category of objects.

[0220] Among them, the first type of sample object may refer to the high-frequency mutation site in the gene, indicating that the mutation frequency of the sample object is high, while the second type of sample object may refer to the low-frequency mutation site in the gene, indicating that the mutation frequency of the sample object is low. In practice, there is less clinical data for low-frequency mutation sites, and the number of samples is small. Therefore, it is necessary to virtualize low-frequency mutation sites to increase the data samples of low-frequency mutation sites, thereby improving the diversity of training samples in the training set of the risk determination model.

[0221] The following is an illustrative example of the training method of the model disclosed in the present invention, using two examples from different application scenarios:

[0222] Example #1 is applied in the medical field, specifically, to predict the prognosis risk of a disease after predicting a gene or site mutation on a gene. Referring to Figure 13, taking the TP53 gene as an example, a schematic diagram of the process of applying the model training method in the medical field is shown, as shown in Figure 13:

[0223] S1: Obtain real data samples of high-frequency mutation sites in the TP53 gene;

[0224] Missense mutation data and clinical data of the TP53 gene were obtained from the public ICGC and MSK databases to form the mutation group data. The data of the TP53 gene without mutation served as the control group data. Then, a statistical analysis method (Cox regression method) was used to obtain the prognostic risk of mutation at each amino acid site of the gene.

[0225] The classification is based on the statistical P value, which is used to measure whether the phenomenon observed in the statistical hypothesis test is a random event. A P value ≤ 0.05 indicates that the prognostic risk of the site is statistically significant, and a P value > 0.05 indicates that the prognostic risk of the site is not statistically significant (a random event). Then, the prognostic risk of sites with statistical significance is used as the real data, and sites without statistical significance are used as sites for which risk needs to be predicted. The distribution of statistical mutation sites of TP53 across the entire amino acid chain can be shown in Figure 14, and the distribution of the obtained special data (i.e., real data samples) can be shown in Figure 15.

[0226] In Figure 14, the dashed boxes indicate sites with P values ​​less than 0.05, while sites outside the dashed boxes all have P values ​​greater than 0.05. The bar chart in the lower half of Figure 14 shows the distribution of TP53 missense mutations at 393 amino acid positions, with the right Y-axis representing sample size. It can be seen that TP53 missense mutations are primarily concentrated in the 100-281 and 326-335 intervals, while most sites outside these intervals have relatively low mutation frequencies. The left Y-axis represents the logarithm of the prognostic risk of the site. A value greater than 0 indicates that the mutation at that site is risky; a value less than 0 indicates that the mutation at that site is beneficial. Therefore, sites with significant prognostic risk are labeled according to this criterion, with a high risk value of 1 and a low risk value of 0.

[0227] Missense mutations in genes can affect protein structure and function. Using 20 online tools, protein impact scores can be derived for missense mutations, serving as a statistical factor for the prognostic risk of a site affecting a disease. These include SIFT_score and Polyphen2_HDIV_score, which measure the impact of point mutations on protein structure / function. For example, SIFT_score values ​​range from 0 to 1. Lower values ​​indicate a deleterious variant, increasing the likelihood of altering protein function; conversely, higher values ​​indicate a lower impact. Clinical guidelines also rank the impact of some high-frequency missense mutations on a scale of {2, 1, 0, -1}, where 2 indicates strong clinically significant evidence of carcinogenesis (FDA-approved or investigational evidence), 1 indicates supportive clinically significant evidence of carcinogenesis, 0 indicates no clinically significant evidence of carcinogenesis, and -1 indicates the variant is likely benign or neutral. Therefore, using the annotation tool Annovar, clinical evidence characteristics can be derived. In addition, location information can also be used as part of the feature. Combining the above feature data and the corresponding prognostic risk value, each data factor in the third data sample is obtained, as shown in FIG15 , which exemplarily shows a distribution diagram of the data factors.

[0228] As shown in Figure 15, the data shown are protein impact score features of the variant sites, such as SIFT_score, Polyphen2_HDIV_score, Polyphen2_HVAR_score, LRT_score, MutationTaster_score, MutationAssessor_score, FATHMM_score, PROVEAN_score, VEST3_score, CADD_raw, CADD_phred, DANN_score, fathmm-MKL_coding_score, MetaSVM_score, MetaLR_score, integrated_fitCons_score, integrated_confidence_value, GERP++_RS, etc.

[0229] Among them, the real data sample is marked with a second label to indicate that it is real data, and is marked with a first label to indicate the first risk category to which it belongs, that is, whether the mutated site has a high risk or a low risk for prognosis.

[0230] S2: prepare noise data samples;

[0231] The noise data is generated using the normal distribution method. The noise data obeys a normal distribution with a mean of 0 and a variance of 1. The noise data is randomly sampled from the normal distribution.

[0232] S3: training a generative adversarial network;

[0233] S31: training the discriminator;

[0234] The real data sample is input into the discriminator, and the discriminator outputs 2-dimensional output data, where the first dimension is the first prediction result, which represents the probability that the discriminator determines whether the input data is true or false (greater than 0.5 indicates real data, less than 0.5 indicates false data), and the error is represented by errorD_real; the second dimension is the second prediction result, which represents the first risk category corresponding to the real data sample input by the discriminator, that is, whether the prognosis is harmful or beneficial, and the error is represented by errorD_real_risk.

[0235] The noise data samples are input into the generator in the generative adversarial network. The generator fits the distribution of real site data to generate synthetic data. Specifically, when the conditional generative adversarial network fits the characteristic data distribution of the mutation site, according to the risk value HR and P value of each site calculated by Cox risk analysis of site variation, we regard the mutation sites with p <= 0.05 as the real sites, simulate the characteristic data distribution of these sites, and thus generate synthetic data samples based on the noise data samples;

[0236] The synthetic data sample fake_data is fed into the discriminator, which similarly determines the authenticity of the input data (errorD_fake) and its riskiness (errorD_fake_risk). During the discriminator training process, the error errorD_real + errorD_fake + errorD_real_risk is used as the total error and fed back to the discriminator network to update the model parameters. Since the synthetic data and real data have the same distribution but subtle differences, and their labels may not be the same, to avoid increasing the error, the errorD_fake_risk is not included in the returned error.

[0237] S32: training generator;

[0238] In the process of inputting the noise data sample into the generator, the synthetic data sample output by the generator is obtained, and the synthetic data sample is input into the discriminator to determine whether it is real data or synthetic data. The error is represented by errorG, and the error is fed back to the generator to update the parameters of the generator, thereby obtaining a generative adversarial network.

[0239] In this way, the generator can be trained while training the discriminator.

[0240] Among them, the discriminator simultaneously judges the authenticity and risk level, and the two loss functions used are both cross entropy loss functions, which can be expressed by the following formula (2): oss = -1 / n∑ i (t[i]*log(o[i])+(1-t[i]log(1-o[i])) Formula (2)

[0241] Where n represents the number of samples, t[i] represents the label of the i-th sample, and o[i] represents the model prediction value of the i-th sample.

[0242] During the training process, the random dropout method is used to avoid over-learning of the model. The leaky_relu method is also used as the activation function, which is expressed as follows:

[0243] The advantage of the LeakyReLU method is that during the back-propagation process, the gradient can also be calculated for the part of the LeakyReLU activation function input that is less than zero, thus avoiding the gradient vanishing problem.

[0244] S4: Prepare a first data sample. The process of obtaining the first data sample can refer to the process of obtaining the real data sample mentioned above.

[0245] S5: Prepare the second data sample:

[0246] The noise data and the category information representing the high and low prognostic risk are input into the generative adversarial network after the S3 step training is completed, and multiple candidate data output by the generative adversarial network and the prediction probability corresponding to the first prediction result output by each candidate data are obtained. In this way, multiple first candidate data with the label of 1, that is, the first risk category is high risk, and multiple second candidate data with the label of 0, that is, the first risk category is low risk can be obtained.

[0247] Suppose that 50,000 noise data points labeled 1 (indicating a high prognostic risk) are fed into the generator to generate candidate data. This candidate data point is then fed into the discriminator to determine if it is true or false. The discriminant probability is then sorted from highest to lowest, yielding the top 50 (top 100) second data samples with the highest probability (indicating that the discriminator believes the input data is true). Similarly, using noise data labeled 0 (indicating a low prognostic risk) as input into the generator to generate candidate data points, this candidate data point is then fed into the discriminator to yield the top 50 / top 100 second data samples with the highest probability.

[0248] S6: Train the preset model to obtain the risk determination model:

[0249] 200 first data samples and 100 second data samples are used as training sets and input into the preset model to determine the corresponding first risk categories. Then, based on the first labels carried by the first data samples and the second data samples, and the first risk categories output by the preset model, the loss value is calculated. Then, based on the loss value, the parameters of the preset model are updated to obtain a risk determination model. The risk determination model can be used to predict the risk corresponding to the input data. For example, if the input is the characteristic distribution data of a certain site, the prognostic risk of the disease caused by the mutation of the site can be obtained.

[0250] Example #2 is applied in the medical field, specifically, to predicting the prognosis and survival of a disease after predicting a gene or site mutation on a gene. Referring to Figure 16 , taking the TP53 gene as an example, a schematic diagram of the process of applying the model training method to another medical field is shown, as shown in Figure 16:

[0251] S1': Obtain a real data sample of a high-frequency mutation site in the TP53 gene. The real data sample carries a third label, which represents the second risk category to which the real data sample belongs. The second risk category represents the prognostic risk of the disease caused by the gene mutation.

[0252] The process of obtaining real data samples can refer to the process in Example #1 above, and will not be repeated here.

[0253] S2': prepare noise data samples;

[0254] The noise data is generated using the normal distribution method. The noise data obeys a normal distribution with a mean of 0 and a variance of 1. The noise data is randomly sampled from the normal distribution.

[0255] S3': training a generative adversarial network;

[0256] S31': training the discriminator;

[0257] The real data sample is input into the discriminator, and the discriminator outputs 2-dimensional output data, where the first dimension is the first prediction result, which represents the probability that the discriminator determines whether the input data is true or false (greater than 0.5 indicates real data, less than 0.5 indicates false data), and the error is represented by errorD_real; the second dimension is the second prediction result, which represents the second risk category corresponding to the real data sample input by the discriminator, that is, whether the prognosis is harmful or beneficial, and the error is represented by errorD_real_risk.

[0258] The noise data samples are input into the generator in the generative adversarial network. The generator fits the distribution of real site data to generate synthetic data. Specifically, when the conditional generative adversarial network fits the characteristic data distribution of the mutation site, according to the risk value HR and P value of each site calculated by Cox risk analysis of site variation, we regard the mutation sites with p <= 0.05 as the real sites, simulate the characteristic data distribution of these sites, and thus generate synthetic data samples based on the noise data samples;

[0259] The synthetic data sample fake_data is fed into the discriminator, which similarly determines the authenticity of the input data (errorD_fake) and its riskiness (errorD_fake_risk). During the discriminator training process, the error errorD_real + errorD_fake + errorD_real_risk is used as the total error and fed back to the discriminator network to update the model parameters. Since the synthetic data and real data have the same distribution but subtle differences, and their labels may not be the same, to avoid increasing the error, the errorD_fake_risk is not included in the returned error.

[0260] S32': training generator;

[0261] In the process of inputting the noise data sample into the generator, the synthetic data sample output by the generator is obtained, and the synthetic data sample is input into the discriminator to determine whether it is real data or synthetic data. The error is represented by errorG, and the error is fed back to the generator to update the parameters of the generator, thereby obtaining a generative adversarial network.

[0262] In this way, the generator can be trained while training the discriminator.

[0263] S4': prepare the first data sample;

[0264] Collect clinical data corresponding to multiple users with mutations at the same site, the clinical data including multiple data factors corresponding to the mutation site and the second risk category corresponding to the first data sample; wherein the first data sample carries a first label, and the first label represents the actual prognosis survival period,

[0265] S5': prepare a second data sample;

[0266] The noise data and the category information representing the high and low prognostic risk are input into the generative adversarial network after the training in step S3', and multiple candidate data output by the generative adversarial network and the prediction probability corresponding to the first prediction result output by each candidate data are obtained. In this way, multiple first candidate data with the label of 1, that is, the second risk category is high risk, and multiple second candidate data with the label of 0, that is, the second risk category is low risk can be obtained.

[0267] Suppose that 50,000 noise data points labeled 1 (indicating a high prognostic risk) are fed into the generator to generate candidate data. This candidate data point is then fed into the discriminator to determine if it is true or false. The discriminant probability is then sorted from highest to lowest, yielding the top 50 (top 100) second data samples with the highest probability (indicating that the discriminator believes the input data is true). Similarly, using noise data labeled 0 (indicating a low prognostic risk) as input into the generator to generate candidate data points, this candidate data point is then fed into the discriminator to yield the top 50 / top 100 second data samples with the highest probability.

[0268] According to the second risk category corresponding to the second data sample, the prognostic survival period corresponding to the second data sample is set, that is, the first label is set.

[0269] S6': Train the preset model to obtain the risk determination model

[0270] 200 first data samples and 100 second data samples are used as training sets and input into the preset model to determine the corresponding second risk categories. Then, based on the first labels carried by the first data samples and the second data samples, and the first risk categories output by the preset model, the loss value is calculated. Then, based on the loss value, the parameters of the preset model are updated to obtain a risk determination model. The risk determination model can be used to predict the prognosis survival period corresponding to the input data. For example, if the input is the characteristic distribution data of a certain site, the prognosis survival period of the disease when the site mutates can be obtained.

[0271] In summary, the model training method disclosed in this disclosure has the following advantages:

[0272] First, since the second data sample is simulated based on noise data, the second data sample can be used to expand the first data sample, so that sample objects with few data samples can also be expanded through the second data sample, thereby improving the diversity and richness of training samples and avoiding the problem of model overfitting caused by insufficient training samples.

[0273] Second, the second data sample is generated by the generator in the trained generative adversarial network, and the generative adversarial network is trained using the third data sample and the noise data sample. Therefore, the authenticity of the second data sample is improved through the generative adversarial network. Therefore, when the preset model is trained with more realistic second data samples and first data samples, the classification accuracy of the risk determination model in the inference stage can be improved.

[0274] Third, since the discriminator in the generative adversarial network not only determines whether the data is true or false when training the generative adversarial network, but is also used to determine the second risk category to which the data belongs, the second data sample generated by the generative adversarial network can be more consistent with the second risk category. Since the second data sample can be data that simulates a small sample class, when the authenticity of the second data sample is improved, the trained risk determination model can effectively improve the recognition accuracy of the risk determination model for the small sample class.

[0275] Fourth, since some target neurons in the generative adversarial network can be randomly inactivated during training, over-learning of the generative adversarial network can be avoided.

[0276] Fifth, since the category information corresponding to the noise data is also input when the noise data is input into the generative adversarial network to generate the second data sample, the generated second data sample can fit the characteristics of the second risk category represented by the category information, thereby improving the authenticity of the second data sample.

[0277] Based on the same inventive concept, the present disclosure further provides a risk determination method. Referring to FIG. 17 , a schematic flow chart of the steps of the risk determination method is shown. As shown in FIG. 17 , the method may specifically include the following steps:

[0278] Step S201: Acquire characteristic distribution data of a target object to be classified; wherein the characteristic distribution data includes at least one data factor affecting the target risk of the sample object in the target direction;

[0279] Step S201: inputting the characteristic distribution data into a risk determination model to obtain a first risk category corresponding to the target risk of the target object;

[0280] Wherein, the risk determination model is obtained based on the training method of the model described above.

[0281] As described above, in some embodiments, the risk determination model can be applied to the medical field, where the target direction can include the direction of gene mutation, and the target risk includes the prognostic risk, or the target direction can include the direction of site mutation, and the target risk includes the prognostic risk. The first risk category can then be a high or low prognostic risk category, characterizing the high or low prognostic risk of a disease when a gene mutates, or the high or low prognostic risk of a disease when a site mutates; wherein the characteristic distribution data can be the characteristic distribution data of the site, for example, the characteristic distribution data shown in FIG15 ; or, the characteristic distribution data can be the characteristic distribution data of a gene, and the characteristic distribution data of the gene can refer to the characteristic distribution data shown in FIG15 .

[0282] In some other embodiments, the risk determination model can be applied to the medical field, where the target direction may include gene mutation direction, and the target risk includes prognostic survival; alternatively, the target direction includes site mutation direction, and the target risk includes prognostic survival. The first risk category can then be a prognostic survival category, characterizing the prognostic survival corresponding to a gene mutation, or the prognostic survival corresponding to a site mutation; wherein the target object can be a gene or a site, and the characteristic distribution data can be the characteristic distribution data of the site, or can be the characteristic distribution data of the gene. In this example, the characteristic distribution data can include clinical data of the corresponding disease, such as the patient's gender, age, medical history, and medication history, thereby enabling the prediction of prognostic survival based on the characteristic distribution data.

[0283] In some other embodiments, the risk determination model can be applied to the field of natural disaster prediction, where the target direction can be a rainstorm attack, and the target risk can be the risk of debris flow after the rainstorm attack. The first risk category can be a high or low debris flow risk, where the target object can be region H, and the characteristic distribution data can be the characteristic distribution data of each monitoring location in region H. In this example, the characteristic distribution data can include data such as the softness of the soil at the monitoring location, the composition of the soil, the slope gradient, and the rock structure.

[0284] Among them, the risk determination model can be obtained by referring to the process described in the embodiment of the training method of the above model, for example, the process shown in the above example #1 and example #2.

[0285] For example, following Example #1 above, the risk determination method of the present disclosure is exemplified as follows:

[0286] After step S6 in the above example #1, a risk determination model can be obtained. Then, for the site T to be predicted (target object), specifically, site T can be a site in any gene other than the TP53 gene, or can be any mutation site in the TP53 gene. The characteristic distribution data of the site T to be predicted is obtained through the process of step S1. Then, the characteristic distribution data is input into the risk determination model to obtain the prognostic risk of the mutation at the site T.

[0287] For example, following Example #2 above, the risk determination method of the present disclosure is exemplified as follows:

[0288] After step S6' in the above example #1, a risk determination model can be obtained. Then, for the site T to be predicted, similarly, site T can be a site in any gene other than the TP53 gene, or can be any mutation site in the TP53 gene. The characteristic distribution data of the site T to be predicted is obtained through the process of step S1. The characteristic distribution data can also be obtained from clinical data of patient P for the corresponding disease, such as age, gender, medication history and medical history. Then, the characteristic distribution data is input into the risk determination model to obtain the prognostic survival result of the site T mutation for patient P.

[0289] By adopting the risk determination method of the embodiment of the present disclosure, since the risk determination model is obtained by training based on the second data sample and the first data sample, the problem of model overfitting caused by insufficient training samples is avoided during the training process, so that the generalization of the risk determination model is enhanced, so that the first risk category of multiple objects can be predicted, thereby improving the accuracy of predicting the first risk category of multiple objects, especially small sample objects.

[0290] Furthermore, since the second data sample is generated by the generator in the trained generative adversarial network, and the generative adversarial network is trained using the third data sample and the noise data sample, the authenticity of the second data sample is improved through the generative adversarial network, and the risk determination model obtained by training with the more realistic second data sample and the first data sample has a higher accuracy, thereby improving the classification accuracy of the target object.

[0291] Based on the same inventive concept, the present disclosure also provides a model training device. FIG18 shows a schematic diagram of the framework structure of the model training device. As shown in FIG18 , the model training device may include the following modules:

[0292] A first acquisition module is configured to acquire a first data sample corresponding to each of a plurality of sample objects, wherein the first data sample includes at least one data factor, and the data factor affects a target risk of the sample object in a target direction;

[0293] A second acquisition module is configured to simulate at least one data factor affecting the target risk of the sample object based on the noise data to obtain a plurality of second data samples;

[0294] A training module is used to train a preset model based on the first data sample and the second data sample to obtain a risk determination model; wherein the risk determination model is used to determine a first risk category corresponding to the target direction of the sample object.

[0295] Exemplarily, the first data sample and the second data sample both carry their own corresponding first labels, where the first labels are used to indicate the first risk category corresponding to the sample object and to supervise the training.

[0296] Exemplarily, a first number of the first data samples is greater than a second number of the second data samples, and a ratio of the first number to the second number is greater than or equal to 1.5 and less than or equal to 3.

[0297] Exemplarily, the noise data includes at least one data group, and different data groups are used to simulate second data samples corresponding to different sample objects; the second acquisition module includes:

[0298] a category information acquiring unit, configured to acquire category information corresponding to each of the data groups, the category information being used to indicate a second risk category of the data group in the target direction;

[0299] a generating unit, configured to generate, based on at least one of the data groups and the corresponding category information, at least one second data sample that meets the risk category indicated by the category information;

[0300] The second risk category is the same as or different from the first risk category.

[0301] Exemplarily, the first risk category is used to indicate the risk level of the sample object in the target direction, and the second risk category is the same as the first risk category;

[0302] Alternatively, the first risk category is used to indicate a controllable range after a risk occurs to the sample object in the target direction, and the second risk category is used to indicate a high or low risk of the sample object in the target direction.

[0303] Exemplarily, the first data sample includes clinical data and mutation genomic data, the first risk category includes at least one of a prognostic risk category, a gene mutation category, and a prognostic survival category; the second risk category includes at least one of a prognostic risk category, a gene mutation category, and a prognostic survival category.

[0304] Exemplarily, the noise data includes at least one data group, and different data groups are used to simulate characteristic data corresponding to different objects; the noise data is generated by the following steps:

[0305] Generate basic data that conforms to normal distribution based on target commands;

[0306] Performing at least one round of sampling on the basic data to obtain data groups corresponding to the at least one round of sampling; wherein, in each round of sampling, data corresponding to at least one description dimension is sampled from the basic data;

[0307] Different description dimensions correspond to different influencing factors of the object related to the target direction.

[0308] Exemplarily, the second acquisition module is specifically configured to input the noise data into a generator to obtain a plurality of second data samples output by the generator; wherein the generator is obtained by training a generative adversarial network using a plurality of third data samples and noise data samples;

[0309] In which, the third data sample carries a second label representing whether the data is real data, the generator in the generative adversarial network is used to generate predicted data based on the noise data sample, and the discriminator in the generative adversarial network is used to determine whether the predicted data is real data.

[0310] Exemplarily, the noise data includes at least one data group, and different data groups are used to simulate feature data corresponding to different objects; the second acquisition module is specifically configured to: input at least one of the data groups and corresponding category information into the generator, and obtain second data samples output by the generator corresponding to the at least one data group;

[0311] The category information is used to indicate a second risk category of the data group in the target direction.

[0312] Exemplarily, the apparatus further includes a generative adversarial network training module, the generative adversarial network training module including:

[0313] A first training unit is configured to train a discriminator in the generative adversarial network using a plurality of the third data samples;

[0314] The second training unit is used to train the generator in the generative adversarial network using the noise data sample.

[0315] Exemplarily, one third data sample corresponds to one sample object, and the third data sample further carries a third label representing a second risk category of the sample object in the target direction; the first training unit includes:

[0316] a first input subunit, configured to input the plurality of third data samples into the discriminator, and obtain a first prediction result and a second prediction result output by the discriminator; wherein the first prediction result is used to indicate whether the third data sample is real data, and the second prediction result is used to indicate a second risk category corresponding to the third data sample;

[0317] A first updating subunit is configured to update parameters of the discriminator based on the first prediction result, the second label, the second prediction result, and the third label.

[0318] Exemplarily, the first training unit further includes:

[0319] a second input subunit, configured to input the noise data sample and category information corresponding to the noise data sample into the generator;

[0320] a third input subunit, configured to input the prediction data corresponding to the noise data sample output by the generator into the discriminator;

[0321] A result acquisition subunit, configured to acquire a first prediction result and a second prediction result output by the discriminator; wherein the first prediction result is used to indicate whether the predicted data is real data, and the second prediction result is used to indicate a second risk category corresponding to the predicted data;

[0322] The second updating subunit is configured to update the parameters of the discriminator based on the first prediction result, or to update the parameters of the discriminator based on the first prediction result, the second prediction result and the category information.

[0323] Exemplarily, the plurality of third data samples include real data samples and synthetic data samples, the second label carried by the real data sample indicates that the real data sample is real data, and the second label carried by the synthetic data sample indicates that the real data sample is false data; the first updating subunit is specifically configured to:

[0324] Updating the loss of the discriminator based on the first prediction result, the second prediction result, the second label, the second prediction result, and the third label corresponding to the real data sample;

[0325] Based on the first prediction result and the second label corresponding to the synthetic data sample, the parameters of the discriminator are updated.

[0326] Exemplarily, the second training unit further includes:

[0327] a fourth input subunit, configured to input the noise data sample and category information corresponding to the noise data sample into the generator;

[0328] a fifth input subunit, configured to input the predicted data corresponding to the noise data sample output by the generator into the discriminator;

[0329] A result acquisition subunit, configured to acquire a first prediction result and a second prediction result output by the discriminator; wherein the first prediction result is used to indicate whether the predicted data is real data, and the second prediction result is used to indicate a second risk category corresponding to the predicted data;

[0330] The third updating subunit is used to update the parameters of the generator based on the first prediction result, or to update the parameters of the generator based on the first prediction result, the second prediction result and the category information.

[0331] Exemplarily, multiple sample objects originate from the same individual, and multiple third data samples include at least one real data sample. The target probability corresponding to the sample object to which the real data sample belongs is less than or equal to 0.5, and the target probability represents the statistical significance of the sample object in causing the target risk to the individual in the target direction.

[0332] Exemplarily, the generative adversarial network includes multiple network layers, each of the network layers includes at least one neuron, and the apparatus further includes:

[0333] The inactivation module is used to inactivate some target neurons among the multiple neurons in at least one training session during the process of training the generative adversarial network.

[0334] Exemplarily, the deactivation module is configured to perform at least one of the following processes:

[0335] In N consecutive training sessions, the target neurons subjected to the inactivation treatment are different;

[0336] In N consecutive training sessions, the target neurons subjected to the inactivation treatment are partially repeated;

[0337] In N consecutive trainings, the target neurons subjected to the inactivation process are located in different network layers;

[0338] In N consecutive trainings, the network layer where the target neurons subjected to the inactivation processing are located is repeated.

[0339] Exemplarily, the deactivation process includes: setting the weight parameter corresponding to the target neuron to 0, or fixing the weight parameter of the target neuron at that time.

[0340] Exemplarily, the second acquisition module includes:

[0341] An input unit, configured to input the noise data into the trained generative adversarial network to obtain a plurality of candidate data output by the generator, and a prediction probability corresponding to each candidate data, wherein the prediction probability is a probability representing that the candidate data is true data;

[0342] A screening unit is used to screen out a plurality of second data samples from a plurality of candidate data based on the prediction probability.

[0343] Exemplarily, the sample objects include sites in genes, and the plurality of first data samples include: data samples of first-category sample objects and data samples of second-category objects;

[0344] The mutation frequency corresponding to the first type of sample objects is higher than the mutation frequency corresponding to the second type of objects.

[0345] Based on the same inventive concept, a risk determination device is provided. FIG. 19 shows a schematic diagram of the framework structure of the risk determination device. As shown in FIG. 19 , the risk determination device may specifically include the following modules:

[0346] A data acquisition module, configured to acquire characteristic distribution data of a target object to be classified; wherein the characteristic distribution data includes at least one data factor affecting the target risk of the sample object in a target direction;

[0347] a data input module, configured to input the characteristic distribution data into a risk determination model to obtain a first risk category corresponding to the target risk of the target object;

[0348] Wherein, the risk determination model is obtained based on the training method of the model described above.

[0349] The device embodiment may refer to the above method embodiment and will not be described in detail here.

[0350] The embodiments of the present disclosure also provide a computer-readable storage medium, which stores a computer program that enables a processor to execute the model training method or risk determination method as described in the embodiments of the present disclosure.

[0351] An embodiment of the present disclosure also discloses an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the computer program implements the model training method or risk determination method as described above.

[0352] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0353] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, commodity, or device that includes the element.

[0354] The above is a detailed introduction to the training method, risk determination method, device, equipment and medium of a model provided by the present disclosure. Specific examples are used in this article to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method and core ideas of the present disclosure. At the same time, for those skilled in the art, according to the ideas of the present disclosure, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present disclosure.

[0355] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0356] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

[0357] References herein to "one embodiment," "an embodiment," or "one or more embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Furthermore, please note that instances of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.

[0358] In the description provided herein, numerous specific details are described. However, it is understood that embodiments of the present disclosure may be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0359] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present disclosure may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.

[0360] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.

Claims

1. A method for training a model, wherein: The method comprises: Acquire a first data sample corresponding to each of the plurality of sample objects, wherein the first data sample includes at least one data factor, and the data factor affects a target risk of the sample object in a target direction; Based on the noise data, simulating at least one data factor affecting the target risk of the sample object to obtain a plurality of second data samples; A preset model is trained based on the first data sample and the second data sample to obtain a risk determination model; wherein the risk determination model is used to determine a first risk category corresponding to the target direction of the sample object.

2. The training method according to claim 1, wherein: A first number of the first data samples is greater than a second number of the second data samples, and a ratio of the first number to the second number is greater than or equal to 1.5 and less than or equal to 3.

3. The training method according to claim 1, wherein: The first data sample and the second data sample both carry their own corresponding first labels, where the first labels are used to indicate the first risk category corresponding to the sample object and to supervise the training.

4. The training method according to any one of claims 1 to 3, wherein: The noise data includes at least one data group, and different data groups are used to simulate second data samples corresponding to different sample objects; The step of simulating at least one data factor affecting the target risk of the sample object based on the noise data to obtain a plurality of second data samples includes: Acquire category information corresponding to each of the data groups, the category information being used to indicate a second risk category of the data group in the target direction; Based on at least one of the data groups and the corresponding category information, at least one of the second data samples conforming to the risk category indicated by the category information is generated.

5. The training method according to claim 4, wherein: The first risk category is used to indicate the risk level of the sample object in the target direction, and the second risk category is the same as the first risk category; Alternatively, the first risk category is used to indicate a controllable range after a risk occurs to the sample object in the target direction, and the second risk category is used to indicate a high or low risk of the sample object in the target direction.

6. The training method according to claim 4, wherein: The first data sample includes clinical data and mutation genome data, the first risk category includes at least one of a prognostic risk category, a gene mutation category, and a prognostic survival category; the second risk category includes at least one of a prognostic risk category, a gene mutation category, and a prognostic survival category.

7. The training method according to claim 1, wherein: The noise data includes at least one data group, and different data groups are used to simulate characteristic data corresponding to different objects; the noise data is generated by the following steps: Generate basic data that conforms to normal distribution based on target commands; Perform at least one round of sampling on the basic data to obtain data groups corresponding to the at least one round of sampling; wherein, in each round of sampling, data corresponding to at least one description dimension is sampled from the basic data; Different description dimensions correspond to different influencing factors of the object related to the target direction.

8. The training method according to any one of claims 1 to 3, wherein: The step of simulating at least one data factor affecting the target risk of the sample object based on the noise data to obtain a plurality of second data samples includes: Inputting the noise data into a generator to obtain a plurality of second data samples output by the generator; wherein the generator is obtained by training a generative adversarial network using a plurality of third data samples and noise data samples; Among them, the third data sample carries a second label representing whether the data is real data, the generator in the generative adversarial network is used to generate predicted data based on the noise data sample, and the discriminator in the generative adversarial network is used to determine whether the predicted data is real data.

9. The training method according to claim 8, wherein: The noise data includes at least one data group, and different data groups are used to simulate characteristic data corresponding to different objects; the inputting the noise data into the generator to obtain a plurality of second data samples output by the generator includes: Inputting at least one of the data groups and corresponding category information into the generator, and obtaining second data samples output by the generator and corresponding to at least one of the data groups; The category information is used to indicate a second risk category of the data group in the target direction.

10. The training method according to claim 8, wherein: The step of training the generative adversarial network includes: Using the plurality of the third data samples, training a discriminator in the generative adversarial network; The generator in the generative adversarial network is trained using the noise data samples.

11. The training method according to claim 10, wherein: One of the third data samples corresponds to one sample object, and the third data sample further carries a third label representing a second risk category of the sample object in the target direction; The step of training the discriminator in the generative adversarial network by using the plurality of the third data samples comprises: Inputting a plurality of the third data samples into the discriminator to obtain a first prediction result and a second prediction result output by the discriminator; wherein the first prediction result is used to indicate whether the third data sample is real data, and the second prediction result is used to indicate a second risk category corresponding to the third data sample; Based on the first prediction result, the second label, the second prediction result and the third label, the parameters of the discriminator are updated.

12. The training method according to claim 10, wherein: Before using the noise data sample to train the generator in the generative adversarial network, the method further includes: Inputting the noise data sample and the category information corresponding to the noise data sample into the generator; Inputting the prediction data corresponding to the noise data sample output by the generator into the discriminator; Obtaining a first prediction result and a second prediction result output by the discriminator; wherein the first prediction result is used to indicate whether the predicted data is real data, and the second prediction result is used to indicate a second risk category corresponding to the predicted data; The parameters of the discriminator are updated based on the first prediction result.

13. The training method according to claim 10, wherein: The plurality of third data samples include real data samples and synthetic data samples, the second label carried by the real data sample indicates that the real data sample is real data, and the second label carried by the synthetic data sample indicates that the real data sample is false data; The updating of the loss of the discriminator based on the first prediction result, the second label, the second prediction result and the third label includes: Based on the first prediction result, the second prediction result, the second label, the second prediction result and the third label corresponding to the real data sample, updating the loss of the discriminator; Based on the first prediction result and the second label corresponding to the synthetic data sample, the parameters of the discriminator are updated.

14. The training method according to claim 8, wherein: The step of training the generator in the generative adversarial network using the noise data sample comprises: Inputting the noise data sample and the category information corresponding to the noise data sample into the generator; Inputting the prediction data corresponding to the noise data sample output by the generator into the discriminator; Obtaining a first prediction result and a second prediction result output by the discriminator; wherein the first prediction result is used to indicate whether the predicted data is real data, and the second prediction result is used to indicate a second risk category corresponding to the predicted data; The parameters of the generator are updated based on the first prediction result, or the parameters of the generator are updated based on the first prediction result, the second prediction result and the category information.

15. The training method according to claim 8, wherein: The multiple sample objects originate from the same individual, the multiple third data samples include at least one real data sample, the target probability corresponding to the sample object to which the real data sample belongs is less than or equal to 0.5, and the target probability represents the statistical significance of the sample object in the target direction to the individual's occurrence of the target risk.

16. The training method according to claim 8, wherein: The generative adversarial network includes a plurality of network layers, each of which includes at least one neuron. In the process of training the generative adversarial network, the method further includes: In at least one training session, some target neurons among the plurality of neurons are inactivated.

17. The training method according to claim 16, wherein: In the at least one training, performing inactivation processing on some neurons among the plurality of neurons includes at least one of the following: In N consecutive trainings, the target neurons subjected to the inactivation treatment are different; In N consecutive trainings, the target neurons subjected to the inactivation treatment are partially repeated; In N consecutive trainings, the target neurons subjected to the inactivation process are in different network layers; In N consecutive trainings, the network layer where the target neurons subjected to the inactivation treatment are located is repeated; Wherein, N is an integer greater than or equal to 2.

18. The training method according to claim 16, wherein: The deactivation process includes: setting the weight parameter corresponding to the target neuron to 0, or fixing the weight parameter of the target neuron in the previous training of the current training.

19. The training method according to claim 8, wherein: The step of inputting the noise data into a generator to obtain a plurality of second data samples output by the generator comprises: The noise data is input into the trained generative adversarial network to obtain multiple candidate data output by the generator, and the prediction probability corresponding to each candidate data, wherein the prediction probability is a probability characterizing the candidate data. is the probability of true data; Based on the prediction probability, a plurality of second data samples are selected from a plurality of candidate data.

20. A method for determining risk, wherein: The method comprises: Acquire characteristic distribution data of the target object to be classified; wherein the characteristic distribution data includes at least one data factor affecting the target risk of the sample object in the target direction; Inputting the characteristic distribution data into a risk determination model to obtain a first risk category corresponding to the target risk of the target object; Wherein, the risk determination model is obtained based on the training method of the model described in any one of claims 1-19 above.

21. The risk determination method according to claim 20, or the model training method according to claim 1, wherein: The target direction includes a gene mutation direction, and the target risk includes a prognostic risk; or, the target direction includes a site mutation direction, and the target risk includes a prognostic risk.

22. A model training device, wherein: The device comprises: A first acquisition module, configured to acquire first data samples corresponding to each of the plurality of sample objects, wherein the first data samples include at least one data factor, and the data factor affects a target risk of the sample object in a target direction; A second acquisition module is used to simulate at least one data factor affecting the target risk of the sample object based on the noise data to obtain a plurality of second data samples; A training module is used to train a preset model based on the first data sample and the second data sample to obtain a risk determination model; wherein the risk determination model is used to determine a first risk category corresponding to the target direction of the sample object.

23. A risk determination device, wherein: The device comprises: A data acquisition module, used to acquire characteristic distribution data of a target object to be classified; wherein the characteristic distribution data includes at least one data factor affecting the target risk of the sample object in a target direction; A data input module, used for inputting the characteristic distribution data into a risk determination model to obtain a first risk category corresponding to the target risk of the target object; Wherein, the risk determination model is obtained based on the training method of the model described in any one of claims 1-19 above, or the risk determination method described in any one of claims 20-21.

24. An electronic device, wherein: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor is executed, the method for training a model as described in any one of claims 1 to 19 or the method for determining a risk as described in any one of claims 20 to 21 is implemented.

25. A computer-readable storage medium, wherein: The computer program stored therein enables the processor to execute the model training method described in any one of claims 1-19, or the risk determination method described in any one of claims 20-21.