Multi-constraint Representation Learning Model for Missing Data Filling and Cancer Diagnosis Model

Through the projection matrix and constraint term optimization of the multi-constraint characterization learning model, the problem of inaccurate interpolation of missing data in ovarian cancer diagnosis is solved, stable and accurate data filling is achieved, and the classification ability of ovarian cancer diagnosis is improved.

CN119089125BActive Publication Date: 2025-07-04SOUTHERN MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411565752.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-05
Publication Date
2025-07-04
Estimated Expiration
2044-11-05

AI Technical Summary

Technical Problem

The existing deleted data interpolation methods are not stable and accurate enough in the diagnosis of ovarian cancer, especially when the proportion of deletions is high, which affects the accuracy of data set information and classification model.

Method used

The multi-constraint characterization learning model is adopted to project missing data to the latent space through the projection matrix, and the feature importance consistency constraint terms, missing position estimation constraint terms and fuzzy relationship constraint terms are used to optimize the redundancy between shared fusion features and the correlation with labels, and improve the accuracy and stability of imputation.

Benefits of technology

It improves the accuracy and stability of missing data filling, and improves the classification ability of ovarian cancer diagnostic models, especially in the case of high missing ratio, which has good identification and classification effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119089125B_ABST
    Figure CN119089125B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-constraint representation learning model for missing data filling and a cancer diagnosis model. The multi-constraint representation learning model includes: a data preprocessing module for preprocessing source data to obtain preprocessed data; a projection module for projecting the preprocessed data according to a projection matrix to obtain filled data and output it, where the projection matrix is obtained after training the multi-constraint representation learning model. Among them, the constraint functions for training the multi-constraint representation learning model include: a projected data constraint term, a feature importance consistency constraint term, a missing position estimation constraint term, and a fuzzy relationship constraint term. It can stably and accurately fill missing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing technology, and specifically relates to a multi-constraint representation learning model and a cancer diagnosis model for missing data filling. Background Art

[0002] With the help of artificial intelligence, we can explore the value of laboratory tests in the diagnosis and prognosis prediction of ovarian cancer, screen out early ovarian cancer patients, improve postoperative living treatment, and conduct postoperative evaluation and monitoring of patients with advanced cancer.

[0003] However, due to various reasons, such as sample quality issues, patient wishes, data transmission and preservation, etc., the laboratory index data detected by routine health examinations are missing, and the missing rate is not low. The missing rate of some items can reach 90%. It is a big problem that missing data cannot be used for classifier training using machine learning. If you simply delete the missing patient data and only keep the complete patient data, the sample size of the patient will be greatly reduced, and at the same time, it will also cause the loss of data set information, which will have a great impact on the subsequent identification of ovarian cancer. Therefore, the current mainstream approach is to interpolate the missing data set first, and then train the classifier to obtain a classification model.

[0004] The current missing value imputation algorithms mainly include: imputation methods based on statistics, imputation methods based on machine learning, and imputation methods based on deep learning. The imputation methods based on statistics are simply filling in the maximum and minimum values. Further, it is to perform imputation of mean, median, and mode based on simple statistics. In addition to the above simple imputation methods, there are currently imputation methods that use singular value decomposition to maintain the distribution unchanged (SOFT), and the most classic and commonly used imputation method that uses global information is multiple imputation (MICE). Its main idea is to generate multiple completed versions of multiple datasets to better estimate and handle the uncertainty of missing data. There is also BPCA (Bayesian Principal Component Analysis), a multiple imputation method based on principal component analysis (PCA) to handle incomplete continuous data, and in order to reflect the uncertainty of parameters from one imputation to the next, the Bayesian method is used to process the PCA model. For the imputation methods based on machine learning, currently commonly used ones include KNN (K-Nearest Neighbors), the missing data imputation method based on random forest (MISS FOREST), XGBOOST (Extreme Gradient Boosting), etc. They use machine learning models to directly predict the missing values. Due to the strong adaptability and prediction ability of deep learning, in recent years, there have been more and more imputation methods based on deep learning. For example, the GAIN (Generative Adversarial Imputation Nets) algorithm proposes using generative adversarial networks for missing data imputation to make the generated data as close as possible to the real distribution. In addition, the imputation algorithm based on autoencoders learns the compressed representation of data using a partially complete dataset, and then decodes the compressed representation of the incomplete data to obtain the complete data. There is also a further improved method of autoencoders, REMASK (REcurrent Masked Autoencoder), which randomly "re-masks" another set of values in addition to the missing values (i.e., the natural mask), and optimizes the autoencoder by reconstructing this set of re-masked values, as well as the imputation algorithm based on the optimal transport (OT) method that uses the optimal transport distance to quantify the differences between different distributions and converts it into a loss function to estimate the missing data values.

[0005] In the imputation methods based on statistics, simply using conventional statistical variables such as the mean and median for imputation will greatly reduce the variability of the data, because all missing values are replaced with the same mean, which makes the imputed data unable to reflect the volatility of the real data. In addition, using conventional statistical variables for imputation may introduce biases. Especially when the data itself has an asymmetric distribution or a multimodal distribution, this method may lead to distorted results, affecting the prediction performance and accuracy of the classification model. The disadvantage of using MICE imputation lies in its high computational complexity. In addition, MICE has a strong dependence on model assumptions. If the model is not properly selected, it may lead to inaccurate imputation results or even introduce systematic errors. Since MICE generates imputation values through multiple models, the stability and consistency of its results may be inferior to other simpler imputation methods. Moreover, the imputation effect of the imputation methods based on statistics is relatively poor for features with a high missing ratio.

[0006] For the imputation methods based on machine learning, on the one hand, the performance depends relatively much on the adjustment of machine learning model parameters. On the other hand, the specific imputation effect is often highly related to the data structure distribution. If the missing ratio is relatively high, the imputation effect may be poor. Finally, these machine learning models often require a certain amount of non-missing samples to build the model, and if the quality of the non-missing samples is not good, the imputation methods based on machine learning will amplify the noise or outliers in the data.

[0007] For the imputation methods based on deep learning, the main disadvantages are that they usually require a large amount of data and computing resources to train complex neural networks, which is particularly time-consuming when dealing with large-scale datasets or high-dimensional data. In addition, these methods are relatively sensitive to the selection of hyperparameters and network architectures. Improper settings may lead to unstable or inaccurate imputation results. Deep learning models are also prone to overfitting, especially in the case of sparse data, and may generate unreliable imputation results. Finally, the current interpretability of deep learning is relatively poor, and most of the prediction processes are "black boxes".

[0008] Therefore, the imputation results of the existing missing data imputation methods are not stable and accurate enough. Summary of the Invention

[0009] The purpose of the present invention is to provide a multi-constraint representation learning model and a cancer diagnosis model for missing data filling, which can stably and accurately fill the missing data.

[0010] The first aspect of the present invention discloses a multi-constraint representation learning model for missing data filling, including:

[0011] A data preprocessing module for preprocessing the source data to obtain preprocessed data;

[0012] A projection module, configured to project the preprocessed data according to a projection matrix, obtain and output the filled data, where the projection matrix is obtained after training the multi-constraint representation learning model;

[0013] Among them, the constraint functions for training the multi-constraint representation learning model include: a projection data constraint term, a feature importance consistency constraint term, a missing position estimation constraint term, and a fuzzy relationship constraint term;

[0014] The feature importance consistency constraint term is used to keep the weight ranking of features consistent with the feature importance ranking determined by multiple feature selection methods during the projection process; the missing position estimation constraint term is used to identify the inconsistency between feature values according to the information of the missing positions during back-projection; the fuzzy relationship constraint term is used to optimize the redundancy between the shared fusion features obtained by projection and the correlation between the shared fusion features and the labels through fuzzy relationships.

[0015] In some embodiments, the expression of the projection data constraint term is: , where is the preprocessed training data, is the projection matrix, is the back-projection matrix, is the latent space to which it is projected, represents the F norm.

[0016] In some embodiments, the feature importance consistency constraint term includes a projection weight ranking unit and a feature importance ranking unit. The projection weight ranking unit is configured to calculate the weight ranking of different features in the projection matrix during the projection process, and the feature importance ranking unit is configured to calculate the correlation between features and labels respectively through multiple feature selection algorithms, and determine the feature importance ranking of each feature according to the correlation between features and labels.

[0017] In some embodiments, calculating the correlation between features and labels respectively through multiple feature selection algorithms, and determining the feature importance ranking of each feature according to the correlation between features and labels includes:

[0018] Calculating the correlation between features and labels respectively according to multiple feature selection algorithms, and ranking according to the correlation between features and labels;

[0019] Using the ranking aggregation method to convert the ranking results into a preference matrix;

[0020] Based on the preference matrix, minimizing the cost of an optimal ranking in the preference matrix through the Manhattan distance to obtain the feature importance ranking.

[0021] In some embodiments, the fuzzy relationship constraint term includes a correlation unit and a redundancy unit. The correlation unit is used to calculate the fuzzy mutual information between the shared fusion feature and the label, and the redundancy unit is used to calculate the fuzzy mutual information between the shared fusion features.

[0022] In some embodiments, the expression of the correlation unit is: , ; the expression of the redundancy unit is: , ; where is the shared fusion feature; is a one-dimensional vector obtained by reducing the dimensionality of all features of the shared fusion feature except by PCA; is a diagonal matrix, and the elements on the diagonal are the sums of the corresponding elements in each row of the fuzzy relationship matrix ; is a diagonal matrix, and the elements on the diagonal are the sums of the corresponding elements in each row of; n is the total number of samples; k is the dimension of the shared fusion feature; is the label.

[0023] In some embodiments, the missing position estimation constraint term includes a mask of missing data and a conversion unit. The conversion unit is used to calculate the loss of transforming the data projected into the latent space into the mask of missing data.

[0024] In some embodiments, the expression of the constraint function of the multi-constraint representation learning model is:

[0025] ,

[0026] where is the preprocessed data; is the projection matrix; is the back-projection matrix; is the latent space to which it is projected; is the mask of missing data; is the feature importance consistency constraint term; is a linear transformation matrix for transforming the data in the latent space into the mask of missing data; and are the fuzzy relationship constraint terms; , , , , , and are trade-off parameters for balancing different terms; is a regularization term.

[0027] The second aspect of the present invention discloses a cancer diagnosis model, including the multi-constraint representation learning model and the cancer classification model of any one of the above.

[0028] In some embodiments, the steps of determining the cancer classification model include:

[0029] Training the multi-constraint representation learning model with a missing data training set to obtain a projection matrix;

[0030] Projecting the missing data training set into a shared fusion feature data set according to the projection matrix;

[0031] Training multiple classifiers based on the shared fusion feature data set, and selecting the optimal classifier as the cancer classification model.

[0032] The beneficial effects of the present invention are as follows: the imputation of missing data is realized by using the projection matrix in the multi-constraint representation learning model. When training the multi-constraint representation learning model so that the projection matrix learns the law of missing data, the consistency of feature importance and the missing position estimation constraint term are added. On the one hand, it avoids the loss of information of important features during the projection process and improves the rationality of the representation learning projection. On the other hand, through the estimation of the missing position during the projection representation learning process, the accuracy of the representation learning projection is improved. And the redundancy between the shared fusion features after projection and the correlation with the label are optimized through the fuzzy relation constraint term, so that the classification ability of the data processed by the representation learning model is greatly improved. It can fill the missing data stably and accurately. Description of the Drawings

[0033] The drawings here show specific examples of the technical solutions of the present invention and form a part of the description together with the specific implementation manners, and are used to explain the technical solutions, principles and effects of the present invention.

[0034] Unless otherwise specified or defined, in different drawings, the same reference numerals represent the same or similar technical features, and for the same or similar technical features, different reference numerals may also be used to represent them.

[0035] Figure 1 is a module schematic diagram of a multi-constraint representation learning model disclosed in an embodiment of the present invention. Detailed Embodiments

[0036] Unless otherwise specified or defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs. In the context of combining the technical solutions of the present invention with real-world scenarios, all technical and scientific terms used herein may also have meanings corresponding to the purpose of implementing the technical solutions of the present invention. The "first, second, ..." used herein is only for differentiating names and does not represent a specific quantity or order. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0037] It should be noted that when an element is considered to be "fixed to" another element, it can be directly fixed to the other element or there can be an intermediate element; when an element is considered to be "connected to" another element, it can be directly connected to the other element or there can be an intermediate element at the same time; when an element is considered to be "mounted on" another element, it can be directly mounted on the other element or there can be an intermediate element at the same time. When an element is considered to be "provided in" another element, it can be directly provided in the other element or there can be an intermediate element at the same time.

[0038] Unless otherwise specified or defined, the "said" and "the" used herein refer to the technical features or technical contents mentioned or described before the corresponding position. The technical features or technical contents can be the same as or similar to the technical features or technical contents they mention. In addition, the terms "comprising" and "having" and any variations thereof used herein are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0039] The imputation results of common missing data imputation methods are always not stable and accurate enough. Especially when dealing with missing conventional laboratory test indicators, due to the characteristics of conventional laboratory test indicators such as many features, a high proportion of missing features in some parts, and strong correlations between features, the accuracy and stability of the imputation results need to be further improved.

[0040] In view of the above problems, the present invention proposes a multi-constraint representation learning model. The projection matrix in the multi-constraint representation learning model is used to project the missing data into the latent space to obtain complete data, thereby realizing the filling of the missing data. When training the multi-constraint representation learning model to enable the projection matrix to learn the rules of the missing data, a feature importance consistency constraint term and a missing position estimation constraint term are added. On the one hand, it avoids the loss of information of important features during the projection process and improves the rationality of the representation learning projection. On the other hand, through the estimation of the missing positions during the projection representation learning process, the accuracy of the representation learning projection is improved. Moreover, the redundancy between the shared fusion features obtained by optimizing the projection through fuzzy relations and the correlation with the labels are optimized, so that the discriminative classification ability of the data processed by the representation learning model is greatly improved.

[0041] In this embodiment, a multi-constraint representation learning model is provided for ovarian cancer laboratory index data to fill the missing data in the ovarian cancer laboratory index data. It should be noted that the multi-constraint representation learning model of the present invention is not limited to filling laboratory index data and can also be used to fill other types of data.

[0042] Specifically, the multi-constraint representation learning model mainly includes two parts: a data preprocessing module and a projection module. Among them, the data preprocessing module is used to perform various preprocessing such as normalization on various source data such as ovarian cancer laboratory index data to obtain preprocessed data. By normalizing each feature of the data set, the multi-constraint representation learning model can be made not to be affected by the data size. Then, the preprocessed data is multiplied by the projection matrix to project the preprocessed data and obtain and output the filled data. Among them, the projection matrix is a data matrix. When training the multi-constraint representation learning model, the projection matrix will be iteratively optimized. After the training is completed, the final projection matrix can be obtained.

[0043] In this embodiment, when training the multi-constraint representation learning model, a constraint function of the multi-constraint representation learning model is constructed according to the projection data constraint term, the feature importance consistency constraint term, the missing position estimation constraint term, and the fuzzy relation constraint term. Among them, the feature importance consistency constraint term is used to keep the weight ranking of the features consistent with the feature importance ranking determined by various feature selection methods during the projection process; the missing position estimation constraint term is used to identify the inconsistency between the feature values according to the information of the missing positions during back-projection; the fuzzy relation constraint term is used to optimize the redundancy between the shared fusion features obtained by the projection through fuzzy relations and the correlation between the shared fusion features and the labels.

[0044] Specifically, referring to Figure 1 , the construction process of the constraint function is as follows:

[0045] First, a matrix is defined Represents the patient data matrix, which is used as training data for training. Among them, n represents the number of patients, m represents the number of features, and normalizes each feature of the data. In order to learn the latent representation of the data to obtain the complete data, it is assumed that there is a latent space , and through the projection matrix the data can be projected into this latent space and the shared fusion features can be learned. Among them, k represents the dimension of the shared fusion features obtained by projecting the data into the latent space. At the same time, through the reconstruction matrix the latent space is reconstructed back to the patient data matrix, and the back-projection process can significantly improve the reliability of data projection in representation learning. According to the bidirectional mapping process, the projection data constraint term can be determined and is defined as follows:

[0046] (1)

[0047] Among them, is the preprocessed training data, is the projection matrix, is the back-projection matrix, is the latent space to which it is projected, represents the F norm. Then, the projection matrix , the back-projection matrix and the latent space are regularized, and the rewritten formula is as follows.

[0048] (2)

[0049] Among them, , and are trade-off parameters used to balance different terms. In this embodiment, it is defined as . , , and have similar meanings. Regularizing these parameters can avoid overfitting problems. Formula (2) is the preliminary framework of the bidirectional mapping of representation learning. This framework encourages projecting incomplete data into the latent space to obtain complete data and learning the projection matrix for new incomplete data.

[0050] However, if data is directly subjected to projection representation learning without constraints, the resulting shared fusion features are unreasonable. This is because in this process, the projection weights of different features are determined according to the specific missing conditions during the learning of data features and the distribution law of the data. Such weight constraints are one-sided as they do not take into account that the importance of different features is often different, and the impact of feature importance on the projection weights should be significant. Otherwise, the shared fusion features may lose the information of important features. Therefore, in this embodiment, by introducing a feature importance consistency constraint term into the constraint function, the projection matrix ranks the weights of different features in the projection process and then, through a constraint term, makes the weight ranking of features in the projection process consistent with the feature importance ranking determined by multiple feature selection methods.

[0051] The feature importance consistency constraint term includes a projection weight ranking unit and a feature importance ranking unit. The projection weight ranking unit is used to calculate the weight ranking of different features in the projection matrix during the projection process. The specific steps are as follows: First, calculate the weights of different features in the projection process. The specific calculation formula is as follows:

[0052] ,

[0053] where is a column vector all of whose elements are 1, is a column vector, is the sum of the weights of the th feature in the projection process.

[0054] Then, calculate the weight ranking of the features. The specific formula is as follows:

[0055] ,

[0056] where represents the weight ranking of the th feature, and the indicator function represents the binary relation . When , . However, because the indicator function is not a continuous function, it cannot be differentiated and optimized in the subsequent optimization process. Therefore, this embodiment introduces the SIGMOD function , which can map to the range of 0 - 1.

[0057] ,

[0058] After the above two-step calculation, the weight ranking of features can finally be obtained .

[0059] The feature importance ranking unit is used to calculate the correlation between features and labels respectively through multiple feature selection algorithms, and determine the feature importance ranking of each feature according to the correlation between features and labels. The specific process is as follows: five different feature selection algorithms, namely random forest, mutual information, logistic regression, gradient boosting tree, and support vector machine, are used. These feature selection algorithms calculate the correlation between features and labels through a certain variable, and rank the importance of features according to the correlation between features and labels; the Kemeny-Young ranking aggregation method is used to convert the ranking results into a preference matrix; then, the cost of an optimal ranking in the preference matrix is minimized through the Manhattan distance to obtain the final feature importance ranking .

[0060] Weight ranking vector and feature importance ranking The elements in the vector correspond one by one, and are all the importance rankings of features in a certain order. Finally, a feature importance consistency constraint term is designed to make the two rankings as consistent as possible. The specific formula of the feature importance consistency constraint term is as follows:

[0061] (3),

[0062] where .

[0063] After organizing formulas (2) and (3), the constraint function of the updated multi-constraint representation learning model is obtained:

[0064] (4).

[0065] Since representation learning is to perform projection learning on missing data, and the proportion of missing features in ovarian cancer laboratory indicators is not low, it is quite important how to make the multi-constraint representation learning model learn the laws of missing data. Through analysis and research, the information of missing data lies not only in the data distribution but also in the specific missing positions of the missing data. Therefore, it is particularly crucial to add the information of the missing positions to the constraint function of the multi-constraint representation learning model. In addition, a key point of representation learning is to learn the correlation between feature values. Although 0 is filled in the blank positions for calculation when projecting missing data, simple projection and back-projection cannot make good use of the missing position information, which may lead to learning incorrect information. Based on the above two points, the missing position estimation constraint term in this embodiment includes the mask and conversion unit of the missing data. By using the mask of the missing data during back-projection to enable the latent space data to restore the specific missing situation of the source data, the mask of the missing data is a matrix consisting of only 0 and 1. The specific distribution of 0 and 1 is determined by the missing positions in the missing dataset. If it is missing, fill it with 0, and fill other positions with 1. When back-projecting according to the information of the missing positions, the inconsistency between eigenvalues can be recognized and the accuracy of the representation learning of the missing data can be improved. The expression of the conversion unit is: , which is used to calculate the loss of converting the data projected into the latent space into a mask of the missing data. Among them, is a linear transformation matrix, whose function is to transform the data in the latent space into a mask of the missing data, so that the shared fusion features in the latent space obtain the information of the mask.

[0066] After applying the missing position estimation constraint term to formula (1), the update is obtained as:

[0067] (5),

[0068] After sorting out formulas (4) and (5), the updated constraint function is as follows:

[0069] (6),

[0070] Since the purpose of the representation learning model is to project the missing data into the latent space to obtain complete data that can be classified. In order for the shared fusion features obtained by projecting the missing data to have better discriminative classification ability, in this embodiment, starting from the relationship between information, the correlation between the shared fusion features and the labels and the redundancy between the shared fusion features are connected through fuzzy information relationships, optimizing and enhancing the correlation between the shared fusion features and the labels and reducing the redundancy between the shared fusion features, so as to achieve the effect of improving the discriminative classification ability of the shared fusion features. Therefore, a fuzzy relationship constraint term is introduced into the constraint function. The fuzzy relationship constraint term includes a correlation unit and a redundancy unit. The correlation unit is used to calculate the fuzzy mutual information between the projected shared fusion features and the labels, and the redundancy unit is used to calculate the fuzzy mutual information between the projected shared fusion features.

[0071] The specific construction process of the fuzzy relationship constraint term is as follows:

[0072] First is the definition of fuzzy information. In this embodiment, each feature of the shared fusion features will have a fuzzy relationship matrix representing a fuzzy equivalence relationship.

[0073] ,

[0074] Among them, is a function for measuring membership relationships, used to measure the similarity between any two samples in The th feature among the shared features and the th sample. Similarly, the label of the patient data also has a similar fuzzy relationship matrix .

[0075] After defining the fuzzy relationship matrix, the fuzzy information entropy of relevant information can be calculated:

[0076] ,

[0077] where is the estimated value of the fuzzy equivalence related to , .

[0078] Furthermore, the fuzzy joint information entropy between two fuzzy equivalence informations is derived , and the calculation formula is:

[0079] ,

[0080] Finally, the fuzzy mutual information between two information vectors is defined as:

[0081] ,

[0082] Through the above definition of the fuzzy mutual information between information vectors, the correlation between the shared fusion feature and the label and the redundancy between shared fusion features can be quantified. The specific formula is:

[0083] (7),

[0084] (8),

[0085] where is the shared fusion feature; is a one-dimensional vector obtained by reducing the dimensionality of all features of the shared fusion feature except through PCA; is a diagonal matrix, and the elements on the diagonal are the sum of the elements in each corresponding row of the fuzzy relationship matrix . Its matrix derivation form is:

[0086] ,

[0087] is the same as and is also a diagonal matrix, and the elements on the diagonal are the sum of each corresponding row in . is similar to it, and its matrix derivation form is:

[0088] ,

[0089] Among them, is a square matrix, whose first column elements are 1 and the rest are 0, while is an identity matrix.

[0090] Finally, simplify and derive formulas (7) and (8) to the entire shared fusion feature to obtain the correlation unit and redundancy unit.

[0091] The expression of the correlation unit is:

[0092] (9),

[0093] The expression of the redundancy unit is:

[0094] (10),

[0095] Combining formulas (6), (9), and (10) to obtain the final constraint function of the multi-constraint representation learning model:

[0096] ,

[0097] Among them, is the data after preprocessing; is the projection matrix; is the back-projection matrix; is the latent space to which it is projected; is the mask of the missing data; is the feature importance consistency constraint term; is a linear transformation matrix used to transform the data in the latent space into the mask of the missing data; and are the fuzzy relationship constraint terms; , , , , , and are the trade-off parameters used to balance different terms; is the regularization term.

[0098] After constructing the constraint function of the multi-constraint representation learning model, the multi-constraint representation learning model is trained and optimized. The training process of this embodiment is as follows: 87 routine laboratory test indicators and age of a total of 714 patients with ovarian cancer (primary ovarian cancer, fallopian tube cancer, and peritoneal cancer) and non-ovarian cancer control groups (benign ovarian lesions and some benign lesions of the fallopian tubes and uterus) are collected from the hospital as features, and the total number of features is 88. The collected data matrix is set as And the pathological examination results are used as labels (ovarian cancer = 1; non-ovarian cancer = 0). For the missing data matrix Perform normalization processing, fill 0 at the position of the missing value, and calculate the mask of the missing data . For the parameters that need to be optimized in the constraint function, perform random initialization; under the guidance of the constraint function, perform gradient descent loop t times or until the loss value of the constraint function is reduced to a certain level: in each loop process, the derivatives of the four parameters with respect to the objective function need to be calculated respectively , , , , and the parameters are updated iteratively using the following several formulas:

[0099]

[0100]

[0101]

[0102]

[0103] Among them, is the learning rate of the gradient descent of several parameters, and the formula of is to make directly obtained.

[0104] The ultimate goal of iteration is to obtain the projection matrix , and this projection matrix can be used for subsequent projection representation learning of new missing data to obtain a new shared fusion feature for discriminant classification.

[0105] The imputed ovarian cancer laboratory index data in this embodiment and the data processed by 10 commonly used missing data imputation methods based on statistical theory, machine learning models, and deep learning models at the present stage are tested. Classification effects are quantified using indicators such as area under the curve (AUC), accuracy (ACC), sensitivity (SEN), and specificity (SPE), and the obtained prediction results are compared. The AUC, ACC, SEN, and SPE of the ovarian cancer laboratory index data tested after imputation in this embodiment are 0.998, 0.983, 0.968, and 0.987 respectively. It shows that the laboratory index data after imputation by the multi-constraint representation learning model has a good prediction effect on ovarian cancer prediction, and AUC, ACC, and SEN are all better than other imputation methods, indicating a better discrimination and classification effect on the processed data. Especially, SEN is better than other imputation methods, which has great clinical significance in practical applications.

[0106] In summary, for the multi-constraint representation learning model of this embodiment, in view of the characteristics of many missing laboratory index features, strong feature correlation, and a high proportion of missing parts of some features, constraint terms for consistent feature importance and missing position estimation are added. On the one hand, it avoids information loss of important features during the projection process and improves the rationality of the representation learning projection. On the other hand, through the estimation of the missing positions during the projection representation learning process, the accuracy of the representation learning projection is improved. And by optimizing the redundancy between the projected shared fusion features and the correlation with the label through fuzzy relations, the discrimination and classification ability of the data processed by the representation learning model is greatly improved. It can handle the situation of a large amount of missing data when using laboratory indexes for early diagnosis of ovarian cancer.

[0107] Based on the above multi-constraint representation learning model, the present invention also provides a cancer diagnosis model, which includes any one of the above multi-constraint representation learning models and a cancer classification model.

[0108] The cancer diagnosis model of this embodiment also optimizes and selects the cancer classification model based on the training results of the multi-constraint representation learning model. First, the multi-constraint representation learning model is trained using the missing data training set to obtain a projection matrix; the missing data training set is projected into a shared fusion feature data set according to the projection matrix; then, various classifiers are trained based on the shared fusion feature data set, and the optimal classifier is selected as the cancer classification model.

[0109] The specific process is as follows: First, the missing data is normalized, then the missing positions are filled with 0, and then the missing data is divided into 5 folds, namely the training set and the test set . Each fold of the missing data training set is trained under the representation learning model to obtain the corresponding projection matrix for each fold , multiply the missing data training set of each fold by the projection matrix to obtain the shared fusion feature data set corresponding to each fold. . Use all the shared fusion feature data sets. Perform training on several conventional classifiers and select the optimal classifier as the cancer classification model.

[0110] During the test of the cancer diagnosis model, the divided test set is multiplied by the projection matrix to calculate the shared fusion feature data set for testing. , and input the shared fusion feature data set for testing into the classifier obtained in the training stage for testing to obtain the predicted output result.

[0111] The purpose of the above embodiments is to exemplarily reproduce and deduce the technical solution of the present invention, and to completely describe the technical solution, purpose and effect of the present invention. Its purpose is to enable the public to understand the disclosed content of the present invention more thoroughly and comprehensively, and it does not limit the protection scope of the present invention.

[0112] The above embodiments are not an exhaustive list based on the present invention. In addition, there may be multiple other embodiments not listed. Any substitution and improvement made on the basis of not violating the concept of the present invention fall within the protection scope of the present invention.

Claims

1. A method for filling missing data based on a multi-constraint representation learning model, characterized in that Including: Preprocess the data of ovarian cancer laboratory test indicators to obtain preprocessed data; Project the preprocessed data according to the projection matrix to obtain filled data and output it, where the projection matrix is obtained after training the multi-constraint representation learning model; Among them, the constraint functions for training the multi-constraint representation learning model include: projection data constraint term, feature importance consistency constraint term, missing position estimation constraint term, and fuzzy relationship constraint term; The feature importance consistency constraint term is used to keep the weight ranking of features consistent with the feature importance ranking determined by multiple feature selection methods during the projection process; the missing position estimation constraint term is used to identify the inconsistency between feature values according to the information of the missing position during back-projection; the fuzzy relationship constraint term is used to optimize the redundancy between the shared fusion features obtained by projection and the correlation between the shared fusion features and the label through fuzzy relationships; The expression of the projection data constraint term is as follows: , where is the preprocessed training data, is the projection matrix, is the back-projection matrix, is the latent space to which it is projected, represents the F-norm.

2. The missing data filling method based on the multi-constraint representation learning model according to claim 1, wherein The feature importance consistency constraint term includes a projection weight ranking unit and a feature importance ranking unit. The projection weight ranking unit is used to calculate the weight ranking of different features in the projection matrix during the projection process, and the feature importance ranking unit is used to calculate the correlation between features and labels respectively through multiple feature selection algorithms, and determine the feature importance ranking of each feature according to the correlation between features and labels.

3. The missing data filling method based on the multi-constraint representation learning model according to claim 2, wherein The calculating the correlation between features and labels respectively through multiple feature selection algorithms, and determining the feature importance ranking of each feature according to the correlation between features and labels includes: Calculating the correlation between features and labels respectively according to multiple feature selection algorithms, and ranking according to the correlation between features and labels; Using the method of ranking aggregation to convert the ranking results into a preference matrix; Based on the preference matrix, minimizing the cost of an optimal ranking in the preference matrix through the Manhattan distance to obtain the feature importance ranking.

4. The missing data filling method based on the multi-constraint representation learning model according to claim 1, wherein The fuzzy relationship constraint term includes a correlation unit and a redundancy unit. The correlation unit is used to calculate the fuzzy mutual information between the shared fusion feature and the label, and the redundancy unit is used to calculate the fuzzy mutual information between the shared fusion features.

5. The method for filling missing data based on the multi-constraint representation learning model according to claim 4, wherein The expression of the correlation unit is as follows: , ; The expression of the redundancy unit is as follows: , ; Among them, is the shared fusion feature; is the one-dimensional vector formed by dimensionality reduction of all features of the shared fusion feature except through PCA; is a diagonal matrix, and the elements on the diagonal are the sum of the elements corresponding to each row in the fuzzy relation matrix ; is a diagonal matrix, and the elements on the diagonal are the sum of the elements corresponding to each row in ; is the total number of samples; is the dimension of the shared fusion feature; is the label.

6. The method for filling missing data based on the multi-constraint representation learning model according to claim 1, wherein The missing position estimation constraint term includes a mask for missing data and a conversion unit. The conversion unit is used to calculate the loss of transforming the data projected into the latent space into the mask of the missing data.

7. The missing data filling method based on the multi-constraint representation learning model according to claim 1, wherein The expression of the constraint function of the multi-constraint representation learning model is: , Among them, is the preprocessed data; is the projection matrix; is the back-projection matrix; is the latent space onto which it is projected; is the mask of the missing data; is the feature importance consistency constraint term; is a linear transformation matrix for transforming the data in the latent space into the mask of the missing data; and are the fuzzy relation constraint terms; 、 、 、 、 、 and are the trade-off parameters for balancing different terms; is the regularization term.

8. A method for selecting a cancer classification model, characterized in that The steps for determining the cancer classification model include: Training the multi-constraint representation learning model with a missing data training set to obtain a projection matrix, where the missing data training set consists of ovarian cancer laboratory test indicator data; Projecting the missing data training set into a shared fusion feature data set according to the projection matrix; Training multiple classifiers based on the shared fusion feature data set, and selecting the optimal classifier as the cancer classification model; Among them, the constraint functions for training the multi-constraint representation learning model include: projection data constraint term, feature importance consistency constraint term, missing position estimation constraint term, and fuzzy relationship constraint term; The feature importance consistency constraint term is used to keep the weight ranking of features consistent with the feature importance ranking determined by multiple feature selection methods during the projection process; the missing position estimation constraint term is used to identify the inconsistency between feature values according to the information of the missing position during back-projection; the fuzzy relationship constraint term is used to optimize the redundancy between the shared fusion features obtained by projection and the correlation between the shared fusion features and the labels through the fuzzy relationship; The expression of the projection data constraint term is as follows: , where is the preprocessed training data, is the projection matrix, is the back-projection matrix, is the latent space to which it is projected, represents the F-norm.

Citation Information

Patent Citations

  • Traffic data filling method based on non-negative low-rank dynamic mode decomposition

    CN110188427A

  • Machine learning-based sepsis early prediction method and system

    CN118507071A