A Method for Nonlinear Association Analysis Based on Joint Structure Constraint and Incomplete Multimodal Data

Through the nonlinear correlation analysis method of incomplete multimodal data based on joint structural constraints, the problems of data loss and linear model limitation in multimodal image data are solved, and the performance improvement of detection of biomarkers and the consideration of complex association of SNP and phenotypes are realized.

CN114187962BActive Publication Date: 2025-06-24SOUTHERN MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111308654.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-05
Publication Date
2025-06-24
Estimated Expiration
2041-11-05

AI Technical Summary

Technical Problem

The prior art has the problem of data missing in multimodal image data. Only complete data is used for modeling, resulting in information loss and degradation of detection performance, while lacking the consideration of inter-modal and in-modal data correlation, and the inability to effectively detect complex associations between SNP and phenotypes using only linear models.

Method used

A nonlinear correlation analysis method of incomplete multimodal data based on joint structural constraints is adopted. By collecting multimodal image data and genetic data, preprocessing and quality control are performed, and the objective function is used to solve the weights of SNP and phenotype on different modalities, and the nonlinear association between SNP and phenotype is constructed through nonlinear transformation.

Benefits of technology

It can detect modal sharing and modal-specific biomarkers, improve the performance of detection of biomarkers, consider the complex association between SNP and phenotype, and solve the problems of data loss and linear model limitation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114187962B_ABST
    Figure CN114187962B_ABST
Patent Text Reader

Abstract

A method for non-linear association analysis based on joint structure constraints and incomplete multi-modal data, which obtains the weights corresponding to multiple modal phenotype data and SNPs through 4 steps. The present invention constructs the non-linear association between SNPs and phenotypes through non-linear transformation, thereby considering the complex association between SNPs and phenotypes, and obtains the modal sharing and modal-specific biomarkers corresponding to different modalities through the contributions of multiple SNPs to phenotypes. The root mean square error of the present invention is significantly better than the value of the root mean square error obtained by the prior art, thereby being able to improve the performance of detecting biomarkers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of application of genetic data structure information and its incomplete multimodal data, and particularly relates to a non-linear association analysis method based on joint structure constraint and incomplete multimodal data. Background Art

[0002] Du et al. (L. Du et al., "Multi-Task Sparse Canonical Correlation Analysis with Application to Multi-Modal Brain Imaging Genetics," IEEE / ACM Transactions on Computational Biology and Bioinformatics, vol. 18, no. 1, pp. 227-239, 2021.) proposed a multi-task sparse canonical correlation analysis (MTSCCA) method, which uses multi-modal image data generated by different imaging technologies that may carry complementary information to identify disease-related SNPs and multi-modal phenotypes. At the same time, this method takes into account the structural associations between genetic data and the sparsity at the individual level of genetic and phenotypic data. Considering this information may improve the detection performance of biomarkers.

[0003] However, there are some problems in the MTSCCA method. First, due to reasons such as imaging quality and high cost, most multi-modal imaging phenotype data have data missing problems. This method removes the missing parts of the samples and only uses the complete multi-modal imaging data for modeling, which may lose some information and thus reduce the detection performance. Second, the MTSCCA method only focuses on the feature information of a single modality and does not consider the correlations between data within and between modalities. Third, the MTSCCA method uses a linear model to identify the linear association between SNPs and phenotypes. However, the association between SNPs and phenotypes is very complex, and it is difficult to detect such complex relationships with a simple linear model.

[0004] Therefore, in view of the deficiencies of the prior art, it is very necessary to provide a non-linear association analysis method based on joint structure constraint and incomplete multimodal data to solve the deficiencies of the prior art. Summary of the Invention

[0005] The purpose of the present invention is to avoid the deficiencies of the prior art and provide a non-linear association analysis method based on joint structure constraint and incomplete multimodal data. This non-linear association analysis method based on joint structure constraint and incomplete multimodal data can detect modality-shared and modality-specific biomarkers.

[0006] The above object of the present invention is achieved by the following technical measures:

[0007] Provide a method for non-linear correlation analysis of joint structure constraints and incomplete multimodal data, including the following steps:

[0008] Step 1: Collect image data of multiple objects respectively. For each object, image data of different modalities corresponding to the object are obtained through multiple imaging methods, and at the same time, genetic data of each object are collected;

[0009] Step 2: Process the image data of different modalities obtained in Step 1 according to the preprocessing method to obtain the processed images; process the genetic data obtained in Step 1 through the control and screening method to obtain the processed genetic data;

[0010] Step 3: Substitute the processed genetic data and the processed images into the objective function of the method for non-linear correlation analysis of joint structure constraints and incomplete multimodal data;

[0011] Step 4: By solving the objective function, the weights of SNPs and phenotypes on different modalities are obtained respectively.

[0012] Preferably, Step 1 is specifically to collect image data of multiple objects. For each object, an MRI image is obtained through magnetic resonance imaging, a PET image is obtained through positron emission tomography, and a DTI image is obtained through diffusion tensor imaging; at the same time, genetic data of each object are collected.

[0013] Preferably, Step 4 is specifically to solve the objective function by the alternating convex search method and the Lagrange multiplier method, and the weights of SNPs and phenotypes on different modalities are obtained respectively.

[0014] Preferably, the preprocessing method in Step 2 includes an MRI image preprocessing method, a PET image preprocessing method, and a DTI image preprocessing method.

[0015] Preferably, the quality control and screening method in Step 2 includes:

[0016] Step a.1: Perform quality control on the genetic data to obtain preprocessed genetic data;

[0017] Step a.2: Fill in and encode the original SNP genotypes in each preprocessed genetic data respectively to obtain the preprocessed genetic data after encoding and enter Step a.3;

[0018] Step a.3: Screen the preprocessed genetic data after encoding through a global independent screening process to obtain the processed genetic data after SNPs screening.

[0019] Preferably, the above-mentioned MRI image preprocessing method includes:

[0020] Step b.1: Use MIPAV software to perform anterior commissure and posterior commissure correction on the MRI images of each object respectively, and proceed to step b.2;

[0021] Step b.2: Use the N3 algorithm to correct the intensity inhomogeneity of the MRI images to obtain intensity-corrected images, and proceed to step b.3;

[0022] Step b.3: Delete the skull region and the cerebellar region, and proceed to step b.4;

[0023] Step b.4: Register the MRI images to the MNI space, and proceed to step b.5;

[0024] Step b.5: Segment the gray matter, white matter, cerebral lateral ventricles, and cerebrospinal fluid tissues in the MRI images to obtain a gray matter segmentation region, a white matter segmentation region, a cerebral lateral ventricle segmentation region, and a cerebrospinal fluid segmentation region, and proceed to step b.6;

[0025] Step b.6: Use the AAL atlas in the MNI space to label to obtain multiple ROIs, and proceed to step b.7;

[0026] Step b.7: Calculate the gray matter tissue volume for each of the multiple ROIs respectively to obtain multiple ROI volume data.

[0027] Preferably, the above-mentioned PET image preprocessing method is to align the PET images of each object with the corresponding MRI images respectively by using affine registration, and then calculate the average gray scale of each ROI as the PET feature.

[0028] Preferably, the above-mentioned DTI image preprocessing method includes:

[0029] Step c.1: The DTI image of each object contains 65 3D images, where the 65 3D images include one b0 image and 64 images with different gradient directions; use the dcm2niix tool to convert the 65 3D images into a 4D image, and generate a b-vector file and a b-value file respectively representing each gradient direction and its scalar value;

[0030] Step c.2: Apply the eddy command in the FSL package of the FMRIB software library to perform eddy current distortion correction on the 4D image in step c.1, and proceed to step c.3;

[0031] Step c.3: Use the BET algorithm in the FSL package to remove the skull from the b.0 image in step c.1, and proceed to step c.4;

[0032] Step c.4: Apply the difiti command in the FSL package, along with the b-vector file and b-value file generated in step c.1, to calculate the fractional anisotropy, which is defined as FA, and proceed to step c.5;

[0033] Step c.5: Register the b0 image to the MNI space through affine transformation, then apply the obtained transformation matrix to FA, and calculate the average density of each region of FA to obtain multiple ROI values.

[0034] Preferably, the above step a.1 includes:

[0035] Step a.1.1: Mark the genetic data of each object and each SNP data in the genetic data, then screen out the SNP data with a SNP detection rate greater than or equal to 95% and the corresponding genetic data, and proceed to step a.1.2;

[0036] Step a.1.2: Conduct a gender check on the genetic data corresponding to multiple objects, remove the genetic data with incorrect gender information and the corresponding MRI images, and proceed to step a.1.3;

[0037] Step a.1.3: Conduct a blood relationship check on the genetic data of each object separately, delete the genetic data of the object with a blood relationship to this object and the corresponding MRI images, and proceed to step a.1.4;

[0038] Step a.1.4: Delete the minor allele frequency in the genetic data, and proceed to step a.1.5;

[0039] Step a.1.5: Conduct a Hardy-Weinberg equilibrium test to obtain the preprocessed genetic data corresponding to the genetic data, and define the preprocessed genetic data as SNP data, and proceed to step a.1.6;

[0040] Step a.1.6: Apply the Minimmac4 software to perform genotype imputation on the SNP data obtained in step a.1.5, and proceed to step a.2.

[0041] Preferably, the above step a..2 is specifically to encode the original SNP genotypes in the preprocessed SNP data corresponding to the genetic data, and define the genetic data as preprocessed genetic data and proceed to step a.3.

[0042] Preferably, the above step a.3 steps are as follows:

[0043] Step a.3.1: Screen the SNP data in the preprocessed genetic data obtained in step a.2 respectively, and screen out the SNP data with a missing value greater than or equal to 5%, and proceed to step a.3.2;

[0044] Step a.3.2: Screen out the SNP data with a minor allele frequency less than or equal to 5%, and proceed to step a.3.3;

[0045] Step a.3.3: Screen out the SNP data with a Hardy-Weinberg equilibrium p-value less than 10 -6 . Define this SNP data as the processed genetic data, and at the same time define the SNP data of this processed genetic data as the processed SNP data;

[0046] Step a.3.4: Use the global sure independence screening procedure to screen the SNP data, and finally select the top 3000 or select the SNP loci with a p-value less than 0.1 as the genetic data input for the model.

[0047] Preferably, the above objective function is as shown in formula (I):

[0048]

[0049] where is the imaging phenotype data of the m-th modality; M is the total number of modalities; is the number of samples of the m-th modality, where n c and are the number of samples of the complete multi-modal phenotype data and the number of samples of the incomplete phenotype data of the m-th modality respectively; is the latent representation of the m-th modality; h is the feature dimension of the latent imaging representation; H c is the common latent representation of the samples with complete multi-modal phenotype data; is the independent latent imaging representation of the samples of the m-th modality in the incomplete multi-modal phenotype data; is the sparse error matrix of the m-th modality; is the correlation matrix of the learned phenotypic latent imaging representation of the m-th modality; is the correlation matrix of the learned phenotypic latent imaging representation; P T is the transpose matrix of P; is an identity matrix; is the SNP correlation matrix corresponding to the m-th modality phenotype; Ω(S) and Ω(Z) are the constraint terms for selecting relevant SNPs and imaging phenotypes; f is a non-linear transformation to construct the non-linear correlation between SNPs and phenotypes; L m = D m - C m is the Laplacian matrix; D m is a diagonal matrix where the i-th diagonal element represents the sum of the i-th row of C m ; C m is the similarity matrix of the m-th modality phenotype data; the (i, j)-th element is where Ym,:i and Y m,:j are the i-th column and the j-th column of Y m respectively, and set σ = 1; is the local fidelity projection.

[0050] Preferably, the above Ω(Z) is obtained by formula (II),

[0051]

[0052] where β1 and β2 are constraint term regulation parameters; is the connection penalty term; is the Laplacian matrix of the phenotypic connection matrix, l 21 is the norm.

[0053] Preferably, the above l 21 norm is obtained by formula (III),

[0054]

[0055] where Z m is the phenotypic correlation coefficient corresponding to the m-th modality; q is the number of phenotypic features; h is the number of features of the latent image representation; z m,ij is the number in the i-th row and j-th column of the correlation coefficient of the m-th modality.

[0056] Preferably, the above Ω(S) is obtained by formula (IV),

[0057]

[0058] where ||X - XU|| 21 is the graph self-expression constraint of SNP; is the group sparse constraint to explore the structural association between SNP groups; α1 and α2 are constraint term regulation parameters; ||U||1 is the sparse constraint of the object.

[0059] Preferably, the above G 21 norm is represented by formula (V),

[0060]

[0061] where the SNP data is divided into K groups p is the number of features of SNP loci.

[0062] Preferably, the above encoding method is to code the number of base pair mutations of the original SNP genotype as 0, 1, or 2 respectively.

[0063] Preferably, the above SNP detection rate is the ratio of the number of objects in which SNP loci are successfully detected to the total number of all objects.

[0064] Preferably, the blood relationship is at least one of parent-child relationship, brother relationship or sister relationship.

[0065] A method for nonlinear correlation analysis of joint structure constraint and incomplete multimodal data according to the present invention includes the following steps: Step 1: Collect image data of multiple objects respectively, where each object obtains image data of different modalities corresponding to the object through multiple imaging methods, and at the same time collect genetic data of each object; Step 2: Process the image data of different modalities obtained in Step 1 according to the preprocessing method to obtain the processed images; Process the genetic data obtained in Step 1 according to the control and screening method to obtain the processed genetic data; Step 3: Substitute the processed genetic data and the processed images into the objective function of the method for nonlinear correlation analysis of joint structure constraint and incomplete multimodal data; Step 4: By solving the objective function, obtain the weights of SNPs and phenotypes on different modalities respectively. Through the above 4 steps, the present invention obtains the weights corresponding to multimodal phenotype data and SNPs, constructs the nonlinear correlation between SNPs and phenotypes through nonlinear transformation, thereby considering the complex correlation between SNPs and phenotypes, and obtains modality-sharing and modality-specific biomarkers corresponding to different modalities through the contributions of multiple SNPs to phenotypes, so as to improve the performance of detecting biomarkers. Description of the Drawings

[0066] The present invention is further illustrated by the accompanying drawings, but the content in the drawings does not constitute any limitation to the present invention.

[0067] Figure 1 It is a flow chart of the method for nonlinear correlation analysis of joint structure constraint and incomplete multimodal data according to the present invention.

[0068] Figure 2 (a) is the original MRI image in the ADNI1 database, Figure 2 (b) is the original PET image in the ADNI1 database, Figure 2 (c) is the original MRI image in the PPMI database, Figure 2 (d) is the original DTI image in the PPMI database.

[0069] Figure 3 (a) is Figure 2 (a) The processed MRI image, Figure 3 (b) is Figure 2 (b) The processed PET image, Figure 3 (c) is Figure 2 (c) The processed MRI image, Figure 3 (d) is Figure 2 (d) The processed DTI image. Detailed implementation manners

[0070] The technical solution of the present invention will be further described in conjunction with the following embodiments.

[0071] Embodiment 1.

[0072] A method for non - linear correlation analysis of joint structure constraints and incomplete multimodal data, as Figure 1 shown, includes the following steps:

[0073] Step 1: Collect the image data of multiple objects respectively. For each object, the image data of different modalities of the corresponding object are obtained through multiple imaging methods. At the same time, use the Human 610 - Quad BeadChip to collect the genetic data of each object;

[0074] Step 2: Process the image data of different modalities obtained in Step 1 according to the pre - processing method to obtain the processed images; process the genetic data obtained in Step 1 through the control and screening method to obtain the processed genetic data;

[0075] Step 3: Substitute the processed genetic data and the processed images into the objective function of the method for non - linear correlation analysis of joint structure constraints and incomplete multimodal data;

[0076] Step 4: By solving the objective function, obtain the weights of SNPs and phenotypes on different modalities respectively.

[0077] Among them, Step 1 is specifically to collect the image data of multiple objects. For each object, the MRI image is obtained through the structural magnetic resonance imaging method, the PET image is obtained through the positron emission tomography method, and the DTI image is obtained through the diffusion tensor imaging method; at the same time, collect the genetic data of each object.

[0078] Among them, Step 4 is specifically to solve the objective function by the alternating convex search method and the Lagrange multiplier method, and obtain the weights of SNPs and phenotypes on different modalities respectively.

[0079] The pre - processing method in Step 2 of the present invention includes the MRI image pre - processing method, the PET image pre - processing method, and the DTI image pre - processing method.

[0080] The quality control and screening method in Step 2 of the present invention includes:

[0081] Step a.1: Perform quality control on the genetic data to obtain the pre - processed genetic data;

[0082] Step a.2: Fill in and encode the original SNP genotypes in each pre - processed genetic data respectively to obtain the encoded pre - processed genetic data and enter Step a.3;

[0083] Step a.3: Screen the preprocessed genetic data after encoding through a global independent screening process to obtain the processed genetic data after SNPs screening.

[0084] Among them, the MRI image preprocessing method includes:

[0085] Step b.1: Perform anterior commissure and posterior commissure correction on the MRI images of each object using MIPAV software, and enter Step b.2;

[0086] Step b.2: Correct the intensity inhomogeneity of the MRI image using the N3 algorithm to obtain an intensity-corrected image, and enter Step b.3;

[0087] Step b.3: Delete the skull region and cerebellum region, and enter Step b.4;

[0088] Step b.4: Register the MRI image to the MNI space, and enter Step b.5;

[0089] Step b.5: Segment the gray matter, white matter, cerebral lateral ventricles, and cerebrospinal fluid tissues in the MRI image to obtain a gray matter segmentation region, a white matter segmentation region, a cerebral lateral ventricle segmentation region, and a cerebrospinal fluid segmentation region, and enter Step b.6;

[0090] Step b.6: Use the AAL atlas in the MNI space to label to obtain multiple ROIs, and enter Step b.7;

[0091] Step b.7: Calculate the gray matter tissue volume for each of the multiple ROIs to obtain multiple ROI volume data.

[0092] Among them, the PET image preprocessing method is to align the PET images of each object with the corresponding MRI images respectively by using affine registration, and then calculate the average gray level of each ROI as the PET feature.

[0093] Among them, the DTI image preprocessing method includes:

[0094] Step c.1: Each object's DTI image contains 65 3D images, and the 65 3D images include a b0 image and 64 images with different gradient directions; use the dcm2niix tool to convert the 65 3D images into a 4D image, and generate a b-vector file and a b-value file respectively representing each gradient direction and its scalar value;

[0095] Step c.2: Apply the eddy command in the FSL package of the FMRIB software library to perform eddy current distortion correction on the 4D image in Step c.1, and enter Step c.3;

[0096] Step c.3: Use the BET algorithm in the FSL package to remove the skull from the b0 image in Step c.1, and proceed to Step c.4;

[0097] Step c.4: Apply the difiti command in the FSL package, along with the b-vector file and b-value file generated in Step c.1, to calculate the fractional anisotropy, where the fractional anisotropy is defined as FA, and proceed to Step c.5;

[0098] Step c.5: Register the b0 image to the MNI space through affine transformation, then apply the obtained transformation matrix to FA, and calculate the average density of each region of FA to obtain multiple ROI values.

[0099] Step a.1 of the present invention includes:

[0100] Step a.1.1: Mark the genetic data of each object and each SNP data in the genetic data, then screen out the SNP data with a SNP detection rate greater than or equal to 95% and the corresponding genetic data, and proceed to Step a.1.2;

[0101] Step a.1.2: Conduct a gender check on the genetic data corresponding to multiple objects, remove the genetic data with incorrect gender information and the corresponding MRI images, and proceed to Step a.1.3;

[0102] Step a.1.3: Conduct a blood relationship check on the genetic data of each object separately, delete the genetic data of the object related by blood to this object and the corresponding MRI images, and proceed to Step a.1.4;

[0103] Step a.1.4: Delete the minor allele frequency in the genetic data, and proceed to Step a.1.5;

[0104] Step a.1.5: Conduct a Hardy-Weinberg equilibrium test to obtain the preprocessed genetic data corresponding to the genetic data, define the preprocessed genetic data as SNP data, and proceed to Step a.1.6;

[0105] Step a.1.6: Use the Minimmac4 software to perform genotype imputation on the SNP data obtained in Step a.1.5, and proceed to Step a.2.

[0106] Among them, Step a..2 is specifically to encode the original SNP genotypes in the preprocessed SNP data corresponding to the genetic data, and define the genetic data as the preprocessed genetic data and proceed to Step a.3.

[0107] Among them, Step a.3 is as follows:

[0108] Step a.3.1: Screen the SNP data in the preprocessed genetic data obtained in step a.2 respectively, and screen out the SNP data with a missing value greater than or equal to 5%, then enter step a.3.2;

[0109] Step a.3.2: Screen out the SNP data with a minor allele frequency less than or equal to 5%, then enter step a.3.3;

[0110] Step a.3.3: Screen out the SNP data with a Hardy-Weinberg equilibrium p-value less than 10 –6 and define this SNP data as the processed genetic data, and at the same time define the SNP data of this processed genetic data as the processed SNP data;

[0111] Step a.3.4: Use the global sure independence screening procedure to screen the SNP data, and finally select the first 3000 or select the SNP loci with a p-value less than 0.1 as the genetic data input for the model.

[0112] It should be noted that the global sure independence screening procedure of the present invention for screening SNP data is proposed by Huang et al. in 2015 ((M. Huang et al., "FVGWAS: Fast voxelwise genome wide association analysis of large-scale imaging genetic data," (in eng), NeuroImage, vol. 118, pp. 613-627, 2015.).

[0113] The objective function of the present invention is shown in formula (I):

[0114]

[0115] where is the imaging phenotype data of the m-th modality; M is the total number of modalities; is the number of samples of the m-th modality, where n c and are respectively the number of samples of the complete multi-modal phenotype data and the number of samples of the incomplete m-th modality phenotype data; is the latent representation of the m-th modality; h is the feature dimension of the latent imaging representation; H c is the common latent representation of the samples with complete multi-modal phenotype data; is the independent latent imaging representation of the samples of the m-th modality in the incomplete multi-modal phenotype data; is the sparse error matrix of the m-th modality; The association matrix for the phenotypic latent image representation of the m-th learned modality; The association matrix for the learned phenotypic latent image representation; P T Is the transpose matrix of P; Is an identity matrix; Is the SNP association matrix corresponding to the m-th modality phenotype; Ω(S) and Ω(Z) are constraint terms for selecting relevant SNPs and image phenotypes; f is a non-linear transformation to construct the non-linear association between SNPs and phenotypes; L m = D m - C m Is the Laplacian matrix; D m Is a diagonal matrix where the i-th diagonal element represents the sum of the i-th row of C; C m ; C m Is the similarity matrix of the m-th modality phenotype data; the (i,j)-th element is where Y m,:i and Y m,:j are the i-th column and j-th column of Y m respectively, and σ = 1 is set; Is the local fidelity projection.

[0116] Since it is considered that the association between SNPs and phenotypes is complex, it is difficult to fit the complex relationship between SNPs and phenotypes with a simple linear model alone. Therefore, the present invention introduces a non-linear transformation to consider this association information. The present invention constructs the non-linear association between SNPs and phenotypes through non-linear transformation, considering the complex association between SNPs and phenotypes, thereby being able to improve the performance of detecting biomarkers.

[0117] It should be noted that when the local fidelity projection is applied to the model, it can keep the neighborhood structure information unchanged before and after projection.

[0118] Ω(Z) of the present invention is obtained through Equation (II),

[0119]

[0120] where β1 and β2 are constraint term regulation parameters; Is the connection penalty term; Is the Laplacian matrix of the phenotypic connection matrix, l 21 Is the norm.

[0121] It should be noted that, The purpose of setting it is to construct the structural information between phenotypes, l 21 The norm is used to remove phenotypes irrelevant to the task to obtain a sparse phenotypic association matrix.

[0122] where, l 21The norm is obtained through formula (III),

[0123]

[0124] where Z m is the phenotypic correlation coefficient corresponding to the m-th modality; q is the number of features of the phenotype; h is the number of features of the latent image representation; z m,ij is the number in the i-th row and j-th column of the correlation coefficient of the m-th modality.

[0125] SNPs within a gene usually perform the same genetic function. In addition, in 2005, linkage disequilibrium proposed by Barrett et al. (Barrett, J.C., Fry, B., Maller, J., Daly, M.J., 2005. Haploview: analysis and visualization of LD and haplotype maps. Bioinformatics 21, 263 - 265.) described the non-random association between alleles at different loci. During meiosis, SNPs with high linkage disequilibrium are linked together through this association. Therefore, these information should be considered in the realistic modeling method of the present invention, and formula (IV) is used to define the SNP data.

[0126] Specifically, Ω(S) is obtained through formula (IV),

[0127]

[0128] where ||X - XU|| 21 is the graph self-expression constraint of the SNP; is the group sparse constraint to explore the structural association between SNP groups; α1 and α2 are constraint term regulation parameters; ||U||1 is the sparse constraint of the object.

[0129] It should be noted that the role of ||U||1 is to remove irrelevant SNP loci. ||X - XU|| 21 is used to construct the structural association between each SNP locus. Since there is a group effect among SNP data, the group sparse constraint is applied, to guide the previous graph self-expression constraint to construct the structural association between SNP groups.

[0130] The G 21 norm of the present invention is represented by formula (V),

[0131]

[0132] where the SNP data is divided into K groups p is the number of features of the SNP locus.

[0133] In step b.5, when selecting the gray matter segmentation region in step b.4 from the intensity-corrected image obtained in step b.3, 90 ROIs of MRI images are obtained through the AAL template anatomical information.

[0134] The encoding method of the present invention is to code the number of base pair mutations of the original SNP genotype as 0, 1, or 2 respectively.

[0135] The SNP detection rate of the present invention is the ratio of the number of objects in which the SNP locus is successfully detected to the total number of all objects.

[0136] The blood relationship of the present invention is at least one of parent-child relationship, brother relationship, or sister relationship.

[0137] The method for non-linear association analysis based on joint structure constraint and incomplete multi-modal data of the present invention first processes the missing data through a multi-constraint joint parallel connection projection method while considering the intra-modal and inter-modal association information and being able to learn the modality-shared and modality-specific information. Secondly, the present invention separately considers the structural association information between SNPs and between imaging phenotypes by adding structural constraints. Finally, the present invention considers the non-linear association between SNPs and phenotypes, as well as the modality-shared and modality-specific biomarkers by introducing a kernel-based non-linear model. Therefore, by applying the present invention, modality-shared and modality-specific biomarkers can be detected.

[0138] The method for non-linear association analysis based on joint structure constraint and incomplete multi-modal data obtains the weights corresponding to multiple modality phenotypes and SNPs through 4 steps, constructs the non-linear association between SNPs and phenotypes through non-linear transformation, thus considering the complex association between SNPs and phenotypes, and obtains the modality-shared and modality-specific biomarkers corresponding to different modalities through the contributions of multiple SNPs to the phenotype, thereby being able to improve the performance of detecting biomarkers.

[0139] Example 2.

[0140] A method for non-linear association analysis based on joint structure constraint and incomplete multi-modal data, as Figure 2 and Figure 3 , includes the steps of: downloading the T1-weighted MRI images and PET images of ADNI 1 from the ADNI database, and downloading the T1-weighted MRI images and DTI images from the PPMI database. Then, candidate genes are screened out by applying the global sure independence screening procedure. In this example, the first 3000 SNP data are selected as genetic data.

[0141] The following details the preprocessing methods for each MRI, PET, and DTI image and genetic data in the database.

[0142] Step 1: Download MRI and PET images as well as genetic data from the ADNI database, and download MRI and DTI images as well as genetic data from the PPMI database.

[0143] Step 2: Process the image data of different modalities obtained in Step 1 according to the preprocessing method respectively to obtain the processed images; process the genetic data obtained in Step 1 by means of the control and screening method to obtain the processed genetic data.

[0144] In this embodiment, first, each MRI image is preprocessed to obtain the processed MRI image, and at the same time, the genetic data corresponding to the MRI image is subjected to quality control and screening to obtain the processed genetic data.

[0145] Among them, the MRI image preprocessing method includes:

[0146] Step b.1: Use the MIPAV software to perform anterior commissure and posterior commissure correction on the MRI images of each object respectively, and enter Step b.2;

[0147] Step b.2: Use the N3 algorithm to correct the intensity inhomogeneity of the MRI image to obtain the intensity-corrected image, and enter Step b.3;

[0148] Step b.3: Apply a robust brain extraction algorithm to remove the skull region, and distort the labeled template on each skull-stripped image to remove the cerebellar region, and enter Step b.4;

[0149] Step b.4: Use the Advanced Normalization Tools to register the MRI image to the MNI space, and enter Step b.5;

[0150] Step b.5: Use the Atropos algorithm for tissue segmentation to segment the gray matter, white matter, cerebral lateral ventricles and cerebrospinal fluid tissues in the MRI image to obtain the gray matter segmentation region, white matter segmentation region, cerebral lateral ventricle segmentation region and cerebrospinal fluid segmentation region, and enter Step b.6;

[0151] Step b.6: Use the AAL atlas in the MNI space to label to obtain 90 ROIs, and enter Step b.7;

[0152] Step b.7: Calculate the gray matter tissue volume for each of the 90 ROIs respectively to obtain a plurality of ROI volume data.

[0153] Therefore, for each MRI image, the MRI image preprocessing method extracts the feature vectors of 90 gray matter tissue volumes as one of the phenotypic data of the objective function of the present invention.

[0154] The PET image preprocessing method is as follows: for each PET image, first align the PET image with the corresponding T1-weighted MRI image through affine registration, and then calculate the average PET intensity value of each ROI as the ROI feature.

[0155] The DTI image preprocessing method includes:

[0156] Step c.1: The DTI image of each subject contains 65 3D images, where the 65 3D images include a b0 image and 64 images with different gradient directions. Use the dcm2niix tool to convert the 65 3D images into a 4D image, and generate a b-vector file and a b-value file respectively representing each gradient direction and its scalar value;

[0157] Step c.2: Apply the eddy command in the FSL package of the FMRIB software library to correct the eddy current distortion of the 4D image, and enter step c.3;

[0158] Step c.3: Use the BET algorithm of the FSL package to remove the skull from the b0 image in step c.1, and enter step c.4;

[0159] Step c.4: Apply the difiti command of the FSL package and the generated files to calculate the fractional anisotropy, i.e., FA;

[0160] Step c.5: Register the b0 image to the MNI space through affine transformation, and then apply the obtained transformation matrix to the FA to calculate the average density of each region of the FA to obtain multiple ROI values.

[0161] The quality control and screening method in step two includes:

[0162] Step a.1 includes:

[0163] Step a.1.1: Mark the genetic data of each subject and each single nucleotide polymorphism (i.e., SNP) data in the genetic data, and then screen out the SNP data with a SNP detection rate greater than or equal to 95% and the corresponding genetic data. Specifically, check the detection rate for each subject and each SNP marker. For example, the SNP detection rate refers to the ratio of the samples in which a certain SNP locus is successfully detected to all samples, generally required to be above 95%, and enter step a.1.2;

[0164] Step a.1.2: Conduct a gender check on the genetic data corresponding to multiple subjects, and remove the genetic data with incorrect gender information and the corresponding MRI images, and enter step a.1.3;

[0165] Step a.1.3: Check the genetic relationship of the genetic data of each object respectively, delete the genetic data of the object related by blood and the corresponding MRI images, and proceed to step a.1.4;

[0166] Step a.1.4: Delete the minor allele frequency in the genetic data, and proceed to step a.1.5;

[0167] Step a.1.5: Conduct a Hardy-Weinberg equilibrium test to obtain the preprocessed genetic data corresponding to the genetic data, define the preprocessed genetic data as SNP data, and proceed to step a.1.6;

[0168] Step a.1.6: Apply the Minimmac4 software to perform genotype imputation on the SNP data obtained in step a.1.5, and proceed to step a.2.

[0169] Step a.2: Encode the original SNP genotypes in each preprocessed genetic data respectively. Specifically, encode the SNP original data (C, T, G, A) as 0, 1, 2, define the genetic data as preprocessed genetic data, and at the same time remove some factors that may cause bias, and proceed to step a.3;

[0170] Step a.3: Screen the preprocessed genetic data after encoding through a global independent screening process to obtain the processed genetic data after SNPs screening. Subsequently, in the further preprocessing process, remove some single nucleotide polymorphisms (SNPs) according to the following conditions.

[0171] Among them, step a.3 is specifically divided into:

[0172] Step a.3.1: Screen the SNP data in the preprocessed genetic data obtained in step a.2 respectively, and screen out the SNP data with a missing value greater than or equal to 5%, and proceed to step a.3.2;

[0173] Step a.3.2: Screen out the SNP data with a minor allele frequency less than or equal to 5%, and proceed to step a.3.3;

[0174] Step a.3.3: Screen out the SNP data with a Hardy-Weinberg equilibrium p-value less than 10 -6 and define the SNP data as the processed genetic data, and at the same time define the SNP data of the processed genetic data as the processed SNP data;

[0175] Step a.3.4: Adopt the global sure independence screening process proposed by Huang et al. to select candidate genes, and obtain 3000 SNP data in the ADNI and PPMI data sets respectively.

[0176] Step 3: After preprocessing, 708 subjects can be obtained from the ADNI database and 512 subjects can be obtained from the PPMI database. Substitute the processed genetic data and processed multimodal images into the objective function of the method for nonlinear association analysis of joint structure constraints and incomplete multimodal data for association analysis. The objective function is the ScCNAA model constructed using image data and genetic data:

[0177]

[0178] where is the imaging phenotype data of the m-th modality; M is the total number of modalities; is the number of samples of the m-th modality, where n c and are the number of samples of the complete multimodal phenotype data and the number of samples of the incomplete m-th modality phenotype data, respectively; is the latent representation of the m-th modality; h is the feature dimension of the latent imaging representation; H c is the common latent representation of the samples with complete multimodal phenotype data; is the independent latent imaging representation of the samples of the m-th modality in the incomplete multimodal phenotype data; is the sparse error matrix of the m-th modality; is the association matrix of the learned phenotypic latent imaging representation of the m-th modality; is the association matrix of the learned phenotypic latent imaging representation; P T is the transpose matrix of P; is an identity matrix; is the SNP association matrix corresponding to the m-th modality phenotype; Ω(S) and Ω(Z) are constraint terms for selecting relevant SNPs and imaging phenotypes; f is a nonlinear transformation to construct the nonlinear association between SNPs and phenotypes; L m = D m - C m is the Laplacian matrix; D m is a diagonal matrix where the i-th diagonal element represents the sum of the i-th row of C m ; C m is the similarity matrix of the m-th modality phenotype data; the (i, j)-th element is where Y m,:i and Y m,:j are the i-th column and j-th column of Y m , respectively, and σ = 1 is set; is the local fidelity projection.

[0179] where Ω(Z) is obtained through Equation (II). Consider the structural association information between brain regions and the sparsity at the individual level through Ω(Z):

[0180]

[0181] Among them, β1 and β2 are constraint term regulation parameters. Represents the connection penalty term, and the structural information between phenotypes can be considered; is the Laplacian matrix of the phenotype connection matrix. l 21 The norm is used to remove phenotypes irrelevant to the task to obtain a sparse phenotype association matrix.

[0182] l 21 The norm is obtained through Equation (III),

[0183]

[0184] For l 21 The role of the norm is to remove phenotype regions irrelevant to the task and only retain phenotype regions relevant to the task.

[0185] Ω(S) is obtained through Equation (IV),

[0186]

[0187] where ||X - XU|| 21 is the graph self-expression constraint of SNPs; is the group sparse constraint to explore the structural association between SNP groups; α1 and α2 are constraint term regulation parameters; ||U||1 is the sparse constraint of the object.

[0188] G 21 The norm is represented by Equation (V),

[0189]

[0190] where the SNP data is divided into K groups The graph self-expression constraint of SNPs is to consider the association information between SNP sites. Here, the present invention applies G 21 to guide the graph self-expression constraint to form a group-guided graph self-expression constraint to consider the intra-group and inter-group structural associations of SNPs. The dimension of SNPs is large, but only a small part of SNPs are relevant to the task. Therefore, the l1 norm is applied to the model to remove SNP sites irrelevant to the task in order to improve the detection performance.

[0191] Step 4: Solve the objective function by the alternating convex search method and the Lagrange multiplier method to obtain the weights of SNPs and ROIs corresponding to different modalities.

[0192] In the model of the present invention, the hyperparameters are determined by selecting the minimum root mean square error (RMSE). The optimal parameters are determined in this set of data: by using the alternating convex search method to solve this objective function, the values of the weights S and Z corresponding to SNP and ROI can be obtained, which respectively correspond to the ROI and SNP features. Since the obtained weights are sparse, the top 20 task-related ROIs and SNPs are selected by sorting the absolute values of the weights from large to small. And the application of the minimum root mean square error, that is, RMSE, is used as the evaluation index of the model to judge whether the model is feasible. The smaller the RMSE, the better the model is considered.

[0193] A comparison was made with other models of the prior art. In the ADNI dataset, the RMSE of the method based on multi-task sparse canonical correlation analysis is 0.13, and the RMSE of the method based on multi-task regression and feature selection is 4.3. The RMSE of the present invention is 0.025. In the PPMI dataset, the RMSE of the method based on multi-task sparse canonical correlation analysis is 0.16, and the RMSE of the methods based on multi-task regression and feature selection are both 5.2. The RMSE of the present invention is 0.045. Therefore, the RMSE of the present invention is the smallest, indicating that the present invention has better effects compared with the prior art. The present invention considers the group structure association of SNPs, so as to be able to detect potential biomarkers of tasks more accurately.

[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A method for non-linear correlation analysis of joint structure constraints and incomplete multi-modal data, characterized in that The steps include: Step 1: Collect the image data of multiple objects respectively. For each object, the image data of different modalities corresponding to the object are obtained through multiple imaging methods, and at the same time, the genetic data of each object are collected. Step 2: Process the image data of different modalities obtained in Step 1 according to the preprocessing method to obtain the processed images; process the genetic data obtained in Step 1 through the control and screening method to obtain the processed genetic data. Step 3: Substitute the processed genetic data and the processed images into the objective function based on the joint structure constraint and the nonlinear correlation analysis method of incomplete multimodal data. Step 4: By solving the objective function, obtain the weights of SNPs and phenotypes on different modalities respectively. The objective function is shown in Equation (I): where is the imaging phenotypic data of the m-th modality; M is the total number of modalities; is the number of samples of the m-th modality, where n c and are the number of samples of the complete multi-modal phenotypic data and the number of samples of the incomplete m-th modality phenotypic data, respectively; is the latent representation of the m-th modality; h is the feature dimension of the latent imaging representation; H c is the common latent representation of the samples with complete multi-modal phenotypic data; is the independent latent imaging representation of the samples of the m-th modality in the incomplete multi-modal phenotypic data; is the sparse error matrix of the m-th modality; is the correlation matrix of the learned phenotypic latent imaging representation of the m-th modality; is the correlation matrix of the learned phenotypic latent imaging representation; P T is the transpose matrix of P; is an identity matrix; is the SNP correlation matrix corresponding to the m-th modality phenotype; Ω(S) and Ω(Z) are the constraint terms for selecting relevant SNPs and imaging phenotypes; f is a non-linear transformation to construct the non-linear correlation between SNPs and phenotypes; L m = D m - C m is the Laplacian matrix; D m is a diagonal matrix where the i-th diagonal element represents the sum of the i-th row of C m ; C m is the similarity matrix of the m-th modality phenotypic data; the (i, j)-th element is where Y m,:i and Y m,:j are the i-th column and the j-th column of Y m , respectively, and σ = 1 is set; is the local fidelity projection; The Ω(Z) is obtained through Equation (II), Among them, β1 and β2 are constraint term regulation parameters; is the connection penalty term; is the Laplacian matrix of the phenotypic connection matrix, l 21 is the norm; the said l 21 The norm is obtained through Equation (III), Among them, Z m is the phenotypic correlation coefficient corresponding to the m-th modality; q is the number of phenotypic features; h is the number of features of the potential image representation; z m,ij is the number in the i-th row and j-th column of the correlation coefficient of the m-th modality; The Ω(S) is obtained through Equation (IV), where ||X - Xu|| 21 is the graph self-expression constraint of SNPs; is the group sparse constraint to explore the structural associations among SNP groups; α1 and α2 are constraint term regulation parameters; ||U||1 is the sparse constraint of the object; The said G 21 norm is represented by Equation (V), Among them, the SNP data is divided into K groups p is the number of features of the SNP locus.

2. The non - linear correlation analysis method based on joint structure constraints and incomplete multi - modal data according to claim 1, wherein: Specifically, Step 1 is to collect the image data of multiple objects. For each object, the MRI image is obtained through the structural magnetic resonance imaging method, the PET image is obtained through the positron emission tomography method, and the DTI image is obtained through the diffusion tensor imaging method; at the same time, the genetic data of each object are collected. Specifically, Step 4 is to solve the objective function through the alternating convex search method and the Lagrange multiplier method, and obtain the weights of SNPs and phenotypes on different modalities respectively.

3. The method for nonlinear correlation analysis of incomplete multimodal data based on joint structure constraint according to claim 2, characterized in that: The preprocessing method in Step 2 includes the MRI image preprocessing method, the PET image preprocessing method, and the DTI image preprocessing method; The quality control and screening method in Step 2 includes: Step a.1: Perform quality control on the genetic data to obtain the preprocessed genetic data; Step a.2: Fill in and encode the original SNP genotypes in each preprocessed genetic data respectively to obtain the encoded preprocessed genetic data and enter Step a.3; Step a.3: Screen the encoded preprocessed genetic data through the global independent screening process to obtain the processed genetic data after SNPs screening.

4. The method for non - linear correlation analysis of joint - structure - constraint - based and incomplete multi - modal data according to claim 3, wherein: The MRI image preprocessing method includes: Step b.1: Use the MIPAV software to perform anterior commissure and posterior commissure correction on the MRI images of each object respectively, and enter Step b.2; Step b.2: Use the N3 algorithm to correct the intensity inhomogeneity of the MRI images to obtain the intensity-corrected images, and enter Step b.3; Step b.3: Delete the skull region and the cerebellum region, and enter Step b.4; Step b.4: Register the MRI images to the MNI space, and enter Step b.5; Step b.5: Segment the gray matter, white matter, cerebral lateral ventricles, and cerebrospinal fluid tissues in the MRI images to obtain the gray matter segmentation region, white matter segmentation region, cerebral lateral ventricle segmentation region, and cerebrospinal fluid segmentation region, and enter Step b.6; Step b.6: Use the AAL atlas in the MNI space to label to obtain multiple ROIs, and enter Step b.7; Step b.7: Calculate the gray matter tissue volume for multiple ROIs respectively to obtain multiple ROI volume data; The PET image preprocessing method is to align the PET images of each object with the corresponding MRI images by using affine registration respectively, and then calculate the average gray value of each ROI as the PET feature; The DTI image preprocessing method includes: Step c.1: The DTI image of each object contains 65 3D images, among which the 65 3D images include a b0 image and 64 images with different gradient directions; Use the dcm2niix tool to convert the 65 3D images into a 4D image, and generate a b-vector file and a b-value file respectively representing each gradient direction and its scalar value; Step c.2: Apply the eddy command of the FSL package in the FMRIB software library to correct the eddy current distortion of the 4D image in Step c.1, and enter Step c.3; Step c.3: Use the BET algorithm of the FSL package to remove the skull from the b.0 image in Step c.1, and enter Step c.4; Step c.4: Apply the difiti command of the FSL package and the b-vector file and b-value file generated in Step c.1 to calculate the fractional anisotropy, where the fractional anisotropy is defined as FA, and enter Step c.5; Step c.5: Register the b0 image to the MNI space through affine transformation, then apply the obtained transformation matrix to FA, and calculate the average density of each region of FA to obtain multiple ROI values.

5. The method for nonlinear correlation analysis of joint structure constraint and incomplete multimodal data according to claim 4, wherein: The said Step a.1 includes: Step a.1.1: Mark the genetic data of each object and each SNP data in the genetic data, and then screen out the SNP data with a SNP detection rate greater than or equal to 95% and the corresponding genetic data, and enter Step a.1.2; Step a.1.2: Conduct a gender check on the genetic data corresponding to multiple objects, and remove the genetic data with incorrect gender information and the corresponding MRI images, and enter Step a.1.3; Step a.1.3: Conduct a blood relationship check on the genetic data of each object respectively, and delete the genetic data of the object with blood relationship to this object and the corresponding MRI images, and enter Step a.1.4; Step a.1.4: Delete the minor allele frequency in the genetic data, and enter Step a.1.5; Step a.1.5: Conduct a Hardy-Weinberg equilibrium test to obtain the preprocessed genetic data corresponding to the genetic data, and define the preprocessed genetic data as SNP data, and enter Step a.1.6; Step a.1.6: Apply the Minimmac4 software to fill in the genotypes of the SNP data obtained in Step a.1.5, and enter Step a.

2.

6. The method for nonlinear association analysis based on joint structure constraint and incomplete multi-modal data according to claim 5, characterized in that: The said Step a.2 is specifically to encode the original SNP genotypes in the preprocessed SNP data corresponding to the genetic data, and define the genetic data as preprocessed genetic data and enter Step a.3; The said Step a.3 steps are: Step a.3.1: Screen the SNP data in the preprocessed genetic data obtained in step a.2 respectively, and screen out the SNP data with a missing value greater than or equal to 5%, then proceed to step a.3.2; Step a.3.2: Screen out the SNP data with a minor allele frequency less than or equal to 5%, then proceed to step a.3.3; Step a.3.3: Screen out SNP data with a Hardy-Weinberg equilibrium p-value less than 10 -6 , define this SNP data as processed genetic data, and at the same time define the SNP data of this processed genetic data as processed SNP data; Step a.3.4: Use the global determination of independence screening process to screen the SNP data, and finally select the first 3000 or the SNP loci with a p-value less than 0.1 as the genetic data input for the model.

7. The method for non-linear correlation analysis of combined structure constraints and incomplete multi-modal data according to claim 6, characterized in that: The encoding method is to code the number of base pair mutations of the original SNP genotype as 0, 1, or 2 respectively; The SNP detection rate is the ratio of the number of objects in which the SNP locus is successfully detected to the total number of all objects; The blood relationship is at least one of parent-child relationship, brother relationship, or sister relationship.

Citation Information

Patent Citations

  • Multimodal molecular imaging strategy for radiation-induced brain injury biomarkers after nasopharyngeal carcinoma radiotherapy

    CN110833414A

  • Multi-modal neural image feature selection method based on sample weight and low-rank constraint

    CN111916162A