Method and system for screening colon cancer diagnosis markers based on transcriptome data

By using dynamic discrimination and dual-layer feature fusion for pathological stratification modeling and pathological subtype difference analysis, the problems of sample heterogeneity and noise influence in the screening of colorectal cancer diagnostic biomarkers were solved, achieving more accurate and stable biomarker screening.

CN121460122APending Publication Date: 2026-02-03固原市人民医院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511555059.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

In the screening of diagnostic biomarkers for colorectal cancer, samples exhibit high heterogeneity in gene expression and pathological characteristics. Traditional methods struggle to accurately classify pathological subtypes, and noisy samples affect the stability and representativeness of biomarker screening results.

Method used

We employed a pathological stratification modeling approach based on dynamic discrimination and dual-layer feature fusion to adaptively identify pathological subtypes of colorectal cancer samples. Through pathological subtype difference analysis, we screened out a set of diagnostic biomarkers with significant and stable expression differences.

Benefits of technology

This improved the accuracy and biological representativeness of diagnostic biomarker screening, and enhanced the robustness and diagnostic specificity of the biomarkers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121460122A_ABST
    Figure CN121460122A_ABST
Patent Text Reader

Abstract

The invention discloses a colon cancer diagnostic marker screening method and system based on transcriptome data, and belongs to the technical field of biological information, and the method comprises the steps of data preparation, pathological hierarchical modeling, diagnostic marker screening and screening report generation. According to the method, pathological hierarchical modeling based on dynamic discrimination and double-layer feature fusion is adopted, on the basis of comprehensively considering gene expression and pathological morphology information, pathological subtypes of colon cancer samples are adaptively recognized, stable and representative subtype features are obtained, and therefore the accuracy and biological representativeness of diagnostic marker screening are improved; diagnostic marker screening based on pathological subtype difference analysis is adopted, and on the premise of considering pathological subtype characteristics, a stable and reliable diagnostic marker set with remarkable expression difference is screened in a targeted manner, so that the biological representativeness, screening robustness and diagnostic specificity of markers are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of biological information technology, and particularly relates to a colon cancer diagnosis marker screening method and system based on transcriptome data. BACKGROUND

[0002] The colon cancer diagnosis marker screening based on transcriptome data aims to obtain colon cancer-specific genes with diagnostic potential and form a colon cancer diagnosis marker set by systematically analyzing the transcriptome data of colon cancer and normal tissues, combining pathological subtype information and sample characteristic stability, and performing differential evaluation and stability screening on candidate genes, so as to provide reliable molecular markers and data basis for early diagnosis, typing and individualized treatment of colon cancer.

[0003] However, in the process of colon cancer diagnosis marker screening, there are technical problems that the samples have high heterogeneity in gene expression and pathological characteristics, traditional single feature or fixed grouping method is difficult to accurately divide pathological subtypes, and noise samples may affect the stability and representativeness of the screening results; there are technical problems that the colon cancer samples are highly heterogeneous, the gene expression of different pathological subtypes is obviously different, and traditional whole-sample differential analysis is easily disturbed by noise or subtype internal difference, resulting in lack of stability and specificity of the screened markers. SUMMARY

[0004] In view of the above problems, the colon cancer diagnosis marker screening method and system based on transcriptome data are provided to overcome the defects of the prior art. The pathological stratification modeling based on dynamic discrimination and double-layer feature fusion is creatively adopted, the pathological subtypes of colon cancer samples are adaptively identified on the basis of comprehensive consideration of gene expression and pathological morphology information, and stable and representative subtype features are obtained, so that the accuracy and biological representativeness of the diagnosis marker screening are improved. The diagnosis marker screening based on pathological subtype differential analysis is creatively adopted, and the diagnosis marker set with significant and stable expression difference is screened out on the premise of considering the pathological subtype features, so that the biological representativeness, screening robustness and diagnostic specificity of the markers are improved.

[0005] The technical solutions adopted by the present application are as follows: The colon cancer diagnosis marker screening method based on transcriptome data provided by the present application comprises the following steps:

[0006] Step S1: data preparation;

[0007] Step S2: pathological stratification modeling;

[0008] Step S3: diagnosis marker screening;

[0009] Step S4: screening report generation.

[0010] Further, in step S1, the data preparation is used to provide a data basis for pathological stratification modeling and diagnostic marker screening, specifically, data acquisition is performed first, and then a sample gene expression matrix and a sample pathological feature matrix are constructed.

[0011] Further, in step S2, the pathological stratification modeling is used to realize adaptive identification and feature extraction of pathological subtypes, specifically, based on the sample gene expression matrix and the sample pathological feature matrix, a pathological stratification modeling based on dynamic discrimination and double-layer feature fusion is adopted to obtain pathological subtype information, including the following steps:

[0012] Step S21: Feature embedding modeling is used to represent the multi-dimensional relationship of samples in the gene expression layer and the pathological feature layer, specifically, gene expression similarity matrix and pathological feature similarity matrix are calculated through a similarity function, and adaptive weighted fusion is performed based on an information entropy balance strategy to generate a sample fusion similarity matrix;

[0013] The gene expression similarity matrix is used to reflect the similarity degree of different samples in the gene expression mode;

[0014] The pathological feature similarity matrix is used to reflect the similarity degree of different samples in the pathological feature mode;

[0015] The sample fusion similarity matrix is used to comprehensively reflect the overall similarity degree of different samples in the double-layer space of gene expression features and pathological features;

[0016] Step S22: Adaptive clustering discrimination is used to divide pathological subtypes based on the sample fusion similarity matrix, specifically, a graph Laplacian matrix is constructed, and feature decomposition is performed to obtain a spectral feature matrix, then a dynamic cluster discrimination function is set, the optimal clustering number is calculated by maximizing the dynamic cluster discrimination function, and the samples are clustered in the spectral feature matrix using a k-means clustering algorithm to obtain sample subtype labels;

[0017] Step S23: Subtype feature integration is used to generate representative expression features of each pathological subtype to suppress the influence of noise samples while maintaining the representativeness of the subtype, specifically, the samples of the same subtype are weighted and integrated to obtain a subtype expression feature matrix;

[0018] The subtype expression feature matrix is used to comprehensively reflect the representative gene expression features of each pathological subtype;

[0019] Step S24: Pathological subtype information generation, specifically, by executing steps S21 to S23, pathological subtype information is generated, including sample subtype labels and subtype expression feature matrix.

[0020] Further, in step S3, the diagnostic marker screening is performed, specifically, based on the pathological subtype information, a diagnostic marker screening based on pathological subtype difference analysis is performed to obtain marker screening information, including the following steps:

[0021] Step S31: subtype difference statistical analysis, used for quantitative evaluation of the expression difference of each gene between different pathological subtypes, specifically, by taking the sample gene expression matrix as the dependent variable and the sample subtype label as the independent variable, a linear model of gene expression difference is established, and the expression difference coefficient is fitted, and then based on the expression difference coefficient and the model residual, the expression difference significance of each gene between subtypes is calculated by F test, and a difference significance matrix is generated;

[0022] Step S32: adaptive difference screening, used for preliminary screening of a candidate difference gene set, specifically, the adaptive difference score of each gene is calculated based on the average expression difference and the expression standard deviation within the subtype, and a candidate difference gene set is screened based on the adaptive threshold;

[0023] Step S33: stability and redundancy evaluation, used for screening a stable and representative gene set, specifically, unstable genes are removed by Bootstrap resampling and gene expression stability score, and redundant genes are removed by calculating the Spearman rank correlation coefficient, to obtain a candidate stable gene set;

[0024] Step S34: comprehensive score sorting, specifically, the comprehensive diagnostic potential score of each gene is obtained by weighting and fusing the difference significance and stability score, and the top A genes are selected according to the comprehensive diagnostic potential score, and a colon cancer diagnostic marker set is generated;

[0025] Step S35: screening information generation, specifically, the marker screening information is generated by executing steps S31 to S34, including the colon cancer diagnostic marker set, the comprehensive diagnostic potential score, the difference significance matrix and the gene expression stability score.

[0026] Further, in step S4, the screening report is generated, specifically, based on the marker screening information, a colon cancer diagnostic marker screening report is generated, the screening results are displayed in a table, and the comprehensive diagnostic potential score, the difference significance, the stability score and the corresponding pathological subtype information of each gene are provided.

[0027] The colon cancer diagnostic marker screening system based on transcriptome data provided by the application comprises a data preparation module, a pathological stratification module, a marker screening module and a report generation module.

[0028] The data preparation module is used for data preparation, and through data preparation, a sample gene expression matrix and a sample pathological feature matrix are obtained, the sample gene expression matrix is sent to the pathological stratification module and the marker screening module, and the sample pathological feature matrix is sent to the pathological stratification module;

[0029] The pathological stratification module is used for pathological stratification modeling, and through pathological stratification modeling, pathological subtype information is obtained, and the pathological subtype information is sent to the marker screening module and the report generation module;

[0030] The marker screening module is used for diagnostic marker screening, and through diagnostic marker screening, marker screening information is obtained, and the marker screening information is sent to the report generation module;

[0031] The report generation module is used for screening report generation, and a colon cancer diagnostic marker screening report is obtained.

[0032] The above scheme has the following beneficial effects:

[0033] (1) In the process of colon cancer diagnostic marker screening, there is high heterogeneity in sample gene expression and pathological features, and traditional single feature or fixed grouping method cannot accurately divide pathological subtypes, and noise samples may affect the stability and representativeness of the marker screening result. To solve the technical problem, the present scheme creatively adopts pathological stratification modeling based on dynamic discrimination and double-layer feature fusion, adaptively identifies the pathological subtypes of colon cancer samples on the basis of comprehensive consideration of gene expression and pathological morphology information, and obtains stable and representative subtype features, thereby improving the accuracy and biological representativeness of diagnostic marker screening.

[0034] (2) In the process of colon cancer diagnostic marker screening, there is high heterogeneity in colon cancer samples, and the gene expression difference between different pathological subtypes is obvious, and traditional whole-sample difference analysis is easily disturbed by noise or subtype internal difference, resulting in lack of stability and specificity of the screened markers. To solve the technical problem, the present scheme creatively adopts diagnostic marker screening based on pathological subtype difference analysis, and under the premise of considering pathological subtype characteristics, a set of diagnostic markers with significant expression difference and stable reliability is screened out, thereby improving the biological representativeness, screening robustness and diagnostic specificity of the markers. BRIEF DESCRIPTION OF DRAWINGS

[0035] Figure 1 A flowchart of a colon cancer diagnostic marker screening method based on transcriptome data provided by the present application is shown in the figure;

[0036] Figure 2 A schematic diagram of a colon cancer diagnostic marker screening system based on transcriptome data provided by the present application is shown in the figure;

[0037] Figure 3 a flowchart for step S2;

[0038] Figure 4 a flowchart for step S3.

[0039] The accompanying drawings are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and are meant to explain the application without limiting the application. DETAILED DESCRIPTION

[0040] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0041] In the description of the present application, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the purpose of facilitating the description of the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.

[0042] Embodiment one, refer to Figure 1 The present application provides a colon cancer diagnosis marker screening method based on transcriptome data, which comprises the following steps:

[0043] Step S1: data preparation;

[0044] Step S2: pathological stratification modeling;

[0045] Step S3: diagnosis marker screening;

[0046] Step S4: screening report generation.

[0047] Embodiment two, refer to Figure 1 This embodiment is based on the above-mentioned embodiment, in step S1, the data preparation provides the data basis for pathological stratification modeling and diagnosis marker screening, specifically, first data acquisition, then constructing sample gene expression matrix and sample pathological feature matrix;

[0048] The data acquisition specifically acquires colon cancer and normal tissue sample data from a public database, wherein each sample includes transcriptome data and pathological image data at the same time;

[0049] The sample gene expression matrix is specifically constructed according to transcriptome data, and the expression value of each gene is logarithmically converted and Z-score standardized to obtain the sample gene expression matrix.

[0050] The sample pathological feature matrix is specifically obtained by extracting features of pathological image data by using a ResNet50 model and performing L2 normalization on the extraction result.

[0051] Embodiment Three, referring to Figure 1 and Figure 3 This embodiment is based on the above-mentioned embodiment, and in step S2, the pathological stratification modeling is used to realize adaptive identification and feature extraction of pathological subtypes, specifically, pathological stratification modeling based on dynamic discrimination and double-layer feature fusion is used according to the sample gene expression matrix and the sample pathological feature matrix to obtain pathological subtype information, including the following steps:

[0052] Step S21: feature embedding modeling is used to represent the multi-dimensional relationship of samples in the gene expression layer and the pathological feature layer, specifically, the gene expression similarity matrix and the pathological feature similarity matrix are calculated by using a similarity function, and adaptive weighted fusion is performed based on an information entropy balance strategy to generate a sample fusion similarity matrix.

[0053] Preferably, the similarity function is a Gaussian kernel function.

[0054] The gene expression similarity matrix is used to reflect the similarity degree of different samples in the gene expression mode, and the calculation formula is:

[0055] ;

[0056] In the formula, S g,ij is the gene expression similarity of sample i and sample j, exp(·) is a natural exponential function, ||·||2 is an L2 norm symbol, g i is the gene expression vector of sample i, the gene expression vector is derived from the sample gene expression matrix, g j is the gene expression vector of sample j, i is the first index of the sample, j is the second index of the sample, i≠j, is the gene expression Gaussian kernel bandwidth, specifically the average Euclidean distance of the gene expression vectors between samples;

[0057] The pathological feature similarity matrix is used to reflect the similarity degree of different samples in the pathological feature mode, and the calculation formula is:

[0058] ;

[0059] In the formula, S c,ijis the pathological feature similarity of sample i and sample j, c i is the pathological feature vector of sample i, which is derived from a sample pathological feature matrix, c j is the pathological feature vector of sample j, is the pathological feature Gaussian kernel bandwidth, specifically the average Euclidean distance between the pathological feature vectors of samples;

[0060] The sample fusion similarity matrix is used to comprehensively reflect the overall similarity degree of different samples in the double-layer space of gene expression features and pathological features, and the calculation formula is:

[0061] ;

[0062] In the formula, S ij is the fusion similarity of sample i and sample j, is the fusion weight;

[0063] The fusion weight is determined by an information entropy balancing strategy, and the calculation formula is:

[0064] ;

[0065] In the formula, Tr(·) is the matrix trace operation, S g is the gene expression similarity matrix, S c is the pathological feature similarity matrix;

[0066] Step S22: adaptive clustering discrimination, used for realizing pathological subtype division based on the sample fusion similarity matrix, specifically, a graph Laplacian matrix is constructed, and feature decomposition is performed to obtain a spectral feature matrix, then a dynamic cluster discrimination function is set, the optimal clustering number is calculated by maximizing the dynamic cluster discrimination function, the samples are clustered in the spectral feature matrix by using a k-means clustering algorithm, and a sample subtype label is obtained;

[0067] The calculation formula of the graph Laplacian matrix is:

[0068] ;

[0069] In the formula, L is the graph Laplacian matrix, D is a degree matrix, the diagonal elements of the degree matrix are the total similarities of each sample, S is the sample fusion similarity matrix, D ii is the total similarity of sample i and other samples;

[0070] The calculation formula of the optimal clustering number is:

[0071] ;

[0072] ;

[0073] In the formula, It is the optimal cluster number. Is Select within the range Take the maximum value of K, where K is the number of clusters. max It is the maximum number of clusters, f dyn (·) is the dynamic cluster discrimination function, where k is the first cluster index. It is the second cluster index, where both the first and second cluster indices are used to represent pathological subtypes, Cen. k It is the center of the k-th cluster. It is the first The center of each cluster, |C k | is the number of samples in the k-th cluster, used to represent the number of samples in the k-th subtype, C k It is the sample set of the k-th cluster, used to represent the sample set belonging to the k-th subtype, x i It is the vector of sample i in the spectral feature matrix. It is the balance coefficient, d k It is the variance of the dispersion within the k-th cluster;

[0074] Step S23: Subtype feature integration, used to generate representative expression features of each pathological subtype, thereby suppressing the influence of noisy samples while maintaining the representativeness of the subtype. Specifically, by weighted integration of samples of the same subtype, a subtype expression feature matrix is ​​obtained.

[0075] The subtype expression feature matrix, used to characterize the representative gene expression patterns of the pathological subtype, is calculated using the following formula:

[0076] ;

[0077] ;

[0078] In the formula, q k W is the feature vector representing the k-th subtype. K It is the normalization factor, w i These are the sample weighting coefficients;

[0079] Preferably, the sample weighting coefficient is set based on the Euclidean distance between the sample and the cluster center and the sequencing quality, and the calculation formula is as follows:

[0080] ;

[0081] In the formula, dist(·) is the Euclidean distance calculation function, and q i The sequencing quality is the sequencing quality of the i-th sample, where the sequencing quality is set as the sequencing depth.

[0082] Step S24: pathological subtype information generation, specifically, generating pathological subtype information including sample subtype labels and subtype expression feature matrix by performing steps S21 to S23.

[0083] By performing the above operations, in view of the technical problems that in the process of colon cancer diagnosis marker screening, there is high heterogeneity in gene expression and pathological characteristics of samples, traditional single feature or fixed grouping method is difficult to accurately divide pathological subtypes, and noise samples may affect the stability and representativeness of the screening results of the marker, the present scheme creatively adopts pathological stratification modeling based on dynamic discrimination and double-layer feature fusion, adaptively identifies the pathological subtypes of colon cancer samples on the basis of comprehensively considering gene expression and pathological morphological information, obtains stable and representative subtype features, and thus improves the accuracy and biological representativeness of the diagnosis marker screening.

[0084] Embodiment four, refer to Figure 1 and Figure 4 This embodiment is based on the above-mentioned embodiments. In step S3, the diagnosis marker screening is specifically diagnosis marker screening based on pathological subtype difference analysis according to the pathological subtype information, and marker screening information is obtained, including the following steps:

[0085] Step S31: subtype difference statistical analysis, used for quantitatively evaluating the expression difference of each gene between different pathological subtypes, specifically, by taking the sample gene expression matrix as the dependent variable and the sample subtype label as the independent variable, a gene expression difference linear model is established, and an expression difference coefficient is fitted. Based on the expression difference coefficient and the model residual, the expression difference significance of each gene between subtypes is calculated by F test, and a difference significance matrix is generated.

[0086] The calculation formula of the gene expression difference linear model is:

[0087] ;

[0088] In the formula, g iu is the expression amount of sample i at gene u, u is the first index of gene, is the average expression amount of gene u, is the expression difference coefficient of gene u in the kth subtype, I(Y i =k) is an indicator function, which takes the value 1 when Y i =k is satisfied, and otherwise takes the value 0, wherein Y i =k indicates that sample i belongs to the kth subtype, is a residual term.

[0089] Step S32: Adaptive differential screening, used for preliminary screening of candidate differentially expressed genes. Specifically, it calculates an adaptive differential score for each gene based on the average expression difference among different subtypes and the expression standard deviation within each subtype, and then selects a candidate differentially expressed gene set based on an adaptive threshold. The calculation formula is as follows:

[0090] ;

[0091] In the formula, G init It is a set of candidate differentially expressed genes, g u It's gene u, ADS u It is an adaptive differential score for gene u. It is an adaptive threshold;

[0092] The formula for calculating the adaptive difference score is as follows:

[0093] ;

[0094] In the formula, It is the average expression characteristic of the kth subtype of gene u. It is the first of gene u Subtype expression characteristic mean, s u,k It is the standard deviation of the expression characteristics of the kth subtype of gene u. It is the first of gene u Subtype expression characteristic standard deviation These are the weighting coefficients between subtypes;

[0095] The formula for calculating the weighting coefficient between subtypes is as follows:

[0096] ;

[0097] In the formula, It is the first Number of samples in the subtype;

[0098] The formula for calculating the adaptive threshold is:

[0099] ;

[0100] In the formula, median(ADS) is the median of the adaptive differential scores for all genes. The adjustment coefficient is used to control the balance between the sensitivity and robustness of the screening. The default value is 1.5. MAD (ADS) is the median absolute deviation of the adaptive differential score for all genes.

[0101] Step S33: stability and redundancy evaluation, for screening a stable and representative gene set, specifically, removing unstable genes by Bootstrap resampling and gene expression stability score, and removing redundant genes by calculating the Spearman rank correlation coefficient, to obtain a candidate stable gene set;

[0102] The Bootstrap resampling, specifically, resampling the candidate differential gene set T times, each time randomly sampling 80% of the sample subset, and recalculating the adaptive differential score of each gene in each sampling;

[0103] The gene expression stability score, the calculation formula is:

[0104] ;

[0105] In the formula, Stab u is the expression stability score of gene u, ADS u (t) is the adaptive differential score of gene u in the t-th resampling, t is the resampling number index, Var{·} is the variance function, and mean{·} is the mean function;

[0106] The calculation formula of the candidate stable gene set is:

[0107] ;

[0108] In the formula, is the second index of the gene, is the gene , is the Spearman rank correlation coefficient of gene u and gene , is a redundancy threshold, and the default value is 0.9, is the adaptive differential score of gene ;

[0109] Step S34: comprehensive score ranking, specifically, by weighting and fusing the differential significance and stability scores, obtaining the comprehensive diagnostic potential score of each gene, and ranking according to the comprehensive diagnostic potential score, selecting the top A genes, and generating a colon cancer diagnostic marker set;

[0110] The calculation formula of the comprehensive diagnostic potential score is:

[0111] ;

[0112] In the formula, Score u is the comprehensive diagnostic potential score of gene u, is a weighting coefficient, and the default value is 0.7, p uis the differential significance of gene u;

[0113] Step S35: screening information generation, specifically, by executing steps S31 to S34, generating marker screening information, including a colon cancer diagnostic marker set, a comprehensive diagnostic potential score, a differential significance matrix, and a gene expression stability score.

[0114] By performing the above operations, in order to solve the technical problems that in the process of colon cancer diagnostic marker screening, there are high heterogeneity of intestinal cancer samples, obvious gene expression difference of different pathological subtypes, and traditional whole sample difference analysis is easily disturbed by noise or internal subtype difference, resulting in lack of stability and specificity of the screened markers, the present application creatively adopts diagnostic marker screening based on pathological subtype difference analysis, and under the premise of considering the pathological subtype characteristics, the diagnostic marker set with significant expression difference and stability and reliability is screened, so that the biological representativeness, screening robustness and diagnostic specificity of the markers are improved.

[0115] Embodiment five, refer to Figure 1 This embodiment is based on the above-mentioned embodiments, in step S4, the screening report generation, specifically, according to the marker screening information, generating a colon cancer diagnostic marker screening report, tabulating the screening results, and providing the comprehensive diagnostic potential score, differential significance, stability score and corresponding pathological subtype information of each gene.

[0116] Embodiment six, refer to Figure 2 This embodiment is based on the above-mentioned embodiments, and the colon cancer diagnostic marker screening system based on transcriptome data provided by the present application comprises a data preparation module, a pathological stratification module, a marker screening module and a report generation module;

[0117] The data preparation module is used for data preparation, and through data preparation, a sample gene expression matrix and a sample pathological feature matrix are obtained, and the sample gene expression matrix is sent to the pathological stratification module and the marker screening module, and the sample pathological feature matrix is sent to the pathological stratification module;

[0118] The pathological stratification module is used for pathological stratification modeling, and through pathological stratification modeling, pathological subtype information is obtained, and the pathological subtype information is sent to the marker screening module and the report generation module;

[0119] The marker screening module is used for diagnostic marker screening, and through diagnostic marker screening, marker screening information is obtained, and the marker screening information is sent to the report generation module;

[0120] The report generation module is used for screening report generation, and a colon cancer diagnostic marker screening report is obtained.

[0121] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting; it is not intended to exclude myriad other embodiments of the present application that other inventors can develop based on the same general inventive concepts embodied by the described embodiments. That is, although the present application is described in terms of particular embodiments and implementations, it is to be understood that the terminology used is for the purpose of descriptive clarity and that it is intended to be limited only by the claims.

[0122] While the embodiments of the application have been shown and described, it is to be understood that the embodiments described are only by way of example and that modifications, changes, substitutions and variations can be made by those skilled in the art without departing from the spirit and scope of the application.

[0123] The above description of the application and its embodiments is not intended to limit the application, as described by the appended claims, to the embodiments described above. Rather, it is intended to cover all adaptations, modifications and variations of the specific embodiments of the application chosen by the inventors as coming within the scope of the application.

Claims

1. A method for screening of diagnostic markers for colon cancer based on transcriptome data, characterized by: The method comprises the following steps: Step S1: data preparation, for providing data basis for pathological stratification modeling and diagnostic marker screening; Step S2: pathological stratification modeling, for realizing adaptive identification and feature extraction of pathological subtypes, specifically, pathological stratification modeling based on dynamic discrimination and double-layer feature fusion is adopted according to sample gene expression matrix and sample pathological feature matrix, and pathological subtype information is obtained; Step S3: diagnostic marker screening, specifically, diagnostic marker screening based on pathological subtype difference analysis is adopted according to the pathological subtype information, and marker screening information is obtained, including the following steps: step S31, subtype difference statistical analysis, step S32, adaptive difference screening, step S33, stability and redundancy evaluation, and step S34, comprehensive score sorting; Step S4: screening report generation.

2. The method of claim 1, wherein the method is based on the colon cancer diagnostic marker screening method using transcriptome data. In step S2, the pathological stratification modeling comprises the following steps: Step S21: feature embedding modeling, for representing the multi-dimensional relationship of samples in the gene expression layer and the pathological feature layer, specifically, gene expression similarity matrix and pathological feature similarity matrix are calculated through a similarity function, and adaptive weighted fusion is performed based on an information entropy balance strategy to generate a sample fusion similarity matrix; The gene expression similarity matrix is used to reflect the similarity degree of different samples in the gene expression mode; The pathological feature similarity matrix is used to reflect the similarity degree of different samples in the pathological feature mode; The sample fusion similarity matrix is used to comprehensively reflect the overall similarity degree of different samples in the double-layer space of gene expression features and pathological features; Step S22: adaptive clustering discrimination, for realizing pathological subtype division based on the sample fusion similarity matrix, specifically, a graph Laplacian matrix is constructed, and feature decomposition is performed to obtain a spectral feature matrix, then a dynamic cluster discrimination function is set, the optimal clustering number is calculated by maximizing the dynamic cluster discrimination function, and k-means clustering algorithm is used to cluster the samples in the spectral feature matrix to obtain sample subtype labels; Step S23: subtype feature integration, for generating representative expression features of each pathological subtype to suppress the influence of noise samples while maintaining the representativeness of the subtypes, specifically, the subtype expression feature matrix is obtained by weighting and integrating samples of the same subtype; The subtype expression feature matrix is used to comprehensively reflect the representative gene expression features of each pathological subtype; Step S24: pathological subtype information generation, specifically, the pathological subtype information is generated by executing steps S21 to S23, including sample subtype labels and subtype expression feature matrix.

3. The method of claim 2, wherein the method is based on the colon cancer diagnostic marker screening method using transcriptome data. In step S31, the subtype difference statistical analysis is used to quantitatively evaluate the expression difference of each gene between different pathological subtypes, specifically, a gene expression difference linear model is established by taking the sample gene expression matrix as the dependent variable and the sample subtype label as the independent variable, the expression difference coefficient is fitted, and the expression difference significance of each gene between subtypes is calculated by F test based on the expression difference coefficient and the model residual, and a difference significance matrix is generated; In step S32, the adaptive difference screening is used for preliminary screening of the candidate differential gene set. Specifically, an adaptive difference score of each gene is calculated based on the average expression difference of each gene between different subtypes and the standard deviation of expression within the subtype, and a candidate differential gene set is screened according to an adaptive threshold. In step S33, the stability and redundancy evaluation is used for screening a stable and representative gene set. Specifically, unstable genes are removed through Bootstrap resampling and gene expression stability score, and redundant genes are removed by calculating the Spearman rank correlation coefficient, to obtain a candidate stable gene set. In step S34, the comprehensive score sorting is specifically performed by weighting and fusing the difference significance and stability score to obtain a comprehensive diagnostic potential score of each gene, and the top A genes are selected according to the comprehensive diagnostic potential score to generate a colon cancer diagnostic marker set. Step S35: screening information generation, specifically, by executing steps S31 to S34, a marker screening information is generated, including a colon cancer diagnostic marker set, a comprehensive diagnostic potential score, a difference significance matrix and a gene expression stability score.

4. The method of claim 3, wherein the method is based on the colon cancer diagnostic marker screening method using transcriptome data. In step S4, the screening report generation is specifically performed according to the marker screening information to generate a colon cancer diagnostic marker screening report, and the screening results are displayed in a table, and the comprehensive diagnostic potential score, the difference significance, the stability score and the corresponding pathological subtype information of each gene are provided.

5. The method of claim 4, wherein the method is based on the transcriptome data. In step S1, the data preparation is used to provide a data basis for pathological stratification modeling and diagnostic marker screening. Specifically, data acquisition is performed first, and then a sample gene expression matrix and a sample pathological feature matrix are constructed.

6. A colon cancer diagnostic marker screening system based on transcriptome data for implementing the colon cancer diagnostic marker screening method based on transcriptome data according to any one of claims 1 to 5, characterized in that: The data preparation module, the pathological stratification module, the marker screening module and the report generation module are included. 7.The colon cancer diagnosis marker screening system based on transcriptome data according to claim 6, characterized in that: The data preparation module is used for data preparation. Through data preparation, a sample gene expression matrix and a sample pathological feature matrix are obtained, and the sample gene expression matrix is sent to the pathological stratification module and the marker screening module, and the sample pathological feature matrix is sent to the pathological stratification module. The pathological stratification module is used for pathological stratification modeling. Through pathological stratification modeling, pathological subtype information is obtained, and the pathological subtype information is sent to the marker screening module and the report generation module. The marker screening module is used for diagnostic marker screening. Through diagnostic marker screening, marker screening information is obtained, and the marker screening information is sent to the report generation module. The report generation module is used for screening report generation to obtain a colon cancer diagnostic marker screening report.