Breast cancer lung metastasis related biomarker screening system

By using differential genetic factor mining, functional association, and validation screening modules, combined with signal pathway analysis and interference elimination, highly specific biomarkers for breast cancer lung metastasis are generated. This solves the problems of unstable screening results and poor applicability in existing technologies, and enables more accurate biomarker screening and clinical application.

CN121601045AInactive Publication Date: 2026-03-03THE FIRST AFFILIATED HOSPITAL OF MEDICAL COLLEGE OF XIAN JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511834168.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-03-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies rely on transcriptomic analysis of single-source samples for screening biomarkers related to lung metastasis in breast cancer. This makes the screening results susceptible to sample heterogeneity and background noise, resulting in poor reproducibility, lack of stability, and limited clinical applicability.

Method used

The differential genetic factor mining module extracts the transcriptional profiles of candidate genetic factors, the functional association module analyzes the weights of signaling pathways, the verification and screening module evaluates transcriptional consistency, and the interference exclusion module removes non-specific factors, generating a list of highly specific biomarkers.

Benefits of technology

It significantly improves the accuracy and clinical applicability of screening biomarkers for lung metastasis in breast cancer, and provides a high-quality set of biomarkers with greater indicative significance of metastasis mechanisms and prognostic value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121601045A_ABST
    Figure CN121601045A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of tumor marker detection, in particular to a breast cancer lung metastasis related biomarker screening system which comprises a differential genetic factor mining module, a function association module, a verification screening module, an interference elimination module and an optimization integration module. According to the method, the candidate genetic factors highly associated with lung metastasis are extracted, and the expression fluctuation characteristics are combined to carry out difference analysis, so that the preliminary screening accuracy is improved, pathological information and pathway annotation are integrated to construct an action network, pathway weights are analyzed, and functional expressions of the candidate factors are quantitatively evaluated; an independent sample is called to verify transcription consistency and clinical relevance, interference factors with poor repeatability are eliminated, non-tumor tissue high-expression interference items are rejected in combination with background transcription characteristics, the specificity and adaptability of screening results are improved, and finally a marker set with metastasis mechanism indicating significance and prognosis prediction value is formed through comprehensive evaluation. And clinical transformation support is provided for identification of lung metastasis of breast cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tumor marker detection technology, and in particular to a screening system for biomarkers related to lung metastasis in breast cancer. Background Technology

[0002] The field of tumor marker detection technology involves the detection and analysis of specific biomolecules related to tumor development and progression in body fluids or tissue samples, thereby assisting in the early diagnosis, classification, efficacy evaluation, and prognosis of tumors. This field mainly includes key aspects such as tumor-related genetic factor expression analysis, proteomics detection, circulating tumor cell identification, microRNA detection, exosome analysis, and metabolite monitoring. Tumor marker detection relies on molecular biology, high-throughput sequencing, immunoassay, and liquid biopsy, and has important value in clinical research and precision medicine, especially in the screening of markers for specific cancer types, which has high specificity and research depth.

[0003] Traditional breast cancer lung metastasis-related biomarker screening systems refer to the differential expression analysis of transcriptome data from existing patient samples to screen for differentially expressed genetic factors highly associated with lung metastasis, particularly the risk and mechanisms of distant metastasis, especially lung metastasis, in breast cancer research. This is then combined with clinical data related to breast cancer metastasis from public databases, and survival analysis and correlation validation are used to further identify potential biomarkers. Traditional methods employ data analysis of single-source samples, relying on existing RNA sequencing data and bioinformatics techniques to identify breast cancer lung metastasis-related factors. The core components include sample data collection, differential genetic factor analysis, survival validation, literature comparison, and database cross-validation.

[0004] Current technologies for screening lung metastasis biomarkers in breast cancer rely on transcriptomic analysis from single-source samples, failing to provide in-depth analysis of the expression background and functional pathways of genetic factors. This makes the screening results susceptible to sample heterogeneity and background noise, and prone to poor reproducibility under diverse sample conditions. Furthermore, the lack of an independent validation mechanism limits the stability assessment of biomarkers, and the absence of background expression elimination strategies also leads to the inclusion of some non-specific factors in the screening results, thereby reducing their clinical interpretability and applicability. For example, factors highly expressed in non-tumor tissues may be mistakenly identified as metastasis-related factors, thus misleading mechanistic studies and prognostic assessments. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a screening system for biomarkers related to lung metastasis in breast cancer.

[0006] To achieve the above objectives, the present invention employs the following technical solution: a breast cancer lung metastasis-related biomarker screening system comprising:

[0007] The differential genetic factor mining module acquires transcriptome data from breast cancer patient samples, extracts the transcription profiles of candidate genetic factors associated with lung metastasis, analyzes the fluctuation range and distribution characteristics of genetic factor transcription values, identifies differentially transcribed genetic factor sets among differentially transcribed samples, and generates a preliminary list of candidate biomarkers.

[0008] Based on the preliminary candidate biomarker list, the functional association module combines breast cancer pathology information and molecular pathway annotation data from public databases to analyze the role weights of breast cancer candidate genetic factors in key signaling pathways and generate functional weight ranking results.

[0009] The validation and screening module, based on the functional weight ranking results, calls independent validation queue sample data to evaluate the transcriptional consistency and clinical relevance of breast cancer candidate genetic factors in the validation queue, eliminates low-repetition genetic factors, and generates biomarker screening results.

[0010] The interference elimination module calls the biomarker screening results, extracts the background transcription amount of breast cancer candidate genetic factors in non-tumor tissues, combines tissue-specific transcription patterns and noise interference characteristics, eliminates non-specific transcription genetic factors, and generates a list of highly specific biomarkers.

[0011] As a further aspect of the present invention, the preliminary candidate biomarker list includes the amplitude of genetic factor transcriptional fluctuation, significant difference score, and transcriptional distribution characteristic value; the functional weight ranking result includes signaling pathway contribution, genetic factor interaction strength, and functional association score; the biomarker screening result includes validation consistency ratio, clinical relevance index, and transcriptional stability score; and the high-specificity biomarker list includes tissue-specific transcriptional score, background interference coefficient, and noise removal ratio.

[0012] As a further aspect of the present invention, the differential genetic factor mining module includes:

[0013] The transcription profile extraction submodule acquires transcriptome data from breast cancer patient samples, extracts the transcription value of each genetic factor and its distribution characteristics among differential samples, analyzes the standard dispersion and fluctuation trend of genetic factor transcription values, determines whether the difference threshold is met, and obtains differential genetic factor trigger signals.

[0014] Based on the differential genetic factor triggering signal, the fluctuation analysis submodule extracts the fluctuation range and distribution characteristics of the transcription values ​​of candidate genetic factors in breast cancer, combines the changing trends of genetic factor transcription values ​​among samples, analyzes the consistency of genetic factor transcription fluctuations, and obtains genetic factor transcription fluctuation parameters.

[0015] The pathway mapping submodule performs functional pathway mapping on candidate genetic factors for breast cancer based on the transcriptional fluctuation parameters of the genetic factors, and generates a preliminary list of candidate biomarkers by combining the distribution location and role weight of the genetic factors in key signaling pathways.

[0016] As a further aspect of the present invention, the functional association module includes:

[0017] Based on the preliminary candidate biomarker list, the pathway annotation submodule extracts signaling pathway annotation information of breast cancer candidate genetic factors in public databases, and combines the upstream and downstream relationships of genetic factors in key pathways to identify the role weight of genetic factors in pathways and obtain a set of pathway contribution values.

[0018] The interaction parsing submodule calls the pathway contribution value set, detects the core nodes and connection density in the genetic factor interaction network based on the interaction relationship and functional association strength between genetic factors, calculates the genetic factor interaction strength value, and obtains the functional association parameter value by combining the interaction results and functional weights.

[0019] The weight ranking submodule calls the function-related parameter values, prioritizes the genetic factors based on their contribution and interaction strength in the pathway, determines whether they are located in the core functional region, maps the functional weight parameters, and generates the functional weight ranking result.

[0020] As a further aspect of the present invention, the verification and screening module includes:

[0021] The consistency assessment submodule collects independent validation cohort sample data based on the functional weight ranking results, extracts the transcription value distribution characteristics of breast cancer candidate genetic factors in the validation cohort, identifies transcriptional consistency trends, and generates validation consistency parameters.

[0022] The correlation analysis submodule calls the validation consistency parameters, combines the correlation between the transcription values ​​of candidate genetic factors in breast cancer and clinical indicators, detects the association strength between genetic factor transcription and clinical phenotype, analyzes the trend of clinical correlation changes, and obtains the clinical correlation evaluation value.

[0023] The screening and optimization submodule calls the clinical relevance evaluation value, combines the transcriptional stability and repeatability assessment results, eliminates low-reproducibility genetic factors, analyzes the screening and optimization extent, corrects the list of candidate genetic factors for breast cancer, and generates biomarker screening results.

[0024] As a further aspect of the present invention, the interference elimination module includes:

[0025] The background transcription extraction submodule calls the biomarker screening results to extract the background transcription amount of breast cancer candidate genetic factors in non-tumor tissues, screens genetic factors that meet the background transcription threshold, and obtains the background transcription change rate.

[0026] The specific detection submodule calls the background transcriptional change rate to detect the transcriptional differences of breast cancer candidate genetic factors in tumor tissues and non-tumor tissues, and calculates the proportion of non-specific transcribed genetic factors to obtain the proportion of non-specific transcription.

[0027] The noise removal submodule collects the transcriptional noise distribution of breast cancer candidate genetic factors in differential tissues based on the non-specific transcription ratio, performs a comparison with tissue-specific transcription patterns and error amplitude, identifies the noise removal coefficient, screens out non-specific transcription genetic factors, and generates a list of high-specificity biomarkers.

[0028] As a further aspect of the present invention, the system also includes an optimization and integration module:

[0029] Based on the list of highly specific biomarkers, the optimization and integration module combines the updated functional weight ranking results and the validation screening status to evaluate the role weight and clinical application potential of breast cancer candidate genetic factors in the mechanism of breast cancer lung metastasis and generate a biomarker evaluation index.

[0030] The biomarker evaluation index includes functional weight contribution, clinical application potential score, and mechanism association strength.

[0031] As a further aspect of the present invention, the optimization and integration module includes:

[0032] The weight evaluation submodule extracts the updated functional weight sequences of breast cancer candidate genetic factors and the initial weight data after background transcription knockout based on the list of highly specific biomarkers. It analyzes the weight change trends of breast cancer candidate genetic factors in the differential functional dimension and performs normalization processing in combination with timestamps to obtain the genetic factor weight change trends.

[0033] The mechanism association comparison submodule analyzes the weight change rate and mechanism association strength at corresponding time points based on the trend of the genetic factor weight change and the weight distribution sequence of the role of candidate genetic factors in the lung metastasis mechanism of breast cancer, and judges the stability of the mechanism association interval to obtain mechanism association characteristics.

[0034] The evaluation strategy generation submodule calls the mechanism-associated features and the weight change rate of the candidate breast cancer genetic factors at the preceding time step, sets the weight parameters according to the mechanism association stability and the magnitude of weight change, and performs weighted processing on each candidate breast cancer genetic factor to generate a biomarker evaluation index.

[0035] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0036] In this invention, by extracting candidate genetic factors highly associated with lung metastasis from the transcriptome and constructing a multi-dimensional differential analysis pathway based on expression fluctuation characteristics during the screening of biomarkers related to lung metastasis in breast cancer, the accuracy of preliminary screening can be significantly improved. By integrating pathological database information and molecular pathway annotations to establish an action network and analyze pathway action weights, quantitative functional assessment of candidate factors can be achieved. Then, by using independent samples to verify their transcriptional consistency and clinical significance, interfering factors with poor reproducibility can be effectively eliminated. By combining background transcriptional characteristics, high-expression interfering items in non-tumor tissues can be further eliminated, thereby improving the specificity and clinical applicability of the final screening results. Finally, based on comprehensive evaluation, a high-quality biomarker set with indications of metastasis mechanisms and prognostic predictive value is formed, providing more clinically translational support for the accurate identification of lung metastasis in breast cancer. Attached Figure Description

[0037] Figure 1 This is a system flowchart of the present invention;

[0038] Figure 2 This is a flowchart of the differential genetic factor mining module in this invention;

[0039] Figure 3 This is a flowchart of the functional association modules in this invention;

[0040] Figure 4 This is a flowchart of the verification and screening module in this invention;

[0041] Figure 5 This is a flowchart of the interference elimination module in this invention;

[0042] Figure 6 This is a flowchart of the optimized integration module in this invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0044] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0045] Please see Figure 1 A breast cancer lung metastasis-related biomarker screening system includes:

[0046] The differential genetic factor mining module acquires transcriptome data from breast cancer patient samples, extracts the transcription profiles of candidate genetic factors associated with lung metastasis, analyzes the fluctuation range and distribution characteristics of genetic factor transcription values, identifies differentially transcribed genetic factor sets among differentially transcribed samples, and generates a preliminary list of candidate biomarkers.

[0047] The functional association module, based on a preliminary candidate biomarker list and combined with breast cancer pathology information and molecular pathway annotation data from public databases, constructs a genetic factor functional association network, analyzes the role weights of breast cancer candidate genetic factors in key signaling pathways, and generates functional weight ranking results.

[0048] The validation and screening module sorts the results by functional weight, calls the sample data of the independent validation cohort, evaluates the transcriptional consistency and clinical relevance of breast cancer candidate genetic factors in the validation cohort, removes low-repetition genetic factors, and generates biomarker screening results.

[0049] The interference elimination module calls the biomarker screening results, extracts the background transcription amount of breast cancer candidate genetic factors in non-tumor tissues, combines tissue-specific transcription patterns and noise interference characteristics, eliminates non-specific transcription genetic factors, and generates a list of highly specific biomarkers.

[0050] The optimization and integration module, based on a list of highly specific biomarkers and combined with updated functional weight ranking results and validation screening status, assesses the role weight and clinical application potential of candidate genetic factors in the lung metastasis mechanism of breast cancer, and generates a biomarker evaluation index.

[0051] The preliminary candidate biomarker list includes the amplitude of genetic factor transcriptional fluctuations, significance difference scores, and transcriptional distribution characteristics. The functional weight ranking results include signaling pathway contribution, genetic factor interaction strength, and functional association score. The biomarker screening results include validation consistency ratio, clinical relevance index, and transcriptional stability score. The high-specificity biomarker list includes tissue-specific transcription score, background interference coefficient, and noise removal ratio. The biomarker evaluation index includes functional weight contribution, clinical application potential score, and mechanism association strength.

[0052] Please see Figure 2 The differential genetic factor mining module includes:

[0053] The transcription profile extraction submodule acquires transcriptome data from breast cancer patient samples, extracts the transcription value of each genetic factor and its distribution characteristics among differential samples, analyzes the standard dispersion and fluctuation trend of genetic factor transcription values, determines whether the difference threshold is met, and obtains differential genetic factor trigger signals.

[0054] Transcriptome data were obtained from breast cancer patient samples. RNA sequencing was performed on 50 breast cancer tumor tissue samples and 50 paired adjacent normal tissue samples. The raw read length data were aligned to the human reference genome and converted to total transcripts per million reads (TPM) to obtain the standardized transcription value of each genetic factor (e.g., gene A) in each sample. For example, the transcription value of gene A in tumor sample T1 was 85 TPM, and the transcription value in normal sample N1 was 12 TPM. The transcription value dataset of gene A in all tumor samples was collected, and its distribution characteristics among differentially expressed samples (tumor and normal) were calculated, including mean, standard deviation, and quartile range. For example, the mean transcription value of gene A in the tumor sample group was 78 TPM, and the standard deviation was [missing data]. The mean transcription value in the normal sample group was 15 TPM, with a standard deviation of 3 TPM. Standard dispersion and fluctuation trend assessment were performed on the transcription value of gene A. The fold change between tumor and normal samples was calculated (e.g., 78 / 15 ≈ 5.2) and the Student's t-test p-value (e.g., 1.2e-6). A fold change threshold of 2.0 and a p-value threshold of 0.05 were determined based on extensive transcriptome research experience. For gene A, a fold change of 5.2 was greater than 2.0, and a p-value of 1.2e-6 was less than 0.05, thus meeting the fold change threshold conditions. By screening all genetic factors that met the above conditions, differential genetic factor trigger signals were obtained. For example, gene A was identified as a differential genetic factor trigger signal.

[0055] The fluctuation analysis submodule extracts the fluctuation range and distribution characteristics of the transcription values ​​of candidate genetic factors in breast cancer based on differential genetic factor triggering signals. Combined with the changing trends of genetic factor transcription values ​​among samples, it analyzes the consistency of genetic factor transcription fluctuations and obtains genetic factor transcription fluctuation parameters.

[0056] Based on the differentially triggered genetic factor signal of gene A, the transcription values ​​of gene A were collected from all 50 breast cancer tumor samples. For example, the maximum value of the transcription value set was 92 TPM, the minimum value was 68 TPM, and its fluctuation range was determined to be [68 TPM, 92 TPM]. The mean of its distribution characteristics was 78 TPM, and the standard deviation was 15 TPM. Then, combined with the changing trend of genetic factor transcription values ​​among samples, the average relative change rate of gene A among different tumor samples was calculated. For example, by comparing the changes in transcription values ​​of adjacent samples, the average relative change rate was found to be 0.05. Subsequently, the consistency of genetic factor transcription fluctuations was analyzed by calculating the gene... The Kendall concordance coefficient of gene A across all 50 tumor samples was used to assess the synergy of its transcriptional pattern across different tumor samples. For example, a result of 0.75 indicates that the transcriptional fluctuation pattern of gene A has high consistency. A consistency higher than 0.70 is defined as high consistency, and a consistency between 0.50 and 0.70 is defined as moderate consistency. The threshold of 0.70 was set based on empirical data on biomarker stability. The transcriptional fluctuation parameters of the genetic factor were obtained, including the fluctuation range of gene A [68 TPM, 92 TPM], the fluctuation trend (mean relative change rate 0.05), and the transcriptional fluctuation consistency score of 0.75.

[0057] The pathway mapping submodule performs functional pathway mapping on candidate genetic factors for breast cancer based on the transcriptional fluctuation parameters of genetic factors, and generates a preliminary list of candidate biomarkers by combining the distribution location and role weight of genetic factors in key signaling pathways.

[0058] Based on the transcriptional fluctuation parameters of gene A (fluctuation range [68 TPM, 92 TPM], fluctuation trend 0.05, consistency score 0.75), functional pathway mapping was performed on candidate genetic factors for breast cancer. First, the EntrezGeneID of gene A was submitted to the KEGG public biological pathway database for querying. For example, the query results showed that gene A is involved in the "breast cancer signaling pathway" (hsa05224). Simultaneously, combining the distribution location and role weight of genetic factors in key signaling pathways, the upstream and downstream regulatory relationships of gene A in the "breast cancer signaling pathway" and its interaction with other pathways were analyzed. The interaction of members, for example, gene A was identified as a downstream effector of ERBB2, and its role weight was set to 0.8 (range 0 to 1) based on its node degree centrality in the pathway (e.g., number of connections of 8) and the importance reported in the literature. The weight evaluation was based on a review of PubMed literature and GeneOntology annotation information. This weight value of 0.8 is higher than the preset high weight judgment threshold of 0.75, and a preliminary candidate biomarker list is generated. For example, based on the high role weight (0.8) of gene A and its important role in key pathways in breast cancer, gene A was included in the preliminary candidate biomarker list.

[0059] Please see Figure 3 The functional modules include:

[0060] The pathway annotation submodule extracts signaling pathway annotation information of breast cancer candidate genetic factors from public databases based on the preliminary candidate biomarker list. Combining the upstream and downstream relationships of genetic factors in key pathways, it identifies the role weight of genetic factors in pathways and obtains a set of pathway contribution values.

[0061] Based on the transcriptional fluctuation parameters of gene A (fluctuation range [68 TPM, 92 TPM], fluctuation trend 0.05, consistency score 0.75), functional pathway mapping was performed on candidate genetic factors for breast cancer. First, the EntrezGeneID of gene A was submitted to the KEGG public biological pathway database for querying. For example, the query results showed that gene A is involved in the "breast cancer signaling pathway" (hsa05224). Simultaneously, combining the distribution location and role weight of genetic factors in key signaling pathways, the upstream and downstream regulatory relationships of gene A in the "breast cancer signaling pathway" and its interaction with other pathways were analyzed. The interaction of members, for example, gene A was identified as a downstream effector of ERBB2, and its role weight was set to 0.8 (range 0 to 1) based on its node degree centrality in the pathway (e.g., number of connections of 8) and the importance reported in the literature. The weight evaluation was based on a review of PubMed literature and GeneOntology annotation information. This weight value of 0.8 is higher than the preset high weight judgment threshold of 0.75, and a preliminary candidate biomarker list is generated. For example, based on the high role weight (0.8) of gene A and its important role in key pathways in breast cancer, gene A was included in the preliminary candidate biomarker list.

[0062] The interaction analysis submodule calls the pathway contribution value set and, based on the interaction relationships and functional association strength between genetic factors, detects the core nodes and connection density in the genetic factor interaction network using the following formula:

[0063] ;

[0064] Calculate the interaction strength of genetic factors, and combine the interaction results with functional weights to obtain the functional association parameter values;

[0065] in, This represents the strength of hereditary factor interactions. Representative genetic factors and The connection weight values ​​between them are based on direct interaction relationships. Representative genetic factors and Through mediating factors Connection coefficients that result in indirect interaction Representing mediating factors Node importance scoring in functional pathways Representative genetic factors and The correlation density value between them in the interaction network Representative genetic factors and Functional similarity score between them Represents the average functional similarity score for all pairs of genetic factors. This represents the total number of mediating factors involved in the calculation;

[0066] Access pathway contribution value sets are retrieved, for example, sets containing gene A's contribution of 0.85 and gene B's contribution of 0.70 in the "breast cancer signaling pathway." Based on the interaction relationships and functional association strength between genetic factors, the core nodes and connection density in the genetic factor interaction network are detected. First, direct physical or genetic interaction information between gene A and gene B is extracted from protein interaction databases such as STRING or BioGRID to determine... The value of the connection weight, for example, the experimentally verified direct physical interaction between gene A and gene B. The value was set to 0.75 (range 0 to 1, representing the strength and reliability of the interaction; 0.75 indicates a strong interaction with experimental evidence), and the interaction between gene A and gene B was identified through mediators. Indirect interactions (such as those between genes C and D) are investigated, and the corresponding connection coefficients are obtained. and mediating factors Node importance scoring in functional pathways For example, gene A interacts indirectly with gene B through mediator C. The importance score of mediator C in breast cancer signaling pathways is 0.6 (range 0 to 1, indicating the strength of indirect interaction). The association density value between gene A and gene B in the interaction network is 0.80 (determined based on their topology and known function in the pathway, e.g., high node degree, appearance in multiple key pathways). It is calculated based on the number of their common neighbors and path lengths in a pre-constructed breast cancer-specific interaction network, for example, The value is 0.9 (ranging from 0 to 1, with higher values ​​indicating stronger network connections), genetic factors. and Functional similarity score between It was calculated by comparing shared Gene Ontology (GO) annotation terminology with KEGG pathway enrichment results, for example, The mean functional similarity score for all pairs of genetic factors is 0.70 (ranging from 0 to 1, with higher values ​​indicating greater functional similarity). It is obtained by averaging the functional similarity scores of all identified pairs of genetic factors, for example, The value is 0.65, representing the total number of mediating factors involved in the calculation. For example, if the mediators are 2 (e.g., mediator C and mediator D), use the formula. ; Calculate the interaction strength between genetic factors, for example, calculate the interaction strength between gene A and gene B:

[0067] Table 1: Parameters for Calculating Interaction Strength Values

[0068] Parameter name symbol Example value illustrate Connection weight value 0.75 The connection weight value between genetic factors A and B based on direct interaction, ranging from 0 to 1, with 0.75 indicating a strong direct interaction. Indirect interaction connection coefficient 0.6 The connection coefficient between genetic factors A and B through the mediator C, ranging from 0 to 1, with 0.6 indicating a moderate indirect interaction. Mediation factor importance score 0.80 The importance score of mediator C in the functional pathway is 0-1, with 0.80 indicating that the factor is highly important in the pathway. Correlation density value 0.9 The association density value between genetic factors A and B in the interaction network, ranging from 0 to 1, with 0.9 indicating a high association density. Functional similarity score 0.70 Functional similarity score between genetic factors A and B, ranging from 0 to 1, with 0.70 indicating high functional similarity. Average functional similarity score 0.65 The average functional similarity score of all pairs of genetic factors. Total number of mediating factors 2 The total number of mediating factors involved in the calculation. Mediator D connection coefficient 0.4 The connection coefficient between genetic factors A and B through the indirect interaction of mediator D, and the node importance score of mediator D in the functional pathway. It is 0.65.

[0069] in, This represents the strength of interaction between genetic factors. It is a dimensionless, comprehensive score reflecting the strength of the functional association between two genetic factors in a biological system; a higher value indicates a stronger association. Representative genetic factors and The connection weights based on direct interaction relationships are obtained through experimentally verified direct physical or genetic interaction data, and are set between 0 and 1. For example, 0.75 indicates a strong direct interaction. Representative genetic factors and Through mediating factors The connection coefficient for indirect interaction is obtained by calculating the reliability and strength of the intermediate path, and is set between 0 and 1. For example, 0.6 represents a moderate indirect interaction. Representing mediating factors The importance score of nodes in functional pathways is obtained by analyzing the topological characteristics (such as centrality and connectivity) of mediators in the pathway network and their known biological functions (such as disease-driving genes). The score is set between 0 and 1, for example, 0.80 indicates high importance. This represents the sum of indirect interactions contributed by all mediating factors. Representative genetic factors and The association density value between two factors in the interaction network reflects the number of common neighbors and path lengths between them. It ranges from 0 to 1; for example, 0.9 indicates high association density. Representative genetic factors and A functional similarity score was calculated by comparing shared Gene Ontology (GO) annotation terms and signaling pathway enrichment results, with a range of 0 to 1. For example, 0.70 indicates high functional similarity. This represents the average functional similarity score for all pairs of genetic factors, used to relativize the functional similarity of a specific pair of factors. This represents the total number of mediating factors involved in the calculation, i.e. and The number of intermediate genetic factors that indirectly interact, and the absolute value sign in the formula. This ensures that the interaction strength value is non-negative, and the denominator contains... The term is used to penalize factor pairs whose functional similarity deviates significantly from the average level. This means that the more unusual the functional similarity of a factor pair, the more cautiously its interaction strength needs to be evaluated functionally. The introduction of this formula makes the contribution of association density to interaction strength non-linear, avoiding the overexpansion caused by high density, while emphasizing the fundamental role of density in the reliability of interaction. The advantage of this formula is that it comprehensively quantifies the association strength between genetic factors by taking into account multiple dimensions such as direct interaction, indirect interaction, interaction network density and functional similarity. This makes the evaluation results closer to the real biological interaction and improves the accuracy of biomarker screening.

[0070] Substitute the parameters from Table 1 into the formula to calculate: First, calculate:

[0071] ;

[0072] Then calculate the absolute value term in the denominator: ;

[0073] Substitute all values ​​into the main formula: ;

[0074] Calculation results The results indicate that the interaction strength between gene A and gene B is 1.346, which is higher than the generally set baseline value of 1.0 for interaction strength (1.0 is set as the threshold for interaction strength with significant biological significance). This means that there is a highly significant functional association between gene A and gene B. This value, as part of the functional association parameter value, will be directly used for priority division in the subsequent weight ranking submodule. Combining the interaction results and functional weights, the functional association parameter value is obtained. This parameter value includes the interaction strength value of 1.346 calculated between gene A and gene B, as well as the specific contribution value in the pathway. For example, the contribution of gene A in the "breast cancer signaling pathway" is 0.85, and the contribution of gene B in the "breast cancer signaling pathway" is 0.70.

[0075] The weight ranking submodule calls the function-related parameter values, prioritizes the genetic factors based on their contribution and interaction strength in the pathway, determines whether they are located in the core functional region, maps the functional weight parameters, and generates the functional weight ranking results.

[0076] The function association parameter values ​​are invoked, for example, the function association parameter values ​​containing the interaction strength value of gene A and gene B (1.346) and the pathway contribution of gene A (0.85) are invoked. Prioritization is based on the contribution and interaction strength of the genetic factors in the pathway. First, for gene A, its contribution of 0.85 in the "breast cancer signaling pathway" is weighted and summed with the average interaction strength value with all genetic factors in the interaction network (assumed to be 1.25). A weighting factor of 0.6 is assigned to the pathway contribution, and 0.4 is assigned to the average interaction strength. Therefore, the overall score of gene A is (0.85⋅0.6) + (1.25⋅0.4) = 1.01, with a weighting factor of 0. The values ​​of 6 and 0.4 were determined based on the training and validation of the machine learning model, aiming to balance the influence of the two indicators. Based on this, it was determined whether the comprehensive score of each genetic factor was located in the core functional region. The core functional region was defined as a comprehensive score greater than 1.00. Gene A's comprehensive score of 1.01 is greater than 1.00, so it is located in the core functional region. Then, the functional weight parameter was mapped, and a high functional weight of 1.0 was assigned to gene A, which is located in the core functional region. This weight parameter will serve as the basis for subsequent analysis to generate functional weight ranking results. For example, all candidate genetic factors are arranged in descending order according to their comprehensive scores, with gene A at the top of the list and a functional weight parameter of 1.0.

[0077] Please see Figure 4 The verification and screening module includes:

[0078] The consistency assessment submodule collects sample data from the independent validation cohort based on the functional weight ranking results, extracts the transcriptional value distribution characteristics of breast cancer candidate genetic factors in the validation cohort, identifies transcriptional consistency trends, and generates validation consistency parameters.

[0079] Based on the ranking results of gene A with a functional weight parameter of 1.0, independent validation cohort sample data were collected. Tumor tissue samples were collected from an independent validation cohort containing 100 breast cancer patients. The samples were sourced from the TCGA database and had a different origin from the initial discovery cohort. The transcription value distribution characteristics of breast cancer candidate genetic factors in the validation cohort were extracted. For gene A, its transcription value (TPM) was extracted from these 100 patient samples, and the mean TPM and standard deviation of the transcription value were calculated to identify the transcriptional consistency trend. The transcription value distribution characteristics of gene A in the initial discovery cohort (mean 78) were used to identify the transcriptional consistency trend. The similarity of transcriptional levels is assessed by comparing the TPM (transcriptional per mille, standard deviation 15TPM) with the transcriptional value distribution characteristics in the validation cohort and calculating the Pearson correlation coefficient between the two distributions. For example, the Pearson correlation coefficient of the transcriptional value distribution of gene A in the two cohorts is 0.88, which is higher than the preset consistency threshold of 0.80 (0.80 is the threshold for judging high transcriptional consistency, which is set according to the general requirements of biomarker validation). This indicates that gene A maintains a highly consistent transcriptional pattern in different patient cohorts, generating validation consistency parameters. For example, the validation consistency parameter of gene A is 0.88.

[0080] The correlation analysis submodule calls the consistency parameter verification, combines the correlation between the transcription values ​​of candidate breast cancer genetic factors and clinical indicators, detects the association strength between genetic factor transcription and clinical phenotype, analyzes the trend of changes in clinical correlation, and uses the following formula:

[0081] ;

[0082] Obtain the clinical relevance evaluation value;

[0083] in, Represents the clinical relevance evaluation value. Representative sample Transcription values ​​of candidate genetic factors for breast cancer in China Representative sample Actual observed values ​​of clinical indicators in China This represents the average of the product of genetic factor transcription values ​​and clinical indicators across all samples. This represents the standard deviation of the product of genetic factor transcription values ​​and clinical indicators across all samples. Represents the total number of samples involved in the calculation;

[0084] Using a validation consistency parameter of 0.88 for gene A, and combining the correlation between the transcriptional values ​​of candidate breast cancer genetic factors and clinical indicators, the association strength between genetic factor transcription and clinical phenotype was detected. First, the transcriptional value of gene A for each patient in the validation cohort was obtained; for example, the transcriptional value of gene A for patient Z1. The transcription value of gene A in patient Z2 was 80 TPM. The target was 72 TPM, and actual observed values ​​of corresponding clinical indicators for patients were obtained, such as disease-free survival (PFS, in months) as a clinical indicator, and the PFS of patient Z1. The progression-free survival (PFS) of patient Z2 was 48 months. For a period of 30 months, with a total of m samples (m=100), calculate the average of the product of genetic factor transcription values ​​and clinical indicators. and standard deviation ,in Representative sample Transcription values ​​of candidate genetic factors for breast cancer in samples, for example, gene A. Transcription values ​​in Representative sample Actual observed values ​​of clinical indicators, such as patient... Disease-free survival (PFS) or tumor size, This represents the average of the product of genetic factor transcription values ​​and clinical indicators across all samples, calculated as follows: , This represents the standard deviation of the product of genetic factor transcription values ​​and clinical indicators across all samples, calculated as follows: , The total number of samples involved in the calculation is represented by m=100. The analysis focuses on the changing trends of clinical relevance. Pearson correlation coefficients are calculated between gene A transcription values ​​and multiple clinical indicators such as disease-free survival, tumor size, and lymph node metastasis status. For example, the correlation coefficient between gene A transcription values ​​and disease-free survival is -0.65, and the correlation coefficient with tumor size is 0.70, indicating that high gene A expression is associated with shorter disease-free survival and larger tumor size. The formula is used... To obtain the clinical relevance evaluation value, taking patients Z1 and Z2 as examples, assuming only these two samples are considered and PFS is the clinical indicator:

[0085] Table 2: Examples of Calculation Parameters for Clinical Relevance Evaluation Values

[0086] Sample ID (TPM) (moon) Z1 80 48 3840 Z2 72 30 2160

[0087] For these two samples:

[0088] Calculate the sum of the product terms: ;

[0089] Calculate the average of the product terms: ;

[0090] Calculate the variance of the product term:

[0091] Squared deviation of sample Z1: ;

[0092] Squared deviation of sample Z2: ;

[0093] variance is ;

[0094] Standard deviation: ;

[0095] Total number of samples: Substitute the value into the formula: ;

[0096] in, This represents a clinical relevance assessment value. This value quantifies the linear association strength between genetic factor transcription values ​​and specific clinical indicators. The larger the absolute value, the stronger the association; a positive value indicates a positive correlation, and a negative value indicates a negative correlation. Representative sample Transcription values ​​of candidate genetic factors for breast cancer were determined using methods such as RNA sequencing, and expressed in TPM or FPKM. Representative sample Actual observed values ​​of clinical indicators, such as the patient's disease-free survival (in months), tumor size (in centimeters), or clinical stage (in numerical codes). This represents the sum of the products of genetic factor transcription values ​​and clinical indicators across all samples. This represents the average of the product of genetic factor transcription values ​​and clinical indicators across all samples. These two statistics represent the standard deviation of the product term of genetic factor transcription values ​​and clinical indicators across all samples. They are used to center and standardize the product term. Used to adjust the impact of sample size on evaluation values, so that research results with different sample sizes are comparable;

[0097] The advantage of this formula lies in its ability to provide a unified and comparable clinical relevance evaluation value by standardizing the product of transcription values ​​and clinical indicators. This allows for the quantification of the association strength between different genetic factors under different clinical indicators, which helps to accurately identify biomarkers closely related to disease progression and prognosis. For example, for all 100 samples, the calculated clinical relevance evaluation value for gene A was 1.85, which is higher than the preset clinical relevance threshold of 1.5 (this threshold was set based on clinicians' expectations for the predictive power of biomarkers and the results of ROC curve analysis), indicating that gene A is significantly associated with the clinical progression of breast cancer.

[0098] The screening and optimization submodule calls the clinical relevance evaluation value, combines the results of transcriptional stability and repeatability assessment, removes low-reproducibility genetic factors, analyzes the screening and optimization extent, corrects the list of candidate genetic factors for breast cancer, and generates biomarker screening results.

[0099] Using the clinical relevance score of 1.85 for gene A, and combining it with the results of transcriptional stability and reproducibility assessments, low-reproducibility genetic factors are eliminated. The transcriptional stability assessment results for gene A are checked; for example, the batch-to-batch reproducibility correlation coefficient for gene A is 0.92, which is higher than 0.70 (0.70 is considered the threshold for high reproducibility, and values ​​below this are considered low reproducibility), indicating high transcriptional stability. Then, based on the preset reproducibility threshold of 0.70, genetic factors with transcriptional reproducibility below this threshold are identified and eliminated. For example, if the reproducibility correlation coefficient for gene X is 0.65, then gene X will be eliminated. The list is generated, and the extent of optimization is analyzed. By comparing the size of the candidate biomarker list before and after removal (e.g., after removing low-reproducibility genetic factors, the list is reduced from 120 to 90, and the average clinical relevance score of the remaining genes increases from 1.65 to 1.78), the list of candidate genetic factors for breast cancer is revised. For example, gene A is retained in the revised candidate list. The biomarker screening results are generated, which is a final list of breast cancer candidate biomarkers that have been evaluated for stability and reproducibility and have high clinical relevance. Gene A is identified as a key member in this screening result.

[0100] Please see Figure 5 The interference elimination module includes:

[0101] The background transcription extraction submodule calls the biomarker screening results to extract the background transcription amount of breast cancer candidate genetic factors in non-tumor tissues, screens genetic factors that meet the background transcription threshold, and obtains the background transcription change rate.

[0102] Using the biomarker screening results containing gene A, the transcriptional value (TPM) of gene A in each normal sample was precisely extracted from the transcriptome data of 50 paired adjacent normal breast tissue samples. For example, the transcriptional value of gene A in normal sample N1 was 12 TPM. Subsequently, the average background transcriptional amount of gene A in all normal samples was calculated, for example, the average was 14 TPM. Genes that met the background transcriptional threshold were screened, and the background transcriptional threshold was set at 20 TPM. This threshold was set based on the statistical distribution characteristics of gene expression in normal tissues. For example, through the analysis of a large number of normal breast tissue samples, it was found that 95% of non-specifically expressed genes had TPM values ​​below 20. The average background transcriptional amount of gene A, 14 TPM, was less than 20 TPM, thus meeting the threshold condition. The background transcriptional variation rate was obtained, and the coefficient of variation of the transcriptional value of gene A in non-tumor tissues was calculated. For example, the standard deviation of gene A in normal tissue samples was 2.1 TPM, and the average was 14 TPM, so the coefficient of variation was 0.15 (i.e., 15%). This variation rate reflects its natural physiological fluctuations among healthy individuals.

[0103] The specific detection submodule calls the background transcriptional change rate to detect the transcriptional differences of breast cancer candidate genetic factors in tumor tissues and non-tumor tissues, and calculates the proportion of non-specific transcribed genetic factors to obtain the non-specific transcriptional proportion value.

[0104] Using a background transcriptional change rate of 15% for gene A, the average transcriptional level of gene A in 50 breast cancer tumor tissue samples (78 TPM) and in 50 paired adjacent normal tissue samples (14 TPM) were obtained. The fold change between the two was calculated (e.g., 78 / 14 ≈ 5.57-fold) and statistically tested (e.g., p-value 1.8e-7). Simultaneously, the proportion of non-specific transcription factors was statistically analyzed. The comparison of the average expression level of gene A in tumor tissues with its average expression level in normal tissues was performed to determine whether the fold change was significantly higher or lower than that in normal tissues, with a fold change greater than 2.0 and a p-value less than 1.8e-7. If the threshold of 0.05 is met, gene A is considered not a non-specific transcription factor. If the differential expression of a certain gene factor is not significant, it is marked as a non-specific transcription factor. By evaluating all candidate genes, the proportion of those marked as non-specific transcription factors is calculated. For example, in the current biomarker screening results, 5% of the genes are identified as non-specific transcription factors, and the non-specific transcription percentage is obtained. For example, the non-specific transcription percentage of gene A is 0%, while the non-specific transcription percentage of the entire candidate list is 5%.

[0105] The noise removal submodule collects the transcriptional noise distribution of breast cancer candidate genetic factors in differential tissues based on the proportion of non-specific transcription, performs a comparison with tissue-specific transcription patterns and error amplitude, identifies the noise removal coefficient, screens out non-specific transcription genetic factors, and generates a list of high-specificity biomarkers.

[0106] Based on the 0% non-specific transcription percentage of gene A, the transcriptional variability of gene A in each tissue type was assessed by analyzing its transcriptional values ​​in breast cancer tumor tissue, normal breast tissue, and other non-breast tissue samples. For example, the standard deviation of transcriptional noise for gene A in tumor tissue was 15 TPM, while it was 3 TPM in normal breast tissue. Comparisons were performed using tissue-specific transcriptional patterns and error magnitudes, comparing the transcriptional values ​​of gene A in tumor tissue with those in other tissues, and considering measurement error (e.g., 5%), to assess whether its specific high expression in tumor tissue exceeded the range of background noise and measurement error. For example, gene A expression in tumors was significantly higher than in normal breast tissue (5.5%). For tumor tissue (7-fold difference) and liver tissue (15.6-fold difference), and with a measurement error greater than 5%, a noise removal coefficient is identified. For gene A, its transcription value in tumor tissue is significantly higher than in all other non-tumor tissues, and the difference fold is greater than 5.0. At the same time, the standard deviation of transcription noise is less than 5 TPM in normal tissues. Therefore, its noise removal coefficient is set to 1.0 (indicating complete retention). Genetic factors with a coefficient lower than 1.0 will be removed. Non-specific transcription genetic factors are screened out. For example, gene A is retained because of its high specificity (noise removal coefficient 1.0). A list of high-specificity biomarkers is generated. For example, gene A is included in the list of high-specificity biomarkers. Biomarkers in this list are highly specifically expressed in breast cancer tissue.

[0107] Please see Figure 6 The optimized and integrated modules include:

[0108] The weight assessment submodule extracts the updated functional weight sequences of breast cancer candidate genetic factors and the initial weight data after background transcription knockout based on a list of highly specific biomarkers. It analyzes the weight change trends of breast cancer candidate genetic factors in the differential functional dimension and performs normalization processing in combination with timestamps to obtain the genetic factor weight change trends.

[0109] Based on a list of highly specific biomarkers including gene A, the updated functional weight sequence was obtained using the functional weight parameter value of gene A (e.g., 1.0). Simultaneously, the initial pathway contribution of gene A before background transcriptional knockout (0.85) was used as the initial weight data. The weight change trends of candidate breast cancer genetic factors on the differential functional dimension were analyzed. By comparing the functional weight of gene A before and after background transcriptional knockout (e.g., an increase from the initial 0.85 to the updated 1.0), the weight increase was quantified as 0.15. Normalization was performed using timestamps, and the weight changes of all candidate genetic factors were correlated with the time point information analyzed. For example, if the latest functional weight of gene A was updated within the last month, its weight change was given a higher influence to ensure the assessment reflects the latest findings. The trend of genetic factor weight changes was obtained; for example, if the weight change trend of gene A showed a significant increase in its functional weight from 0.85 to 1.0, with a trend score of 0.15, it indicates that its value as a potential biomarker has been further confirmed after screening and optimization.

[0110] The mechanism association comparison submodule analyzes the weight change rate and mechanism association strength at corresponding time points based on the trend of genetic factor weight changes and the weight distribution sequence of the role of candidate genetic factors in the lung metastasis mechanism of breast cancer, and judges the stability of the mechanism association interval to obtain mechanism association characteristics.

[0111] Based on the changing trends of genetic factor weights and the distribution sequence of the role weights of candidate breast cancer genetic factors in the lung metastasis mechanism, this study reviewed published literature on breast cancer metastasis to determine the known role strength of gene A in key biological processes of lung metastasis, such as cell migration, invasion, and angiogenesis, and quantified it into a weight value. For example, the role weight of gene A in lung metastasis is 0.75 (range 0 to 1, indicating importance in the metastasis mechanism). The changing trend of gene A's functional weight (e.g., from 0.85 to 1.0, with a change rate of 0.15) was correlated with this lung metastasis role weight (0.75) to assess whether the two changed synergistically and to determine the mechanism. The stability of the mechanism association interval is assessed by monitoring the changes in the functional weight of gene A under different biological backgrounds (e.g., primary tumor and metastatic lesions) and time points, as well as whether its role weight in the lung metastasis mechanism remains stable. If the functional weight of gene A continuously increases and its role weight in the lung metastasis mechanism remains above 0.70, then its mechanism association interval is considered stable. This stability threshold of 0.70 is set based on the general stability of gene functions in breast cancer metastasis studies. Mechanism association characteristics are obtained. For example, the mechanism association characteristics of gene A are: functional weight change rate of 0.15, lung metastasis role weight of 0.75, and the mechanism association interval is stable.

[0112] The evaluation strategy generation submodule calls the mechanism association features and the weight change rate of breast cancer candidate genetic factors at previous time points, sets the weight parameters according to the stability of mechanism association and the magnitude of weight change, and performs weighted processing on each breast cancer candidate genetic factor to generate a biomarker evaluation index.

[0113] The mechanism association features of gene A (functional weight change rate 0.15, lung metastasis weight 0.75, mechanism association interval stable) and the weight change rate of gene A at previous time steps (e.g., in the assessment after removing low-repetition genetic factors, the weight change rate of gene A is 0.10) are used. Weighting parameters are set according to the stability of mechanism association and the magnitude of weight change. First, based on the stability (stable) of gene A's mechanism association, the stability weighting parameter is set to 0.6 (range 0 to 1, representing the influence of this factor on the final evaluation index). Simultaneously, based on the weight change magnitude of gene A (0.15), the weight magnitude weighting parameter is set to 0.4 (range 0 to 1). These weighting parameters are determined based on expert experience and machine learning model training and cross-validation. For example, To maximize prediction accuracy, parameters were optimized using a support vector machine model. Each candidate breast cancer genetic factor was weighted. The weight change rate of gene A (0.15) was multiplied by the weight amplitude parameter (0.4), and the stability parameter related to its mechanism (0.6) was multiplied by the stability parameter (0.6). These two products were then added together. For example, the weighted result for gene A was (0.15⋅0.4) + (0.6⋅0.6) = 0.42. This generated a biomarker evaluation index, a comprehensive score reflecting the overall potential of candidate breast cancer genetic factors as biomarkers. For example, the biomarker evaluation index for gene A was 0.42. A higher index indicates greater reliability and clinical application value as a biomarker.

[0114] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A screening system for breast cancer lung metastasis-related biomarkers, characterized in that, The system includes: The differential genetic factor mining module acquires transcriptome data from breast cancer patient samples, extracts the transcription profiles of candidate genetic factors associated with lung metastasis, analyzes the fluctuation range and distribution characteristics of genetic factor transcription values, identifies differentially transcribed genetic factor sets among differentially transcribed samples, and generates a preliminary list of candidate biomarkers. Based on the preliminary candidate biomarker list, the functional association module combines breast cancer pathology information and molecular pathway annotation data from public databases to analyze the role weights of breast cancer candidate genetic factors in key signaling pathways and generate functional weight ranking results. The validation and screening module, based on the functional weight ranking results, calls independent validation queue sample data to evaluate the transcriptional consistency and clinical relevance of breast cancer candidate genetic factors in the validation queue, eliminates low-repetition genetic factors, and generates biomarker screening results. The interference elimination module calls the biomarker screening results, extracts the background transcription amount of breast cancer candidate genetic factors in non-tumor tissues, combines tissue-specific transcription patterns and noise interference characteristics, eliminates non-specific transcription genetic factors, and generates a list of highly specific biomarkers.

2. The breast cancer lung metastasis-related biomarker screening system according to claim 1, characterized in that, The preliminary candidate biomarker list includes the amplitude of genetic factor transcriptional fluctuation, significant difference score, and transcriptional distribution characteristic value. The functional weight ranking results include signaling pathway contribution, genetic factor interaction strength, and functional association score. The biomarker screening results include validation consistency ratio, clinical relevance index, and transcriptional stability score. The high-specificity biomarker list includes tissue-specific transcription score, background interference coefficient, and noise removal ratio.

3. The breast cancer lung metastasis-related biomarker screening system according to claim 1, characterized in that, The differential genetic factor mining module includes: The transcription profile extraction submodule acquires transcriptome data from breast cancer patient samples, extracts the transcription value of each genetic factor and its distribution characteristics among differential samples, analyzes the standard dispersion and fluctuation trend of genetic factor transcription values, determines whether the difference threshold is met, and obtains differential genetic factor trigger signals. Based on the differential genetic factor triggering signal, the fluctuation analysis submodule extracts the fluctuation range and distribution characteristics of the transcription values ​​of candidate genetic factors in breast cancer, combines the changing trends of genetic factor transcription values ​​among samples, analyzes the consistency of genetic factor transcription fluctuations, and obtains genetic factor transcription fluctuation parameters. The pathway mapping submodule performs functional pathway mapping on candidate genetic factors for breast cancer based on the transcriptional fluctuation parameters of the genetic factors, and generates a preliminary list of candidate biomarkers by combining the distribution location and role weight of the genetic factors in key signaling pathways.

4. The breast cancer lung metastasis-related biomarker screening system according to claim 3, characterized in that, The functional association module includes: Based on the preliminary candidate biomarker list, the pathway annotation submodule extracts signaling pathway annotation information of breast cancer candidate genetic factors in public databases, and combines the upstream and downstream relationships of genetic factors in key pathways to identify the role weight of genetic factors in pathways and obtain a set of pathway contribution values. The interaction parsing submodule calls the pathway contribution value set, detects the core nodes and connection density in the genetic factor interaction network based on the interaction relationship and functional association strength between genetic factors, calculates the genetic factor interaction strength value, and obtains the functional association parameter value by combining the interaction results and functional weights. The weight ranking submodule calls the function-related parameter values, prioritizes the genetic factors based on their contribution and interaction strength in the pathway, determines whether they are located in the core functional region, maps the functional weight parameters, and generates the functional weight ranking result.

5. The breast cancer lung metastasis-related biomarker screening system according to claim 4, characterized in that, The verification and screening module includes: The consistency assessment submodule collects independent validation cohort sample data based on the functional weight ranking results, extracts the transcription value distribution characteristics of breast cancer candidate genetic factors in the validation cohort, identifies transcriptional consistency trends, and generates validation consistency parameters. The correlation analysis submodule calls the validation consistency parameters, combines the correlation between the transcription values ​​of candidate genetic factors in breast cancer and clinical indicators, detects the association strength between genetic factor transcription and clinical phenotype, analyzes the trend of clinical correlation changes, and obtains the clinical correlation evaluation value. The screening and optimization submodule calls the clinical relevance evaluation value, combines the transcriptional stability and repeatability assessment results, eliminates low-reproducibility genetic factors, analyzes the screening and optimization extent, corrects the list of candidate genetic factors for breast cancer, and generates biomarker screening results.

6. The breast cancer lung metastasis-related biomarker screening system according to claim 5, characterized in that, The interference cancellation module includes: The background transcription extraction submodule calls the biomarker screening results to extract the background transcription amount of breast cancer candidate genetic factors in non-tumor tissues, screens genetic factors that meet the background transcription threshold, and obtains the background transcription change rate. The specific detection submodule calls the background transcriptional change rate to detect the transcriptional differences of breast cancer candidate genetic factors in tumor tissues and non-tumor tissues, and calculates the proportion of non-specific transcribed genetic factors to obtain the proportion of non-specific transcription. The noise removal submodule collects the transcriptional noise distribution of breast cancer candidate genetic factors in differential tissues based on the non-specific transcription ratio, performs a comparison with tissue-specific transcription patterns and error amplitude, identifies the noise removal coefficient, screens out non-specific transcription genetic factors, and generates a list of high-specificity biomarkers.

7. The breast cancer lung metastasis-related biomarker screening system according to claim 1, characterized in that, The system also includes an optimization and integration module: Based on the list of highly specific biomarkers, the optimization and integration module combines the updated functional weight ranking results and the validation screening status to evaluate the role weight and clinical application potential of breast cancer candidate genetic factors in the mechanism of breast cancer lung metastasis and generate a biomarker evaluation index. The biomarker evaluation index includes functional weight contribution, clinical application potential score, and mechanism association strength.

8. The breast cancer lung metastasis-related biomarker screening system according to claim 7, characterized in that, The optimization and integration module includes: The weight evaluation submodule extracts the updated functional weight sequences of breast cancer candidate genetic factors and the initial weight data after background transcription knockout based on the list of highly specific biomarkers. It analyzes the weight change trends of breast cancer candidate genetic factors in the differential functional dimension and performs normalization processing in combination with timestamps to obtain the genetic factor weight change trends. The mechanism association comparison submodule analyzes the weight change rate and mechanism association strength at corresponding time points based on the trend of the genetic factor weight change and the weight distribution sequence of the role of candidate genetic factors in the lung metastasis mechanism of breast cancer, and judges the stability of the mechanism association interval to obtain mechanism association characteristics. The evaluation strategy generation submodule calls the mechanism-associated features and the weight change rate of the candidate breast cancer genetic factors at the preceding time step, sets the weight parameters according to the mechanism association stability and the magnitude of weight change, and performs weighted processing on each candidate breast cancer genetic factor to generate a biomarker evaluation index.

Citation Information

Cited By

  • Pathogenic gene prediction method, device and equipment based on phenotypic fingerprints and medium

    CN121838892A