Individualized breast cancer risk detection method based on HRR pathway related genes

By integrating HRR pathway-related genes and clinical data on breast cancer, identifying and screening single nucleotide polymorphism sites, and constructing a personalized risk model, this approach addresses the issues of insufficient data integration and lack of individualized risk assessment in existing technologies, thereby achieving precise breast cancer risk detection and management.

CN121034633APending Publication Date: 2025-11-28钱学庆
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511217247.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

In current breast cancer risk detection, there is insufficient integration of HRR pathway-related gene data with clinical data, single nucleotide polymorphism site screening lacks functional and clinically relevant precision, risk assessment models ignore site interactions, risk calculation and grading lack individualized logic, early warning and intervention measures are not targeted enough, and the effect tracking, evaluation and adjustment mechanisms are imperfect, resulting in low effectiveness of risk management.

Method used

By collecting functional information of HRR pathway-related genes and clinical data on breast cancer from multi-source databases, we identify and screen single nucleotide polymorphism sites that are functionally and clinically associated, perform genotyping and combined analysis, construct individualized risk models, calculate and classify breast cancer risk, formulate differentiated early warning and intervention measures, and track and adjust the intervention effects in real time.

Benefits of technology

It has enabled precise breast cancer risk assessment and individualized management, improved the specificity of detection and the targeted nature of prevention and control, ensured that intervention measures are continuously adapted to individual circumstances, and promoted the upgrade of breast cancer prevention and control from static to dynamic management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034633A_ABST
    Figure CN121034633A_ABST
Patent Text Reader

Abstract

The invention discloses an individual breast cancer risk detection method based on HRR pathway related genes, relates to the technical field of breast cancer risk detection, and aims to solve the problems of inaccurate breast cancer risk detection effect and unobvious subsequent treatment effect. According to the method, genetic typing and clinical data are combined to establish an adaptive model, relative risks are accurately calculated and graded, differential early warning and intervention schemes are formulated for low, medium and high risks, precision from risk detection to intervention is achieved, excessive medical treatment or insufficient intervention is avoided, prevention and control pertinence is improved, a dynamic tracking and optimization mechanism is established, and the risk detection accuracy is improved. Tracking frequency and indexes are set according to risk levels, intervention effects are evaluated regularly, schemes are adjusted, a'detection-intervention-tracking-optimization 'closed loop is formed, it is ensured that intervention measures continuously adapt to individual conditions, long-term risk management and control effectiveness is improved, and breast cancer prevention and control are promoted to be upgraded from static management to dynamic management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of breast cancer risk detection technology, specifically to a personalized breast cancer risk detection method based on HRR pathway-related genes. Background Technology

[0002] Current breast cancer risk detection methods suffer from several shortcomings: insufficient integration of HRR pathway-related gene data with clinical data; lack of functional and clinically relevant precision in single nucleotide polymorphism (SNP) site screening; and non-standardized subtyping methods. Risk assessment models neglect site interactions, risk calculation and grading lack individualized logic, early warning and intervention measures are not targeted enough, and the mechanisms for effect tracking, evaluation, and adjustment are inadequate, resulting in low effectiveness of risk management. Summary of the Invention

[0003] The purpose of this invention is to provide a personalized breast cancer risk detection method based on HRR pathway-related genes, which can solve the problems in the prior art.

[0004] To achieve the above objectives, the present invention provides the following technical solution: A personalized breast cancer risk assessment method based on HRR pathway-related genes includes: Genetic data collection was conducted from different data sources, including functional information of HRR pathway-related genes and clinical data on breast cancer. Single nucleotide polymorphisms (SNPs) were identified from the collected genetic data. Genotyping was performed on the identified SNPs. A combined analysis of the identified SNPs was then conducted, and a risk model was constructed based on the analysis results. The genotyping and the constructed risk model were combined to calculate the relative risk of breast cancer. Breast cancer risk levels were classified based on the relative risk calculation results. Different levels of risk warnings were issued based on the classified breast cancer risk levels. Intervention measures were formulated based on the risk warning results. Real-time effectiveness tracking and evaluation of the formulated intervention measures were conducted, and the intervention plan was adjusted based on the evaluation results.

[0005] Preferably, gene data collection includes functional information of HRR pathway-related genes and clinical data on breast cancer from different data ports, including: Functional information of HRR pathway-related genes includes basic gene information, pathway function annotation, expression patterns, and information on the association between variants and function. Functional information of HRR pathway-related genes was obtained from public gene databases, pathway databases, variant databases, and expression databases. Clinical data on breast cancer includes basic individual information, disease-related information, family history and risk factors, and data on treatment and prognosis; Clinical data on breast cancer were obtained from hospital electronic medical record systems, tumor registry systems, biobanks, clinical trial databases, and literature and public clinical databases. The collected functional information of HRR pathway-related genes and clinical data of breast cancer were preprocessed and standardized. Complete genetic data is obtained after data preprocessing and data standardization.

[0006] Preferably, the single nucleotide polymorphism sites in the collected gene data are identified, including: First, identify the core gene range in the functional information of HRR pathway-related genes in the gene data, including DNA damage recognition, repair signal transduction, and homologous recombination execution; Based on the core gene range, reference sequences are obtained from public gene databases, including the chromosomal location of the gene, gene structural details, and corresponding protein sequences and functional domains; Based on the obtained reference sequence, single nucleotide polymorphism sites are obtained from the variation database, including the single nucleotide polymorphism number, chromosomal location, allele and population frequency data; Screening criteria were used to exclude single nucleotide polymorphisms (SNPs) that had no functional association from the single nucleotide polymorphism sites, and the single nucleotide polymorphisms that affected gene function were obtained after the exclusion. The screening criteria included excluding single nucleotide polymorphisms based on location characteristics, variation type, and functional prediction. Using literature databases and clinical data on breast cancer, we validated the association between single nucleotide polymorphisms affecting gene function and breast cancer risk. Single nucleotide polymorphisms that are not clinically relevant were excluded based on the validation results; Population frequency and genotyping feasibility analysis were performed on the excluded clinically unrelated single nucleotide polymorphisms; Among them, population frequency analysis is for single nucleotide polymorphisms that retain a minor allele frequency of ≥1% in the target population; genotyping feasibility analysis is for excluding single nucleotide polymorphisms that are difficult to genotype using conventional methods. The final recognition site for single nucleotide polymorphisms was obtained after analysis.

[0007] Preferably, the identified single nucleotide polymorphism sites are genotyped, including: The genotyping method is to use PCR-RFLP or AS-PCR. PCR-RFLP identifies and cuts specific DNA sequences using restriction endonucleases, and determines genotypes based on differences in the length of the fragments after digestion. Based on the reference sequence of the single nucleotide polymorphism (SNP) site, PCR-RFLP primers for the SNP site were designed. After the primers were designed, they were synthesized. The purity of the synthesized primers was detected by agarose gel electrophoresis, and the concentration was adjusted to 10 μM for later use. After synthesis, the reaction system is prepared, including template DNA, upstream primer, downstream primer, Taq DNA polymerase, dNTP mixture, 10×PCR buffer and sterile deionized water; After the reaction system was prepared, the PCR-RFLP reaction procedure was carried out. During the process, the corresponding restriction endonuclease was selected according to the enzyme digestion characteristics of the single nucleotide polymorphism site, and the agarose gel was electrophoretically detected. Finally, genotype determination is made based on the implementation results and electrophoresis results, including wild-type homozygotes, mutant homozygotes, and heterozygotes; AS-PCR uses specific primers targeting different alleles at single nucleotide polymorphism sites to directly determine genotype by the presence or absence of AS-PCR products. First, two forward-specific primers were designed for the two alleles at the single nucleotide polymorphism site, namely primer A and primer B. After designing the forward specific primers, set up two AS-PCR reaction systems, and add A specific primer and B specific primer to the two AS-PCR reaction systems respectively; After the positive specific primers were added, the AS-PCR reaction procedure was performed, and the agarose gel electrophoresis was performed simultaneously. Finally, genotype determination is made based on the implementation results and electrophoresis results, including wild-type homozygotes, mutant homozygotes, and heterozygotes; Among them, wild-type homozygotes showed amplification bands only in the A-specific primer system; mutant homozygotes showed amplification bands only in the G-specific primer system; and heterozygotes showed amplification bands in both systems.

[0008] Preferably, the identified single nucleotide polymorphism sites are jointly analyzed, and a risk model is constructed based on the analysis results, including: The identified single nucleotide polymorphism (SNP) sites and breast cancer clinical data were used as independent variables, and the breast cancer clinical data were used as the dependent variable. A regression model was used, and the SNP site variable was added to the regression model after the breast cancer clinical data was included. After the regression model was added, the independent association of the SNP sites in the breast cancer clinical data was evaluated. Interaction analysis was performed based on the independent association. Risk models were constructed based on the interaction analysis results. The risk models were constructed as follows: a regression model was used as the risk model; if the number of single nucleotide polymorphism sites in the interaction analysis results exceeded the preset value or there were complex interactions, a random forest model was used as the risk model. Once the risk model is confirmed, it is trained and evaluated, including discrimination, calibration and clinical usability assessments. The model is then optimized based on the evaluation results to obtain the final risk model.

[0009] Preferably, genotyping and the constructed risk model are combined to calculate the relative risk of breast cancer, including: Transform individual genotyping data into standardized variables that can be identified by risk models; Standardized variables were integrated with corresponding breast cancer clinical data to obtain an individual characteristic dataset. The individual characteristic dataset is input into the risk model, and the risk model determines the absolute risk value of breast cancer based on the input individual characteristic dataset; Once the absolute risk value for breast cancer is determined, a reference population is identified, and this reference population is used as the benchmark for the relative risk value. Finally, the individual relative risk is derived from the relative risk and the absolute risk value of breast cancer, using the following formula: Individual relative risk = absolute risk value of breast cancer ÷ relative risk value.

[0010] Preferably, breast cancer risk levels are classified based on the relative risk calculation results, including: The classification logic is confirmed based on individual relative risk as the core indicator, combined with reference population benchmarks, clinical relevance, and existing clinical guidelines. Based on the confirmed grading logic, individual relative risk is classified into three levels: low risk, medium risk, and high risk. Low risk is defined as a relative risk value ≤ 1; medium risk is defined as a relative risk value > 1 and ≤ a preset threshold; and high risk is defined as a relative risk value > a medium risk threshold. Match individual relative risk with the classified individual relative risk level; After matching is completed, the final risk level data for breast cancer risk is obtained.

[0011] Preferably, risk warnings are issued at different levels based on the classified breast cancer risk levels, including: The warning content is designed based on the breast cancer risk level, and includes risk interpretation, scientific basis, and action recommendations. After the warning content is designed, the warning method is confirmed. Low risk warnings are delivered in the form of written reports or online platform messages; medium risk warnings are delivered in the form of written reports, SMS reminders or telephone notifications, and it is recommended to check the report and contact a doctor as soon as possible; high risk warnings are delivered simultaneously through multiple channels, including written reports, emergency telephone notifications, and face-to-face explanations. Once the warning method is confirmed, corresponding interpretations are provided for individuals at different risk levels. For low-risk individuals, an online FAQ document is provided to answer questions about the warning content; for medium-risk individuals, a consultation hotline is opened, where professional nurses or doctors answer questions about screening details and risk changes; for high-risk individuals, a genetic counselor or breast specialist is arranged to provide one-on-one interpretation, explaining the risk mechanism, intervention options, and data on risk assessments of family members. The corresponding interpretations provide different levels of risk warnings.

[0012] Preferably, intervention measures are formulated based on the results of risk warnings, including: Based on the early warning results of different risk levels, the core objectives of the intervention were identified. The core objective for low-risk cases was to maintain a low-risk status; the core objective for medium-risk cases was to reduce the progression of the risk; and the core objective for high-risk cases was to proactively reduce the probability of disease onset. Intervention frameworks were developed based on the identified core objectives. The low-risk intervention framework included screening management, lifestyle guidance, and regular follow-up; the medium-risk intervention framework included intensive screening, intervention for controllable factors, clinical monitoring, and genetic counseling; and the high-risk intervention framework included intensive screening, risk blocking measures, pathway-targeted intervention, and family risk management. Once the intervention framework is established, individual-specific factors are incorporated and optimized in a personalized manner. These individual-specific factors include genetic, clinical, personal, and social factors. After personalized optimization, an implementation plan is developed for each intervention, including screening, medication, lifestyle, and surgical interventions. Once approved, the final intervention plan is obtained. If the plan fails to pass the review, optimization and adjustments continue until it is approved.

[0013] Preferably, the intervention measures are tracked and evaluated in real time, and the intervention plan is adjusted based on the evaluation results, including: First, establish tracking indicators, including risk control indicators, compliance indicators, and safety indicators; The follow-up plan is then developed. The frequency of low-risk intervention follow-up is once every 6-12 months, through online questionnaires and annual report submissions, with the responsible parties being community doctors or health managers. The frequency of medium-risk intervention follow-up is once every 3-6 months, through offline follow-ups, screening results entering into the system, and telephone follow-ups, with the responsible parties being breast clinic nurses and specialists. The frequency of high-risk intervention follow-up is once a month to once every 3 months, through multidisciplinary team follow-ups and real-time updates of electronic health records, with the responsible parties being genetic counselors coordinating and multidisciplinary team collaboration. The intervention effect is tracked according to the tracking plan and tracking indicators. The intervention effect is comprehensively evaluated every 12-24 months, and the intervention is determined as to whether the target is achieved based on the evaluation results. The evaluation results include target achieved, partial target achieved, and target not achieved. The intervention plan will be adjusted according to the different assessment results. If the target is met, the existing intervention plan will be maintained and the follow-up interval will be extended. If the target is partially met, the unmet targets will be optimized. If the target is not met, the cause will be analyzed and the intervention plan will be reformulated. The final evaluation results and the revised plan will be entered into the system.

[0014] Preferably, the identified single nucleotide polymorphism sites are jointly analyzed, and a risk model is constructed based on the analysis results, including: The collected clinical data on breast cancer were divided into training, validation and test sets based on stratified sampling. Risk folds of single nucleotide polymorphism sites were obtained from literature databases; Based on validated clinical association data from the ClinVar database and GWASCatalog, a predefined single nucleotide polymorphism site-breast cancer risk mapping table was constructed. K-means clustering was used to cluster the allele count data of the training set to obtain the clustering results; the clustering results were then labeled with genotypes based on a predefined single nucleotide polymorphism (SNP) site-breast cancer risk mapping table to generate a single nucleotide polymorphism site genotype-breast cancer risk label mapping table; single nucleotide polymorphism sites in the mapping table whose risk labels are consistent with the clustering results were retained. Using functional nodes of the HRR pathway as target modules, identified single nucleotide polymorphism (SNP) sites are mapped to their corresponding target modules. Pearson correlation coefficients between SNP sites within the target modules are calculated based on training set data. SNP pairs with Pearson correlation coefficients greater than or equal to a preset correlation threshold are identified as strongly cooperating SNP pairs. Based on these selected strongly cooperating SNP pairs, a cooperative network of HRR pathway SNP sites is constructed. Calculate the pathway association strength in the cooperative network of single nucleotide polymorphism sites in the HRR pathway; Based on the clinical association between single nucleotide polymorphism (SNP) genotype and breast cancer risk, a risk score was calculated for each SNP site. The overall risk value of the target module is calculated based on the functional importance weight of the target module, the strength of pathway association, and the risk score of each single nucleotide polymorphism site. Calculate the total risk value of breast cancer based on the comprehensive risk value of the target module; The total risk value of breast cancer in the training set is associated with the clinical label of breast cancer to construct a risk value-clinical label dataset; a logistic regression model is trained based on the training set and the risk value-clinical label dataset to obtain a risk model; the performance of the risk model is evaluated using a validation set; when the performance of the risk model meets the requirements, it is finally validated using a test set; the risk model is obtained when the validation is successful.

[0015] Preferably, intervention measures are formulated based on the results of risk warnings, including: Acquire basic individual information and simultaneously extract the individual's risk-related health behavior records for the past 3 years; define risk scenarios, which are uniquely assigned as mutually exclusive low-risk, medium-risk, or high-risk scenarios based on risk warning results; count the number of compliant health behaviors of an individual under their assigned risk scenarios; retain all risk warning individual data to obtain the first individual data; Obtain the intervention permission association table. When an individual's risk scenario does not match the intervention permission association table, delete the corresponding individual data that does not match and obtain the second individual data. Obtain the preset first blank database; store the first individual data and the second individual data into the first blank database based on the format of individual ID-risk scenario-number of compliance behaviors-basic characteristics; Once all the filtered data has been entered, the first blank database will be used as the risk scenario binding individual database. Acquire historical intervention records and simultaneously collect corresponding clinical characteristic data, aligning the timestamps of both. Extract the entire process trajectory of the intervention behavior record. When the entire process trajectory is uninterrupted and the clinical characteristic data fluctuates within the preset normal range, the corresponding intervention behavior record is regarded as a valid intervention behavior record; otherwise, the corresponding record is removed. Obtain a pre-defined second blank database; store effective intervention behavior records in the second blank database based on the format of behavior ID-individual ID-intervention trajectory-clinical characteristic curve-risk level; once all the filtered data has been entered, use the second blank database as a clinical characteristic verification behavior library; Risk level-risk behavior association pairs are extracted from the clinical feature verification behavior database, and candidate strategies for each risk level-risk behavior combination are collected; the implementation cost of the candidate strategies is calculated; and the implementation benefits of the candidate strategies are calculated. Calculate the cost-benefit ratio based on the implementation costs and benefits of candidate strategies; The clinical feature data distribution of historical individuals is extracted from the clinical feature verification behavior database, and an appropriate feature range is established for each candidate strategy; candidate strategies whose appropriate feature range covers a preset proportion of historical individuals are preferentially retained. From the remaining candidate strategies, for each risk level-risk behavior combination, the strategy with the smallest cost-benefit ratio is selected as the optimal strategy; Iterate through all risk level-risk behavior combinations to obtain the optimal strategy for each combination; input each optimal strategy into a third blank database based on the format of risk level-risk behavior-optimal strategy-cost details-benefit score-adaptation feature range; when all data has been input, use the third blank database as a cost-benefit optimization strategy library. The clinical characteristic verification behavior database and cost-benefit optimization strategy database are used as the intervention strategy database. Based on the risk warning results, the risk scenario binding individual database is queried to obtain the individual risk level and basic characteristics. Then, the intervention strategy database is queried according to the individual risk level and basic characteristics to determine the appropriate intervention strategy.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention provides a personalized breast cancer risk detection method based on HRR pathway-related genes. It integrates HRR pathway gene functional information with breast cancer clinical data, accurately screens single nucleotide polymorphism sites with functional association and clinical significance from multi-source databases, retains sites with minor allele frequencies ≥1%, improves the reliability of the data basis, provides high-quality targets for subsequent subtyping and risk assessment, reduces interference from invalid data, and enhances detection specificity.

[0017] 2. This invention provides a personalized breast cancer risk detection method based on HRR pathway-related genes, constructs a personalized risk assessment system, establishes an adaptation model by combining genotyping and clinical data, accurately calculates and classifies relative risk, and formulates differentiated early warning and intervention plans for low, medium and high risks, so as to achieve precision from risk detection to intervention, avoid over-treatment or insufficient intervention, and improve the targeted nature of prevention and control.

[0018] 3. The present invention provides a personalized breast cancer risk detection method based on HRR pathway-related genes, which establishes a dynamic tracking and optimization mechanism, sets tracking frequency and indicators according to risk level, regularly evaluates the intervention effect and adjusts the plan, forming a closed loop of "detection-intervention-tracking-optimization", ensuring that intervention measures are continuously adapted to individual conditions, improving the effectiveness of long-term risk management, and promoting the upgrade of breast cancer prevention and control from static to dynamic management. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the personalized breast cancer risk detection steps of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Please see Figure 1 This embodiment provides the following technical solution: A personalized breast cancer risk assessment method based on HRR pathway-related genes includes: Genetic data collection was conducted from different data sources, including functional information of HRR pathway-related genes and clinical data on breast cancer. Single nucleotide polymorphisms (SNPs) were identified from the collected genetic data. Genotyping was performed on the identified SNPs. A combined analysis of the identified SNPs was then conducted, and a risk model was constructed based on the analysis results. The genotyping and the constructed risk model were combined to calculate the relative risk of breast cancer. Breast cancer risk levels were classified based on the relative risk calculation results. Different levels of risk warnings were issued based on the classified breast cancer risk levels. Intervention measures were formulated based on the risk warning results. Real-time effectiveness tracking and evaluation of the formulated intervention measures were conducted, and the intervention plan was adjusted based on the evaluation results.

[0022] Genetic data collection was conducted from different data sources, including functional information of HRR pathway-related genes and clinical data on breast cancer, including: Functional information of HRR pathway-related genes includes basic gene information, pathway function annotation, expression patterns, and information on the association between variants and function. Functional information of HRR pathway-related genes was obtained from public gene databases, pathway databases, variant databases, and expression databases. Clinical data on breast cancer includes basic individual information, disease-related information, family history and risk factors, and data on treatment and prognosis; Clinical data on breast cancer were obtained from hospital electronic medical record systems, tumor registry systems, biobanks, clinical trial databases, and literature and public clinical databases. The collected functional information of HRR pathway-related genes and clinical data of breast cancer were preprocessed and standardized, respectively; after data preprocessing and standardization, complete gene data were obtained.

[0023] Specifically, HRR pathway-related gene functional information is obtained from various professional databases, including public gene databases and pathway databases, covering dimensions such as basic gene information, variants, and functional associations. Breast cancer clinical data integrates resources from multiple channels, such as hospital electronic medical records and tumor registry systems, including individual basic information and treatment prognosis. This multi-source data fusion comprehensively captures the association between genes and diseases, improving data richness and reliability. Gene functional information closely aligns with HRR pathway characteristics, while clinical data focuses on key breast cancer indicators, avoiding interference from irrelevant information. This targeted data collection accurately identifies core variables related to breast cancer risk, laying a high-quality data foundation for subsequent analysis and improving the specificity of risk detection. Through data preprocessing and standardization, format differences between data from different sources can be eliminated, outliers corrected, and data consistency and comparability ensured. Complete gene data processed according to standards reduces the impact of systematic errors on subsequent analysis, improves the accuracy of risk model construction, and the combination of comprehensive and standardized gene functional information with clinical data accurately reflects the uniqueness of individual gene characteristics and disease background, making subsequent risk calculations and classifications more aligned with individual circumstances, providing a reliable basis for individualized early warning and intervention.

[0024] The single nucleotide polymorphism sites in the collected genetic data were identified, including: First, identify the core gene range in the functional information of HRR pathway-related genes in the gene data, including DNA damage recognition, repair signal transduction, and homologous recombination execution; Based on the core gene range, reference sequences are obtained from public gene databases, including the chromosomal location of the gene, gene structural details, and corresponding protein sequences and functional domains; Based on the obtained reference sequence, single nucleotide polymorphism sites are obtained from the variation database, including the single nucleotide polymorphism number, chromosomal location, allele and population frequency data; Screening criteria were used to exclude single nucleotide polymorphisms (SNPs) that had no functional association from the single nucleotide polymorphism sites, and the single nucleotide polymorphisms that affected gene function were obtained after the exclusion. The screening criteria included excluding single nucleotide polymorphisms based on location characteristics, variation type, and functional prediction. Using literature databases and clinical data on breast cancer, we validated the association between single nucleotide polymorphisms affecting gene function and breast cancer risk. Single nucleotide polymorphisms that are not clinically relevant were excluded based on the validation results; Population frequency and genotyping feasibility analysis were performed on the excluded clinically unrelated single nucleotide polymorphisms; Among them, population frequency analysis is for single nucleotide polymorphisms that retain a minor allele frequency of ≥1% in the target population; genotyping feasibility analysis is for excluding single nucleotide polymorphisms that are difficult to genotype using conventional methods. The final recognition site for single nucleotide polymorphisms was obtained after analysis.

[0025] Specifically, core genes involved in DNA damage recognition, repair signal transduction, and homologous recombination are identified to define a precise scope for subsequent screening, avoid interference from irrelevant genes, and ensure that identified single nucleotide polymorphism (SNP) sites are closely related to HRR pathway function. This enhances the functional association basis of the sites. Reference sequences containing details such as chromosomal location are obtained from public gene databases, and complete SNP information is obtained from variant databases, providing solid data support for site identification and ensuring the accuracy and comprehensiveness of site information. Sites without functional association are excluded based on location characteristics, variant type, and functional prediction. Sites without clinical association are further excluded by combining literature and clinical data for verification. This multi-layered screening effectively eliminates sites without clinical association. By removing irrelevant sites, the functional relevance and clinical significance of retained sites were significantly improved, reducing noise in subsequent analyses. Sites with a minor allele frequency of ≥1% were retained through population frequency analysis, ensuring that the sites are representative of the target population. The genotyping feasibility analysis excluded sites that were difficult to genotype routinely, taking into account the practicality of clinical applications. This ensured that the finally identified sites were both biologically significant and could meet actual testing needs. The dual checks of functional association screening and clinical relevance verification, combined with population frequency and genotyping feasibility analysis, resulted in single nucleotide polymorphism identification sites that combined functional importance, clinical relevance, and testing feasibility, laying a high-quality foundation for subsequent genotyping and risk model construction.

[0026] Genotyping of identified single nucleotide polymorphism sites includes: The genotyping method is to use PCR-RFLP (restriction fragment length polymorphism) or AS-PCR (allele-specific PCR). PCR-RFLP identifies and cuts specific DNA sequences using restriction endonucleases, and determines genotypes based on differences in the length of the fragments after digestion. Based on the reference sequence of the single nucleotide polymorphism (SNP) site, PCR-RFLP primers for the SNP site were designed. After the primers were designed, they were synthesized. The purity of the synthesized primers was detected by agarose gel electrophoresis, and the concentration was adjusted to 10 μM for later use. After synthesis, the reaction system is prepared, including template DNA, upstream primer, downstream primer, Taq DNA polymerase, dNTP mixture, 10×PCR buffer and sterile deionized water; After the reaction system was prepared, the PCR-RFLP reaction procedure was carried out. During the process, the corresponding restriction endonuclease was selected according to the enzyme digestion characteristics of the single nucleotide polymorphism site, and the agarose gel was electrophoretically detected. Finally, genotype determination is made based on the implementation results and electrophoresis results, including wild-type homozygotes, mutant homozygotes, and heterozygotes; AS-PCR uses specific primers targeting different alleles at single nucleotide polymorphism sites to directly determine genotype by the presence or absence of AS-PCR products. First, two forward-specific primers were designed for the two alleles at the single nucleotide polymorphism site, namely primer A and primer B. After designing the forward specific primers, set up two AS-PCR reaction systems, and add A specific primer and B specific primer to the two AS-PCR reaction systems respectively; After the positive specific primers were added, the AS-PCR reaction procedure was performed, and the agarose gel electrophoresis was performed simultaneously. Finally, genotype determination is made based on the implementation results and electrophoresis results, including wild-type homozygotes, mutant homozygotes, and heterozygotes; Among them, wild-type homozygotes showed amplification bands only in the A-specific primer system; mutant homozygotes showed amplification bands only in the G-specific primer system; and heterozygotes showed amplification bands in both systems.

[0027] Specifically, two classic genotyping methods, PCR-RFLP and AS-PCR, are employed. In PCR-RFLP, primer purity is checked by agarose gel electrophoresis, and the concentration is uniformly adjusted to 10 μM to ensure consistent reaction conditions and reduce experimental errors. Both methods detect products by agarose gel electrophoresis, and genotype can be directly determined based on differences in fragment length or the presence or absence of amplified bands, clearly distinguishing between wild-type homozygotes, mutant homozygotes, and heterozygotes. AS-PCR provides clear definitions of the band characteristics of different genotypes, avoiding subjective judgment bias and improving the accuracy of genotyping results. Neither method requires complex high-end equipment; both can be performed using conventional PCR and electrophoresis equipment, resulting in lower experimental costs and facilitating widespread application in different laboratories. Meanwhile, design schemes targeting the enzyme digestion characteristics or allele features of different single nucleotide polymorphism (SNP) sites can meet diverse genotyping needs, demonstrating outstanding practicality. Accurate genotyping results are a key prerequisite for conducting joint analysis of SNP sites and constructing risk models. Standardized operating procedures and reliable results ensure the scientific rigor of downstream analyses, helping to improve the overall accuracy of breast cancer risk detection. The tables for PCR-RFLP and AS-PCR are as follows:

[0028] To address the shortcomings of existing technologies in breast cancer risk assessment, such as the ineffective integration of HRR pathway single nucleotide polymorphism sites with clinical data, the neglect of site interactions in models, and the lack of precise individualized logic in risk calculation and grading, please refer to [link to relevant documentation]. Figure 1 This embodiment provides the following technical solution: The identified single nucleotide polymorphism sites were jointly analyzed, and a risk model was constructed based on the analysis results, including: The identified single nucleotide polymorphism sites and breast cancer clinical data were used as independent variables, and breast cancer clinical data were used as dependent variables. A regression model was used, and single nucleotide polymorphism (SNP) site variables were added to the clinical data of breast cancer included in the regression model. After incorporating the regression model, the independent association of single nucleotide polymorphism sites in breast cancer clinical data was assessed. Interaction analysis was conducted based on independent associations; Based on the interaction analysis results, a risk model was constructed, which adopted a regression model as the risk model. If the number of single nucleotide polymorphism sites in the interaction analysis results exceeds the preset value or there are complex interactions, then the random forest model will be used as the risk model. Once the risk model is confirmed, model training is performed, and the trained model is evaluated, including discrimination assessment, calibration assessment, and clinical usability assessment. The model is optimized based on the evaluation results, and the final risk model is obtained after optimization.

[0029] Specifically, single nucleotide polymorphism (SNP) sites and breast cancer clinical data are used as independent variables, while breast cancer clinical data is also used as the dependent variable. This clearly establishes a framework for the association between gene characteristics and disease phenotypes, incorporating both genetic factors at the gene level and clinical phenotypic information, making the analysis more comprehensive and providing rich predictive variables for the model. The basic regression model can effectively assess the independent association and interaction of SNP sites and is suitable for simple variable relationships. When the number of SNP sites exceeds the preset value or complex interactions exist, the model is switched to a random forest model. Leveraging its advantages in handling high-dimensional data and complex interactions, the model's adaptability to data features is ensured, improving predictive accuracy. The model is evaluated through three dimensions: discrimination, calibration, and clinical applicability. The evaluation model not only verifies the model's predictive efficacy (discrimination) and the degree of agreement between the predicted results and reality (calibration), but also focuses on the model's application value in clinical scenarios. It avoids "purely theoretical" models and ensures that the final model can effectively assist clinical decision-making. Based on the evaluation results, the model is optimized to form a closed loop of "construction-evaluation-optimization". This can continuously correct model biases, enhance its stability and reliability, and make the final risk model more in line with actual application needs. It provides accurate tool support for subsequent relative risk calculation of breast cancer. By integrating the joint analysis of gene loci and clinical data, the constructed risk model can capture the unique genetic and clinical characteristics of individuals, laying a key foundation for personalized breast cancer risk detection and helping to achieve accurate risk stratification and early warning.

[0030] The process combines genotyping with a constructed risk model to calculate the relative risk of breast cancer, including: Transform individual genotyping data into standardized variables that can be identified by risk models; Standardized variables were integrated with corresponding breast cancer clinical data to obtain an individual characteristic dataset. Individual characteristic datasets are input into the risk model, and the risk model determines the absolute risk value of breast cancer based on the input individual characteristic datasets; Once the absolute risk value for breast cancer is determined, a reference population is identified, and this reference population is used as the benchmark for the relative risk value. Finally, the individual relative risk is derived from the relative risk and the absolute risk value of breast cancer, using the following formula: Individual relative risk = absolute risk value of breast cancer ÷ relative risk value.

[0031] Specifically, individual genotyping data is converted into standardized variables identifiable by the risk model, eliminating data format interference. This data is then integrated with clinical data to form an individual characteristic dataset, achieving a fusion of genetic and clinical information and providing comprehensive and consistent reliable input for risk calculation. First, the absolute risk value of an individual's breast cancer is confirmed through the risk model. Then, relative risk is calculated based on a reference population, forming a dual "absolute + relative" calculation model. This reflects both the individual's actual risk and allows individuals to intuitively understand their own risk through group comparison, enhancing the interpretability of the results. The formula "Individual relative risk = Absolute risk value of breast cancer ÷ Relative risk" establishes a quantitative relationship between the two. The calculation is transparent and combines previous risk models and genetic data, ensuring that the results accurately reflect the comprehensive impact of genetics and clinical background on risk. Integrating the calculation results can characterize individual risk differences, providing quantitative indicators for risk grading, early warning, and intervention formulation, forming a closed-loop risk detection logic and improving the scientific rigor and practicality of individualized assessment.

[0032] Breast cancer risk levels are classified based on the relative risk calculation results, including: The classification logic is confirmed based on individual relative risk as the core indicator, combined with reference population benchmarks, clinical relevance, and existing clinical guidelines. Based on the confirmed grading logic, individual relative risk is classified into three levels: low risk, medium risk, and high risk. Low risk is defined as a relative risk value ≤ 1; medium risk is defined as a relative risk value > 1 and ≤ a preset threshold; and high risk is defined as a relative risk value > a medium risk threshold. Match individual relative risk with the classified individual relative risk level; After matching is completed, the final risk level data for breast cancer risk is obtained.

[0033] Specifically, the approach centers on individual relative risk, combining reference population benchmarks, clinical relevance, and clinical guidelines to achieve multi-dimensional consideration: reference population benchmarks ensure the grading aligns with group characteristics, clinical relevance ensures the grading closely reflects actual disease manifestations, and guidelines guarantee professional and standardized protocols. This multi-faceted approach enhances the rigor of the grading system and avoids bias from single indicators. Risk is clearly categorized into low, medium, and high levels, with clear quantitative standards: a relative risk value ≤1 indicates low risk, >1 and ≤ a preset threshold indicates medium risk, and > a medium risk threshold indicates high risk. This facilitates accurate assessment in practice, allowing healthcare professionals to quickly identify risks and individuals to intuitively understand their own risk levels. Based on the precisely calculated relative risk values, individuals are matched with corresponding levels to ensure the risk level accurately reflects the individual's situation, providing a precise basis for subsequent risk warnings and intervention formulation. This grading is a crucial link between risk calculation and intervention, transforming abstract risk values ​​into concrete levels, making warnings more targeted and interventions more precise, and promoting a complete and efficient closed loop in risk detection and management.

[0034] To address the shortcomings of existing technologies in breast cancer risk detection, such as the lack of personalized analysis based on HRR pathway genes, insufficient targeting of early warning methods and interventions, and inadequate effectiveness tracking, evaluation, and adjustment mechanisms, which lead to low effectiveness in risk management, please refer to [link to relevant documentation]. Figure 1 This embodiment provides the following technical solution: Different levels of risk warnings are issued based on the classified breast cancer risk levels, including: The warning content is designed based on the breast cancer risk level, and includes risk interpretation, scientific basis, and action recommendations. After the warning content is designed, the warning method is confirmed. Low risk warnings are delivered in the form of written reports or online platform messages; medium risk warnings are delivered in the form of written reports, SMS reminders or telephone notifications, and it is recommended to check the report and contact a doctor as soon as possible; high risk warnings are delivered simultaneously through multiple channels, including written reports, emergency telephone notifications, and face-to-face explanations. Once the warning method is confirmed, corresponding interpretations are provided for individuals at different risk levels. For low-risk individuals, an online FAQ document is provided to answer questions about the warning content; for medium-risk individuals, a consultation hotline is opened, where professional nurses or doctors answer questions about screening details and risk changes; for high-risk individuals, a genetic counselor or breast specialist is arranged to provide one-on-one interpretation, explaining the risk mechanism, intervention options, and data on risk assessments of family members. The corresponding interpretations provide different levels of risk warnings.

[0035] Specifically, early warning content is designed for different risk levels, including risk interpretation, scientific basis, and action recommendations. This ensures individuals clearly understand their risk status and underlying causes, while also providing clear directions for subsequent actions, preventing warnings from becoming mere formalities and ensuring they effectively serve their risk warning and guidance functions. Low-risk individuals receive written reports or online messages for brevity and efficiency; medium-risk individuals receive additional SMS or telephone reminders to enhance awareness; and high-risk individuals receive information through multiple channels simultaneously to ensure timely delivery. This tiered delivery approach aligns with the urgency of different risk levels while also allocating resources rationally, avoiding excessive disruption to low-risk individuals and ensuring sufficient attention for high-risk individuals. Low-risk individuals receive an online FAQ document to meet basic consultation needs; medium-risk individuals have access to a professional consultation hotline to answer questions about screening details; and high-risk individuals receive one-on-one consultations with genetic counselors or specialists to deeply analyze risk mechanisms and intervention options. The tiered interpretation service takes into account the different needs of individuals at different risk levels, ensuring the professionalism and relevance of the interpretation, improving individuals' understanding and acceptance of early warning information, and effectively reducing individuals' cognitive biases about risk through precise content, appropriate methods and professional interpretation. This promotes timely intervention measures for high-risk individuals and health management for medium- and low-risk individuals, and drives the transformation of risk warning from "information transmission" to "behavioral change", laying a good foundation for the implementation of subsequent intervention measures.

[0036] Intervention measures will be formulated based on the results of risk warnings, including: Based on the early warning results of different risk levels, the core objectives of the intervention were identified. The core objective for low-risk cases was to maintain a low-risk status; the core objective for medium-risk cases was to reduce the progression of the risk; and the core objective for high-risk cases was to proactively reduce the probability of disease onset. Intervention frameworks were developed based on the identified core objectives. The low-risk intervention framework included screening management, lifestyle guidance, and regular follow-up; the medium-risk intervention framework included intensive screening, intervention for controllable factors, clinical monitoring, and genetic counseling; and the high-risk intervention framework included intensive screening, risk blocking measures, pathway-targeted intervention, and family risk management. Once the intervention framework is established, individual-specific factors are incorporated and optimized in a personalized manner. These individual-specific factors include genetic, clinical, personal, and social factors. After personalized optimization, an implementation plan is developed for each intervention, including screening, medication, lifestyle, and surgical interventions; Once the implementation plan is finalized, the medical team reviews it. The medical team includes a breast surgeon, a genetic counselor, an oncology pharmacist, and a nutritionist. If the plan is approved, the final intervention plan will be obtained. If the plan is not approved, the plan will continue to be optimized and adjusted until it is approved.

[0037] Specifically, core objectives are established based on risk levels: low-risk focuses on maintaining low risk, medium-risk emphasizes reducing risk progression, and high-risk aims to proactively reduce the probability of disease onset, avoiding insufficient or excessive intervention and enhancing targeting. This forms a progressive intervention system: low-risk involves basic screening management and lifestyle guidance; medium-risk involves enhanced screening and intervention for controllable factors; and high-risk involves intensive screening and risk blocking measures, balancing prevention, monitoring, and management. Customized plans are incorporated based on genes, clinical characteristics, and individual preferences, such as adjusting targeting strategies based on gene mutations and selecting lifestyle interventions according to individual wishes, improving acceptability and implementability. Specific intervention plans are developed and reviewed by a multidisciplinary team including breast surgeons and genetic counselors. Plans that do not pass review are optimized to ensure scientific rigor and safety, forming a closed loop with previous risk level classification and early warning systems to promote effective prevention and control.

[0038] Real-time effectiveness tracking and evaluation are conducted based on the established intervention measures, and the intervention plan is adjusted according to the evaluation results, including: First, establish tracking indicators, including risk control indicators, compliance indicators, and safety indicators; The follow-up plan is then developed. The frequency of low-risk intervention follow-up is once every 6-12 months, through online questionnaires and annual report submissions, with the responsible parties being community doctors or health managers. The frequency of medium-risk intervention follow-up is once every 3-6 months, through offline follow-ups, screening results entering into the system, and telephone follow-ups, with the responsible parties being breast clinic nurses and specialists. The frequency of high-risk intervention follow-up is once a month to once every 3 months, through multidisciplinary team follow-ups and real-time updates of electronic health records, with the responsible parties being genetic counselors coordinating and multidisciplinary team collaboration. The intervention effect is tracked according to the tracking plan and tracking indicators. The intervention effect is comprehensively evaluated every 12-24 months, and the intervention is determined as to whether the target is achieved based on the evaluation results. The evaluation results include target achieved, partial target achieved, and target not achieved. The intervention plan will be adjusted according to the different assessment results. If the target is met, the existing intervention plan will be maintained and the follow-up interval will be extended. If the target is partially met, the unmet targets will be optimized. If the target is not met, the cause will be analyzed and the intervention plan will be reformulated. The final evaluation results and the revised plan will be entered into the system.

[0039] Specifically, a three-dimensional indicator system is established, encompassing risk control, adherence, and safety, to comprehensively measure intervention effectiveness and avoid the limitations of a single indicator. Follow-up plans are developed according to risk levels: low-risk cases are tracked online by community doctors every 6-12 months; medium-risk cases are followed up offline and by phone by specialists every 3-6 months; and high-risk cases are followed up monthly to every 3 months by a multidisciplinary team, with real-time updates to electronic records, balancing intervention intensity and resource efficiency. A comprehensive evaluation is conducted every 12-24 months, with dynamic adjustments based on whether the targets are met (partially met, non-metres). For those meeting the targets, the plan is maintained and the interval extended; for those partially meeting the targets, non-metres are optimized; and for those not meeting the targets, the plan is revised to avoid rigidity. Each stage is closely integrated, forming a logical closed loop with previous intervention measures, shifting risk control from "static intervention" to "dynamic management." The responsible parties for tracking each risk level are clearly defined, with high-risk cases emphasizing multidisciplinary collaboration to ensure execution and accuracy, and enhance the long-term value of individualized interventions.

[0040] The identified single nucleotide polymorphism sites were jointly analyzed, and a risk model was constructed based on the analysis results, including: The collected clinical data on breast cancer were divided into training, validation and test sets based on stratified sampling. Risk folds of single nucleotide polymorphism sites were obtained from literature databases; Based on validated clinical association data from the ClinVar database and GWASCatalog, a predefined single nucleotide polymorphism site-breast cancer risk mapping table was constructed. K-means clustering was used to cluster the allele count data of the training set to obtain the clustering results; the clustering results were then labeled with genotypes based on a predefined single nucleotide polymorphism (SNP) site-breast cancer risk mapping table to generate a single nucleotide polymorphism site genotype-breast cancer risk label mapping table; single nucleotide polymorphism sites in the mapping table whose risk labels are consistent with the clustering results were retained. Using functional nodes of the HRR pathway as target modules, identified single nucleotide polymorphism (SNP) sites are mapped to their corresponding target modules. Pearson correlation coefficients between SNP sites within the target modules are calculated based on training set data. SNP pairs with Pearson correlation coefficients greater than or equal to a preset correlation threshold are identified as strongly cooperating SNP pairs. Based on these selected strongly cooperating SNP pairs, a cooperative network of HRR pathway SNP sites is constructed. Calculate the pathway association strength in the cooperative network of single nucleotide polymorphism sites in the HRR pathway; Based on the clinical association between single nucleotide polymorphism (SNP) genotype and breast cancer risk, a risk score was calculated for each SNP site. The overall risk value of the target module is calculated based on the functional importance weight of the target module, the strength of pathway association, and the risk score of each single nucleotide polymorphism site. Calculate the total risk value of breast cancer based on the comprehensive risk value of the target module; The total risk value of breast cancer in the training set is associated with the clinical label of breast cancer to construct a risk value-clinical label dataset; a logistic regression model is trained based on the training set and the risk value-clinical label dataset to obtain a risk model; the performance of the risk model is evaluated using a validation set; when the performance of the risk model meets the requirements, it is finally validated using a test set; the risk model is obtained when the validation is successful.

[0041] In this embodiment, ClinVar is a free and open public database maintained by the National Center for Biotechnology Information in the United States, which focuses on recording the association between human genetic variations and phenotypes (such as diseases) and providing supporting evidence.

[0042] In this embodiment, GWASCatalog is a genome-wide association study (GWAS) database maintained by the European Institute for Bioinformatics (EMBL-EBI) to integrate GWAS research data from around the world.

[0043] In this embodiment, the allele count data of the samples, the genotype of each sample at the single nucleotide polymorphism site, 0 / 1 / 2 correspond to AA / Aa / aa respectively.

[0044] In this embodiment, the single nucleotide polymorphism site genotype-breast cancer risk label is illustrated by, for example, the aa genotype of BRCA1rs799917 has a 2.3 times higher risk of developing breast cancer than the AA genotype.

[0045] In this example, the functional nodes are divided into DNA damage recognition, repair signal transduction, and homologous recombination execution.

[0046] In this embodiment, the pathway association strength in the HRR pathway single nucleotide polymorphism site cooperating network is calculated, including: ; in, This indicates the strength of pathway associations in the co-operative network of single nucleotide polymorphism sites in the HRR pathway; Indicates the number of strongly cooperating single nucleotide polymorphism (SNP) sites within the target module; This represents the risk fold of the i-th strongly cooperating single nucleotide polymorphism pair; This represents the average correlation coefficient of strongly cooperating single nucleotide polymorphism (SNP) site pairs.

[0047] In this embodiment, based on the clinical association between single nucleotide polymorphism (SNP) genotype and breast cancer risk, a risk score is calculated for each SNP, including: ; in, This represents the risk score for the i-th single nucleotide polymorphism site; This represents the breast cancer risk multiple for the genotype at the i-th single nucleotide polymorphism site; β represents the expression effect weighting coefficient, which is determined by fitting the training set data, and its range is 0 ≤ β ≤ 1. , representing the expression change rate of the gene containing the i-th single nucleotide polymorphism site; This indicates the expression level of the gene containing the single nucleotide polymorphism site in breast cancer tissue; This indicates the expression level of the same gene in normal breast tissue.

[0048] In this embodiment, based on the functional importance weight of the target module, the strength of pathway association, and the risk score of each single nucleotide polymorphism site, the comprehensive risk value of the target module is calculated, including: ; in, This represents the overall risk value of the m-th target module; This represents the functional importance weight of the m-th target module; This represents the path association strength of the m-th target module; This represents the average correlation coefficient between the i-th single nucleotide polymorphism (SNP) site and other SNP sites within the target module; This represents the total number of single nucleotide polymorphism sites within the target module.

[0049] In this embodiment, the functional importance weights are initially determined based on the contribution of the target modules in the HRR pathway to DNA repair efficiency through a systematic literature review; the initial weights are then corrected using training set data to ensure that the weights are consistent with the biological rationality of breast cancer risk; and the functional importance weights of the three target modules are then determined.

[0050] In this embodiment, the total risk value of breast cancer is calculated based on the comprehensive risk value of the target module, including: ; in, This represents the overall risk value for breast cancer. 3 represents the overall risk value of the m-th target module; 3 represents the number of core modules in the HRR pathway, namely DNA damage recognition, repair signal transduction, and homologous recombination execution.

[0051] The working principle and beneficial effects of the above technical solution are as follows: Stratified sampling is used to divide breast cancer clinical data into training, validation, and test sets, ensuring a consistent ratio of diseased / non-disease samples and improving the model's generalization ability. Evidence of the association between single nucleotide polymorphism (SNP) sites and breast cancer is obtained from authoritative databases such as ClinVar, establishing a predefined "SNP→risk" mapping. K-means clustering is performed on the allele count data in the training set, combined with back-labeling of the predefined mapping, to screen high-confidence SNPs whose clustering and risk labels are consistent. These are mapped to the HRR pathway, and Pearson correlation coefficients are calculated to identify "strongly cooperating SNP pairs," constructing a pathway-specific SNP cooperating network. Pathway association strength, individual SNP risk scores, and module-wide risk values ​​are calculated, ultimately integrating to obtain an individualized total genetic risk. Using the total risk value of the training set as a feature, a dataset is constructed with real clinical labels. A logistic regression model is used to learn nonlinear relationships. After evaluation on the validation set and validation on the test set, an interpretable and reproducible breast cancer genetic risk prediction model is output.

[0052] Intervention measures will be formulated based on the results of risk warnings, including: Acquire basic individual information and simultaneously extract the individual's risk-related health behavior records for the past 3 years; define risk scenarios, which are uniquely assigned as mutually exclusive low-risk, medium-risk, or high-risk scenarios based on risk warning results; count the number of compliant health behaviors of an individual under their assigned risk scenarios; retain all risk warning individual data to obtain the first individual data; Obtain the intervention permission association table. When an individual's risk scenario does not match the intervention permission association table, delete the corresponding individual data that does not match and obtain the second individual data. Retrieve the preset first blank database; The first and second individual data are stored in the first blank database based on the format of individual ID-risk scenario-number of compliance behaviors-basic characteristics; Once all the filtered data has been entered, the first blank database will be used as the risk scenario binding individual database. Acquire historical intervention records and simultaneously collect corresponding clinical characteristic data, aligning the timestamps of both. Extract the entire process trajectory of the intervention behavior record. When the entire process trajectory is uninterrupted and the clinical characteristic data fluctuates within the preset normal range, the corresponding intervention behavior record is regarded as a valid intervention behavior record; otherwise, the corresponding record is removed. Retrieve the preset second blank database; Effective intervention behavior records are stored in a second blank database based on the format of behavior ID-individual ID-intervention trajectory-clinical characteristic curve-risk level; once all the filtered data has been entered, the second blank database is used as a clinical characteristic verification behavior library. Risk level-risk behavior association pairs are extracted from the clinical feature verification behavior database, and candidate strategies for each risk level-risk behavior combination are collected; the implementation cost of the candidate strategies is calculated; and the implementation benefits of the candidate strategies are calculated. Calculate the cost-benefit ratio based on the implementation costs and benefits of candidate strategies; The clinical feature data distribution of historical individuals is extracted from the clinical feature verification behavior database, and an appropriate feature range is established for each candidate strategy; candidate strategies whose appropriate feature range covers a preset proportion of historical individuals are preferentially retained. From the remaining candidate strategies, for each risk level-risk behavior combination, the strategy with the smallest cost-benefit ratio is selected as the optimal strategy; Iterate through all risk level-risk behavior combinations to obtain the optimal strategy for each combination; input each optimal strategy into the third blank database based on the format of risk level-risk behavior-optimal strategy-cost details-benefit score-adaptation feature range. Once all data has been entered, the third blank database will be used as a cost-benefit optimization strategy library. The clinical characteristic verification behavior database and cost-benefit optimization strategy database will be used as an intervention strategy database. Based on the risk warning results, the risk scenario is bound to the individual database to obtain the individual risk level and basic characteristics. Then, based on the individual risk level and basic characteristics, the intervention strategy database is queried to determine the appropriate intervention strategy.

[0053] In this embodiment, the individual's basic information includes: age, BRCA gene status, and family history of breast cancer; risk-related health behavior records include breast cancer screening time, breast cancer screening results, exercise frequency, and history of estrogen drug use.

[0054] In this embodiment, an intervention permission association table is obtained, as shown in Table 1. When an individual's risk scenario does not match the intervention permission association table: Table 1

[0055] Example of a mismatch scenario: Individual's current medical institution: a community health service center (a primary healthcare institution); Risk scenario requirement: high-risk scenario requires PARP inhibitor treatment at a tertiary oncology hospital; Matching examination: Risk scenario = high-risk scenario; Minimum required medical institution level = tertiary oncology hospital; Actual medical institution level = primary (community health service center); Result: Severe mismatch; Processing result: The individual's data is automatically marked as "permission mismatch" by the system, deleted from the second individual dataset, and a referral process is triggered.

[0056] In this embodiment, the number of compliant health behaviors of an individual in the risk scenarios assigned to them is counted for subsequent strategy adaptation and is not used as a data filtering condition.

[0057] In this embodiment, the intervention behavior record includes the entire process of screening, diagnosis, treatment, and follow-up, such as mammography reports and biopsy results; clinical characteristic data include breast density, estrogen levels, and nodule BI-RADS classification.

[0058] In this embodiment, the entire process trajectory of the intervention behavior record is recorded, for example: mammography → BI-RADS 3 → 6-month follow-up → BI-RADS 2 → annual screening. There is no interruption, meaning there are no fragmented records of "screening only without follow-up".

[0059] In this embodiment, the implementation cost of the candidate strategy is calculated, for example: economic cost (e.g., molybdenum ¥200 / time, surgery ¥30,000 / time); physical cost (e.g., tamoxifen thrombosis rate 3%, surgical complication rate 5%); time cost (e.g., screening takes 2 hours, rehabilitation takes 30 days).

[0060] In this embodiment, calculating the implementation benefits of candidate strategies includes: B = (Probability of survival of risk in 1-5 years) × 80 + (Value of improvement in quality of life) × 20 Here, B represents the implementation benefit of the candidate strategy; the 5-year risk survival probability is calculated based on the breast cancer risk assessment model / BRCA mutation risk, and the quality of life improvement value is assessed through the SF-36 scale; BRCA pathogenic variants mainly refer to harmful mutations on the BRCA1 and BRCA2 genes, which significantly increase an individual's risk of developing breast cancer, ovarian cancer, and other cancers.

[0061] In this embodiment, the cost-benefit ratio is the ratio of implementation costs to implementation benefits.

[0062] The working principle and beneficial effects of the above technical solution are as follows: Through a three-layer screening process (unique risk scenario allocation of mutually exclusive low / medium / high risk, compliance health behavior verification, and intervention permission matching), it ensures that the individual risk level is unique and accurate, avoiding excessive intervention for low-risk individuals (such as preventative surgery) and insufficient intervention for high-risk individuals (such as missed screening), allowing intervention resources to focus on the real risk groups. Simultaneously, through full-process trajectory verification (eliminating fragmented records), clinical characteristic range screening (ensuring the physiological compliance of data), and filtering of invalid / abnormal intervention records, it ensures the authenticity and traceability of historical data upon which the strategy library relies, improving the scientific rigor of intervention plans. Furthermore, through cost-benefit ratio quantitative calculation (integrating multiple costs and linking risk probability with quality of life), the optimal strategy is selected, avoiding high-cost, inefficient solutions (such as blindly using expensive targeted drugs); the strategy's applicability range is defined by combining historical clinical characteristics (such as prioritizing MRI for high breast density groups), balancing universality and individual differences. Finally, a "risk individual library → verification behavior library → optimized strategy library" linkage is constructed, replacing traditional manual decision-making, shortening the intervention plan development time, and improving clinical response speed.

[0063] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.

Claims

1. A personalized breast cancer risk detection method based on HRR pathway-related genes, characterized in that, include: Gene data collection was conducted from different data ports to gather functional information of HRR pathway-related genes and clinical data on breast cancer. Single nucleotide polymorphism sites were identified in the collected gene data; the identified single nucleotide polymorphism sites were then genotyped. The identified single nucleotide polymorphism sites are then jointly analyzed, and a risk model is constructed based on the analysis results. The genotyping and the constructed risk model are combined, and the relative risk of breast cancer is calculated. The risk level of breast cancer is classified based on the relative risk calculation results. Different levels of risk warning are issued based on the classified breast cancer risk level. Intervention measures are formulated based on the risk warning results; the effects of the formulated intervention measures are tracked and evaluated in real time, and the intervention plan is adjusted based on the evaluation results.

2. The personalized breast cancer risk detection method based on HRR pathway-related genes according to claim 1, characterized in that, Genetic data collection was conducted from different data sources, including functional information of HRR pathway-related genes and clinical data on breast cancer, including: Functional information of HRR pathway-related genes includes basic gene information, pathway function annotation, expression patterns, and information on the association between variants and function. Functional information of HRR pathway-related genes was obtained from public gene databases, pathway databases, variant databases, and expression databases. Clinical data on breast cancer includes basic individual information, disease-related information, family history and risk factors, and data on treatment and prognosis; Clinical data on breast cancer were obtained from hospital electronic medical record systems, tumor registry systems, biobanks, clinical trial databases, and literature and public clinical databases. The collected functional information of HRR pathway-related genes and clinical data of breast cancer were preprocessed and standardized. Complete genetic data is obtained after data preprocessing and data standardization.

3. The personalized breast cancer risk detection method based on HRR pathway-related genes according to claim 2, characterized in that, The single nucleotide polymorphism sites in the collected genetic data were identified, including: First, identify the core gene range in the functional information of HRR pathway-related genes in the gene data, including DNA damage recognition, repair signal transduction, and homologous recombination execution; Based on the core gene range, reference sequences are obtained from public gene databases, including the chromosomal location of the gene, gene structural details, and corresponding protein sequences and functional domains; Based on the obtained reference sequence, single nucleotide polymorphism sites are obtained from the variation database, including the single nucleotide polymorphism number, chromosomal location, allele and population frequency data; Screening criteria were used to exclude single nucleotide polymorphisms (SNPs) that had no functional association from the single nucleotide polymorphism sites, and the single nucleotide polymorphisms that affected gene function were obtained after the exclusion. The screening criteria included excluding single nucleotide polymorphisms based on location characteristics, variation type, and functional prediction. Using literature databases and clinical data on breast cancer, we validated the association between single nucleotide polymorphisms affecting gene function and breast cancer risk. Single nucleotide polymorphisms that are not clinically relevant were excluded based on the validation results; Population frequency and genotyping feasibility analysis were performed on the excluded clinically unrelated single nucleotide polymorphisms; Among them, population frequency analysis is for single nucleotide polymorphisms that retain a minor allele frequency of ≥1% in the target population; genotyping feasibility analysis is for excluding single nucleotide polymorphisms that are difficult to genotype using conventional methods. The final recognition site for single nucleotide polymorphisms was obtained after analysis.

4. The personalized breast cancer risk detection method based on HRR pathway-related genes according to claim 3, characterized in that, Genotyping of identified single nucleotide polymorphism sites includes: The genotyping method is to use PCR-RFLP or AS-PCR. PCR-RFLP identifies and cuts specific DNA sequences using restriction endonucleases, and determines genotypes based on differences in fragment length after digestion. Based on the reference sequence of the single nucleotide polymorphism (SNP) site, PCR-RFLP primers for the SNP site were designed. After the primers were designed, they were synthesized. The purity of the synthesized primers was detected by agarose gel electrophoresis, and the concentration was adjusted to 10 μM for later use. After synthesis, the reaction system is prepared, including template DNA, upstream primer, downstream primer, Taq DNA polymerase, dNTP mixture, 10×PCR buffer and sterile deionized water; After the reaction system was prepared, the PCR-RFLP reaction procedure was carried out. During the process, the corresponding restriction endonuclease was selected according to the enzyme digestion characteristics of the single nucleotide polymorphism site, and the agarose gel was electrophoretically detected. Finally, genotype determination is made based on the implementation results and electrophoresis results, including wild-type homozygotes, mutant homozygotes, and heterozygotes; AS-PCR uses specific primers targeting different alleles at single nucleotide polymorphism sites to directly determine genotype by the presence or absence of AS-PCR products. First, two forward-specific primers were designed for the two alleles at the single nucleotide polymorphism site: primer A and primer B. After designing the forward specific primers, two AS-PCR reaction systems were set up, with A-specific primers and B-specific primers added to the two AS-PCR reaction systems respectively; After the positive specific primers were added, the AS-PCR reaction procedure was performed, and the agarose gel electrophoresis was performed simultaneously. Finally, genotype determination is made based on the implementation results and electrophoresis results, including wild-type homozygotes, mutant homozygotes, and heterozygotes; Among them, wild-type homozygotes showed amplification bands only in the A-specific primer system; mutant homozygotes showed amplification bands only in the G-specific primer system; and heterozygotes showed amplification bands in both systems.

5. A personalized breast cancer risk detection method based on HRR pathway-related genes according to claim 4, characterized in that, The identified single nucleotide polymorphism sites were jointly analyzed, and a risk model was constructed based on the analysis results, including: The identified single nucleotide polymorphism sites and breast cancer clinical data were used as independent variables, and breast cancer clinical data were used as dependent variables. A regression model was used, and single nucleotide polymorphism (SNP) site variables were added to the clinical data of breast cancer included in the regression model. After incorporating the regression model, the independent association of single nucleotide polymorphism sites in breast cancer clinical data was assessed. Interaction analysis was conducted based on independent associations; Based on the interaction analysis results, a risk model was constructed, which adopted a regression model as the risk model. If the number of single nucleotide polymorphism sites in the interaction analysis results exceeds the preset value or there are complex interactions, then the random forest model will be used as the risk model. Once the risk model is confirmed, model training is performed, and the trained model is evaluated, including discrimination assessment, calibration assessment, and clinical usability assessment. The model is optimized based on the evaluation results, and the final risk model is obtained after optimization. Meanwhile, the collected breast cancer clinical data were divided into training, validation and test sets based on stratified sampling. Risk folds of single nucleotide polymorphism sites were obtained from literature databases; Based on validated clinical association data from the ClinVar database and GWASCatalog, a predefined single nucleotide polymorphism site-breast cancer risk mapping table was constructed. K-means clustering was used to cluster the allele count data of the training set to obtain the clustering results; the clustering results were then labeled with genotypes based on a predefined single nucleotide polymorphism (SNP) site-breast cancer risk mapping table to generate a single nucleotide polymorphism site genotype-breast cancer risk label mapping table; single nucleotide polymorphism sites in the mapping table whose risk labels are consistent with the clustering results were retained. Using functional nodes of the HRR pathway as target modules, identified single nucleotide polymorphism (SNP) sites are mapped to their corresponding target modules. Pearson correlation coefficients between SNP sites within the target modules are calculated based on training set data. SNP pairs with Pearson correlation coefficients greater than or equal to a preset correlation threshold are identified as strongly cooperating SNP pairs. Based on these selected strongly cooperating SNP pairs, a cooperative network of HRR pathway SNP sites is constructed. Calculate the pathway association strength in the cooperative network of single nucleotide polymorphism sites in the HRR pathway; Based on the clinical association between single nucleotide polymorphism (SNP) genotype and breast cancer risk, a risk score was calculated for each SNP site. The overall risk value of the target module is calculated based on the functional importance weight of the target module, the strength of pathway association, and the risk score of each single nucleotide polymorphism site. Calculate the total risk value of breast cancer based on the comprehensive risk value of the target module; The total risk value of breast cancer in the training set is associated with the clinical label of breast cancer to construct a risk value-clinical label dataset; a logistic regression model is trained based on the training set and the risk value-clinical label dataset to obtain a risk model; the performance of the risk model is evaluated using a validation set; when the performance of the risk model meets the requirements, it is finally validated using a test set; the risk model is obtained when the validation is successful.

6. A personalized breast cancer risk detection method based on HRR pathway-related genes according to claim 5, characterized in that, The process combines genotyping with a constructed risk model to calculate the relative risk of breast cancer, including: Transform individual genotyping data into standardized variables that can be identified by risk models; Standardized variables were integrated with corresponding breast cancer clinical data to obtain an individual characteristic dataset. Individual characteristic datasets are input into the risk model, and the risk model determines the absolute risk value of breast cancer based on the input individual characteristic datasets; Once the absolute risk value for breast cancer is determined, a reference population is identified, and this reference population is used as the benchmark for the relative risk value. Finally, the individual relative risk is derived from the relative risk and the absolute risk value of breast cancer, using the following formula: Individual relative risk = absolute risk value of breast cancer ÷ relative risk value.

7. The personalized breast cancer risk detection method based on HRR pathway-related genes according to claim 6, characterized in that, Breast cancer risk levels are classified based on the relative risk calculation results, including: The classification logic is confirmed based on individual relative risk as the core indicator, combined with reference population benchmarks, clinical relevance, and existing clinical guidelines. Based on the confirmed grading logic, individual relative risk is classified into three levels: low risk, medium risk, and high risk. Low risk is defined as a relative risk value ≤ 1; medium risk is defined as a relative risk value > 1 and ≤ a preset threshold; and high risk is defined as a relative risk value > a medium risk threshold. Match individual relative risk with the classified individual relative risk level; After matching is completed, the final risk level data for breast cancer risk is obtained.

8. A personalized breast cancer risk detection method based on HRR pathway-related genes according to claim 7, characterized in that, Different levels of risk warnings are issued based on the classified breast cancer risk levels, including: The warning content is designed based on the breast cancer risk level, and includes risk interpretation, scientific basis, and action recommendations. After the warning content is designed, the warning method is confirmed. Low risk warnings are delivered in the form of written reports or online platform messages; medium risk warnings are delivered in the form of written reports, SMS reminders or telephone notifications, and it is recommended to check the report and contact a doctor as soon as possible; high risk warnings are delivered simultaneously through multiple channels, including written reports, emergency telephone notifications, and face-to-face explanations. Once the warning method is confirmed, corresponding interpretations are provided for individuals at different risk levels. For low-risk individuals, an online FAQ document is provided to answer questions about the warning content; for medium-risk individuals, a consultation hotline is opened, where professional nurses or doctors answer questions about screening details and risk changes; for high-risk individuals, a genetic counselor or breast specialist is arranged to provide one-on-one interpretation, explaining the risk mechanism, intervention options, and data on risk assessments of family members. The corresponding interpretations provide different levels of risk warnings.

9. A personalized breast cancer risk detection method based on HRR pathway-related genes according to claim 8, characterized in that, Intervention measures will be formulated based on the results of risk warnings, including: Based on the early warning results of different risk levels, the core objectives of the intervention were identified. The core objective for low-risk cases was to maintain a low-risk status; the core objective for medium-risk cases was to reduce the progression of the risk; and the core objective for high-risk cases was to proactively reduce the probability of disease onset. Intervention frameworks were developed based on the identified core objectives. The low-risk intervention framework included screening management, lifestyle guidance, and regular follow-up; the medium-risk intervention framework included intensive screening, intervention for controllable factors, clinical monitoring, and genetic counseling; and the high-risk intervention framework included intensive screening, risk blocking measures, pathway-targeted intervention, and family risk management. Once the intervention framework is established, individual-specific factors are incorporated and optimized in a personalized manner. These individual-specific factors include genetic, clinical, personal, and social factors. After personalized optimization, an implementation plan is developed for each intervention, including screening, medication, lifestyle, and surgical interventions; Once the implementation plan is finalized, the medical team reviews it. If the plan is approved, the final intervention plan will be obtained. If the plan is not approved, the plan will continue to be optimized and adjusted until it is approved. Simultaneously, basic individual information is obtained, and individual risk-related health behavior records for the past three years are extracted; risk scenarios are defined, and each risk scenario is uniquely assigned as a mutually exclusive low-risk, medium-risk, or high-risk scenario based on the risk warning result; the number of compliant health behaviors of an individual under their assigned risk scenario is counted; all individual data under risk warning are retained to obtain the first individual data; Obtain the intervention permission association table. When an individual's risk scenario does not match the intervention permission association table, delete the corresponding individual data that does not match and obtain the second individual data. Retrieve the preset first blank database; The first and second individual data are stored in the first blank database based on the format of individual ID-risk scenario-number of compliance behaviors-basic characteristics; Once all the filtered data has been entered, the first blank database will be used as the risk scenario binding individual database. Acquire historical intervention records and simultaneously collect corresponding clinical characteristic data, aligning the timestamps of both. Extract the entire process trajectory of the intervention behavior record. When the entire process trajectory is uninterrupted and the clinical characteristic data fluctuates within the preset normal range, the corresponding intervention behavior record is regarded as a valid intervention behavior record; otherwise, the corresponding record is removed. Retrieve the preset second blank database; Effective intervention behavior records are stored in a second blank database based on the format of behavior ID-individual ID-intervention trajectory-clinical characteristic curve-risk level; once all the filtered data has been entered, the second blank database is used as a clinical characteristic verification behavior library. Risk level-risk behavior association pairs are extracted from the clinical feature verification behavior database, and candidate strategies for each risk level-risk behavior combination are collected; the implementation cost of the candidate strategies is calculated; and the implementation benefits of the candidate strategies are calculated. Calculate the cost-benefit ratio based on the implementation costs and benefits of candidate strategies; The clinical feature data distribution of historical individuals is extracted from the clinical feature verification behavior database, and an appropriate feature range is established for each candidate strategy; candidate strategies whose appropriate feature range covers a preset proportion of historical individuals are preferentially retained. From the remaining candidate strategies, for each risk level-risk behavior combination, the strategy with the smallest cost-benefit ratio is selected as the optimal strategy; Iterate through all risk level-risk behavior combinations to obtain the optimal strategy for each combination; input each optimal strategy into the third blank database based on the format of risk level-risk behavior-optimal strategy-cost details-benefit score-adaptation feature range. Once all data has been entered, the third blank database will be used as a cost-benefit optimization strategy library. The clinical characteristic verification behavior database and cost-benefit optimization strategy database will be used as an intervention strategy database. Based on the risk warning results, the risk scenario is bound to the individual database to obtain the individual risk level and basic characteristics. Then, based on the individual risk level and basic characteristics, the intervention strategy database is queried to determine the appropriate intervention strategy.

10. A personalized breast cancer risk detection method based on HRR pathway-related genes according to claim 9, characterized in that, Real-time effectiveness tracking and evaluation are conducted based on the established intervention measures, and the intervention plan is adjusted according to the evaluation results, including: First, establish tracking indicators, including risk control indicators, compliance indicators, and safety indicators; The follow-up plan is then developed. The frequency of low-risk intervention follow-up is once every 6-12 months, through online questionnaires and annual report submissions, with the responsible parties being community doctors or health managers. The frequency of medium-risk intervention follow-up is once every 3-6 months, through offline follow-ups, screening results entering into the system, and telephone follow-ups, with the responsible parties being breast clinic nurses and specialists. The frequency of high-risk intervention follow-up is once a month to once every 3 months, through multidisciplinary team follow-ups and real-time updates of electronic health records, with the responsible parties being genetic counselors coordinating and multidisciplinary team collaboration. The intervention effect is tracked according to the tracking plan and tracking indicators. The intervention effect is comprehensively evaluated every 12-24 months, and the intervention is determined as to whether the target is achieved based on the evaluation results. The evaluation results include target achieved, partial target achieved, and target not achieved. The intervention plan will be adjusted according to the different assessment results. If the target is met, the existing intervention plan will be maintained and the follow-up interval will be extended. If the target is partially met, the unmet targets will be optimized. If the target is not met, the cause will be analyzed and the intervention plan will be reformulated. The final evaluation results and the revised plan will be entered into the system.

Citation Information

Cited By

  • Multi-level service point linkage gene screening system and method and storage medium

    CN121601258A

  • Multi-level service point linkage gene screening system, methods and storage media

    CN121601258B