Bioinformation comparative analysis method for gastric cancer pathogenic gene sequence

By constructing a multidimensional topological reference map of gastric cancer and a dynamic weighted multimodal graph comparison algorithm, the problem that linear reference genomes cannot identify pathogenic variants of gastric cancer was solved, achieving high sensitivity and high specificity of variant identification and improving the accuracy of clinical diagnosis.

CN121811972AActive Publication Date: 2026-04-07FUJIAN PROVINCIAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-05
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In the bioinformatics analysis of gastric cancer pathogenic gene sequences, existing technologies cannot effectively characterize nonlinear topological structures using linear reference genomes, and rigid penalty mechanisms cannot be dynamically adjusted, leading to missed and false detections of complex pathogenic variants.

Method used

A multidimensional topological reference map and mutation feature vector library for gastric cancer were constructed. The comparison model was dynamically adjusted based on etiological prior knowledge. A dynamic weighted multimodal graph comparison algorithm was adopted, combined with path optimization and Bayesian inference models, to identify high-confidence variants and update the topological reference map.

Benefits of technology

It significantly improved the sensitivity and specificity of identifying complex pathogenic variants, reduced the false positive rate, and enhanced the adaptability of the comparison model and the reliability of clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811972A_ABST
    Figure CN121811972A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bioinformatics and genome sequencing data processing, in particular to a gastric cancer pathogenic gene sequence biological information comparative analysis method, which comprises the following steps: acquiring gene sequencing data of a target sample, loading a preset gastric cancer multi-dimensional topological reference map, and loading a preset mutation feature vector library; generating characteristic frequency distribution, and matching the characteristic frequency distribution with a preset etiological fingerprint spectrum to determine a pathogenic background type of the target sample; constructing a dynamic weight multi-modal diagram comparison model; calculating the fitting degree of the optimal comparison path and a theoretical variation path based on etiological deduction, and generating a pathogenic confidence score; generating a corrected second topological reference graph; according to the method, the limitation of linear comparison in processing complex structure variation is effectively overcome, and the variation recognition sensitivity is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bioinformatics and genome sequencing data processing technology, specifically a bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences. Background Technology

[0002] In the bioinformatics analysis of gastric cancer pathogenic gene sequences, the genome of the target sample usually has high-frequency structural variation features and carries complex etiological imprints left by specific pathogenic factors such as Helicobacter pylori, EB virus infection or microsatellite instability. For comparative analysis of this type of data, existing schemes generally use a standard linear reference genome as the comparison benchmark and rely on a pre-set fixed penalty matrix to perform physical string matching. Although this linear comparison strategy is relatively mature in handling routine germline variations, it has significant limitations when dealing with complex cancer genomes such as gastric cancer due to its lack of perception of the specific biological background of the sample. Specifically, the linear reference system cannot effectively represent nonlinear topological structures such as inversions and translocations, leading to the failure of the reference system. At the same time, the rigid penalty mechanism cannot dynamically adjust the weights according to specific pathological logics such as oxidative damage or repair defects, making it difficult for the algorithm to accurately identify low-abundance mutations that conform to specific etiological patterns in sequencing noise, resulting in missed detections and false detections under high background noise. Therefore, how to construct a nonlinear dynamic comparison model that can integrate etiological prior knowledge and improve the sensitivity and specificity of identifying complex pathogenic variations by adaptively adjusting the topological structure and comparison parameters has become an urgent technical problem to be solved. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention provides a bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences. Specifically, the technical solution of this invention includes: The bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences includes the following steps: Step 1: Obtain gene sequencing data of the target sample, load the preset gastric cancer multidimensional topology reference map, and load the preset mutation feature vector library; wherein, the gastric cancer multidimensional topology reference map is a non-linear network data structure containing bifurcations and loops, nodes represent gene sequence fragments, and edges represent connection relationships or mutation jump paths; the mutation feature vector library contains position-specific penalty matrices corresponding to different pathogenic factors; Step 2: Perform feature scanning on gene sequencing data to generate feature frequency distribution, and match it with the preset etiological fingerprint map to determine the pathogenic background type of the target sample; based on the pathogenic background type, call the matching penalty matrix from the mutation feature vector library to dynamically adjust the edge weights and node transfer probabilities in the gastric cancer multidimensional topological reference graph, and construct a dynamic weighted multimodal graph comparison model. Step 3: Input the gene sequencing data into the dynamic weighted multimodal graph alignment model, use the path optimization algorithm to obtain the optimal alignment path, and calculate the degree of fit between the optimal alignment path and the theoretical variation path based on etiology to generate a pathogenicity confidence score. Step 4: Identify the list of structural variations based on the pathogenicity confidence score, and extract the unrecorded structural variations with high confidence from the list of structural variations. Use these as new topological nodes or edges to update the gastric cancer multidimensional topological reference graph in reverse, so as to generate the corrected second topological reference graph.

[0004] Optionally, step one includes: S11. Collect raw sequencing read data of the target sample using high-throughput sequencing equipment; S12. Construct a multidimensional topological reference graph of gastric cancer and configure it as a directed graph structure that integrates known gastric cancer structural variation breakpoint information. S13. Construct a mutation feature vector library to store base mismatch penalty value vectors and vacancy penalty vectors for different etiological scenarios such as Helicobacter pylori infection, EB virus infection, chemotherapy drug induction, and microsatellite instability.

[0005] Optionally, the determination of the pathogenic background type in step two includes: S21, Adopt Frequency analysis algorithms process raw sequencing read data to generate characteristic frequency distribution data; S22. Calculate the similarity between the characteristic frequency distribution data and each preset type in the etiological fingerprint, and determine the type with the highest similarity as the pathogenic background type; the pathogenic background type includes at least microsatellite instability, virus-positive type, chromosome instability, and chemotherapy drug-induced type; S23. In response to the determined pathogenic background type, generate a weight adjustment instruction. The weight adjustment instruction is used to define a reduction or increase of the comparison penalty in a specific coordinate region of the gastric cancer multidimensional topological reference map.

[0006] Optionally, the dynamic adjustment process in step two includes: S24. If the pathogenic background type is determined to be microsatellite instability, then retrieve the corresponding feature vector, reduce the insertion and deletion penalty of the homopolymer region in the gastric cancer multidimensional topological reference map, and generate the first dynamic weight matrix. S25. If the pathogenic background type is determined to have oxidative damage characteristics, the corresponding feature vector is retrieved, and the mismatch penalty of specific base transversion is reduced in the gastric cancer multidimensional topological reference map to generate a second dynamic weight matrix. S26. Assign the values ​​of the generated first dynamic weight matrix or second dynamic weight matrix to the corresponding edges of the gastric cancer multidimensional topological reference graph to complete the parameter configuration of the dynamic weighted multimodal graph comparison model.

[0007] Optionally, step three also includes a classification evaluation of the optimal alignment path: S31. Preset a first confidence threshold and a second confidence threshold, wherein the first confidence threshold is less than the second confidence threshold; S32. Compare the pathogenicity confidence score with the first confidence threshold and the second confidence threshold, and execute the following judgment logic: If the pathogenicity confidence score is less than the first confidence threshold, the optimal alignment path is determined to be sequencing noise or background interference and is marked as invalid alignment. If the pathogenicity confidence score is greater than or equal to the first confidence threshold and less than the second confidence threshold, the optimal alignment path is determined to be a suspected variant and marked as a region to be verified. If the pathogenicity confidence score is greater than or equal to the second confidence threshold, the optimal alignment path is determined to be a high-confidence pathogenic mutation and marked as a valid variant.

[0008] Optionally, step three also includes: S33. For paths marked as valid mutations, parse the topological structure data; S34. Identify single nucleotide polymorphisms, insertions, deletions, and fusion gene breakpoints in the optimal alignment path; S35. Summarize all identified variant types and construct a structural variant list; the structural variant list includes the genomic location coordinates of the variant, the variant type, and the corresponding etiological association label.

[0009] Optionally, step four includes: S41. Traverse the list of structural variations and filter out novel structural variations that do not exist in the initial structure of the gastric cancer multidimensional topological reference map. S42. Extract the sequence information of the novel structural variation as a new node, and extract the connection relationship of the novel structural variation as a new edge; S43. Embed the new nodes and edges into the gastric cancer multidimensional topological reference graph to form a modified second topological reference graph, so that the subsequent comparison process can directly use the new edges as a shortcut path.

[0010] Optionally, step four also includes a report generation step: S44. Based on the list of structural variations, generate a visual gene variation report containing variation sites and pathogenicity risk levels; S45. If the pathogenic background type is virus-positive, then the virus sequence integration site information will be output in the report. S46. If the pathogenic background type is microsatellite instability, then the report will output a prompt message indicating related gene repair defects.

[0011] Optionally, the method may also include processing steps for regions marked as to be verified: S36. Extract the local sequence of the region to be verified, and use the Bayesian inference model in combination with the pathogenic background type to perform a secondary evaluation to obtain the corrected pathogenicity probability. S37. If the correction probability of pathogenicity is greater than the preset confirmation threshold, the region to be verified will be upgraded and marked as a valid mutation; if the correction probability of pathogenicity is less than or equal to the preset confirmation threshold, the region to be verified will be downgraded and marked as an invalid alignment.

[0012] Optionally, the embedding process of S43 specifically includes: S431. Perform cluster analysis on novel structural variations, count their frequency of occurrence in a pre-set historical population sample database, and generate a frequency ranking table. S432. If the frequency of occurrence of novel structural variations is higher than the preset frequency threshold, they are updated as general atlas nodes to the base layer of the gastric cancer multidimensional topological reference graph as permanent reference paths. S433. If the frequency of occurrence of novel structural variations is less than or equal to a preset frequency threshold, they are stored as specific map nodes in a temporary variation layer and configured to be activated and loaded only when samples of the same pathogenic background type are detected.

[0013] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention replaces the traditional linear reference genome by constructing a nonlinear network data structure containing bifurcation and loops. This topology map integrates known gastric cancer-specific inversions, translocations, and large fragment insertion / deletion breakpoint information, which can accurately characterize the nonlinear structure of the genome. This allows the alignment algorithm to directly search using a preset shortcut path, effectively overcoming the limitations of linear alignment in handling complex structural variations and significantly improving the sensitivity of variation identification. 2. This invention utilizes etiological fingerprinting to identify the pathogenic background of samples and dynamically adjusts the penalty weight of the alignment model accordingly. For specific scenarios such as microsatellite instability or oxidative damage, it adaptively reduces the penalty for insertions, deletions, or specific base mismatches in the corresponding regions. This biological background-aware alignment strategy breaks the rigid mode of traditional fixed penalty and can significantly improve the detection rate of low-abundance mutations that conform to specific pathological logic while suppressing random sequencing noise. 3. This invention establishes a graded evaluation mechanism based on pathogenicity confidence scores and introduces a Bayesian inference model for secondary evaluation of the region to be validated. By quantifying the significance of the alignment path being superior to the random background and combining the prior probability of pathogenicity to perform probability correction on the region to be validated, the algorithm can effectively distinguish between sequencing noise and recessive mutations. This strategy minimizes false positives while uncovering weak pathogenic signals, ensuring the reliability of clinical diagnosis. 4. This invention designs a hierarchical embedding strategy that can automatically extract novel structural variants with high confidence and update them in reverse to the reference map. The system updates the base layer as a general path or stores it in the temporary layer as a specific path according to the frequency of occurrence of the variant in the population. This mechanism enables the alignment model to have continuous evolution capabilities, and continuously enhances the ability to directly capture unknown variants as sample data accumulates, taking into account both the universality of the map and the retrieval efficiency. Attached Figure Description

[0014] The present invention will be further explained below with reference to the accompanying drawings and embodiments: Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments. Example 1:

[0016] Please see Figure 1 The bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences includes the following steps: Step 1: Obtain gene sequencing data of the target sample, load the preset gastric cancer multidimensional topology reference map, and load the preset mutation feature vector library; wherein, the gastric cancer multidimensional topology reference map is a non-linear network data structure containing bifurcations and loops, nodes represent gene sequence fragments, and edges represent connection relationships or mutation jump paths; the mutation feature vector library contains position-specific penalty matrices corresponding to different pathogenic factors; Step 2: Perform feature scanning on gene sequencing data to generate feature frequency distribution, and match it with the preset etiological fingerprint map to determine the pathogenic background type of the target sample; based on the pathogenic background type, call the matching penalty matrix from the mutation feature vector library to dynamically adjust the edge weights and node transfer probabilities in the gastric cancer multidimensional topological reference graph, and construct a dynamic weighted multimodal graph comparison model. Step 3: Input the gene sequencing data into the dynamic weighted multimodal graph alignment model, use the path optimization algorithm to obtain the optimal alignment path, and calculate the degree of fit between the optimal alignment path and the theoretical variation path based on etiology to generate a pathogenicity confidence score. Step 4: Identify the list of structural variations based on the pathogenicity confidence score, and extract the unrecorded structural variations with high confidence from the list of structural variations. Use these as new topological nodes or edges to update the gastric cancer multidimensional topological reference graph in reverse, so as to generate the corrected second topological reference graph.

[0017] This embodiment discloses a bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences, aiming to solve the problems of missed detection and false detection caused by reference system failure and rigid penalty points when the existing linear reference genome alignment scheme is used to process the genome of gastric cancer, which has high frequency of structural variations and complex etiological imprints. The core logic of this embodiment lies in constructing a dynamic weighted multimodal graph alignment model; this model no longer treats alignment as a simple physical string matching, but elevates it to a biologically context-aware sequence reconstruction process; the specific steps are as follows: Step 1: Initialize the multidimensional topology environment: The system acquires the gene sequencing data of the target sample. The system loads a preset multidimensional topological reference map of gastric cancer. and mutation feature vector library; The so-called gastric cancer multidimensional topological reference graph refers to a non-linear network data structure that differs from standard linear human genomes such as GRCh38. It integrates known gastric cancer-specific structural variation breakpoint information. In this graph, nodes represent gene sequence segments, and edges represent connection relationships or non-regular connections between fused genes and other variation jump paths. The so-called mutation feature vector library refers to a location-specific penalty matrix containing different pathogenic factors such as Helicobacter pylori and Epstein-Barr virus. The database serves to provide prior etiological knowledge for subsequent comparison processes. Step 2: Construct a dynamic weighted multimodal graph alignment model: The system... Feature scanning is performed to generate a feature frequency distribution, which is then matched with a pre-defined etiological fingerprint to determine the pathogenic background type of the target sample. ; Based on certainty The system retrieves the matching penalty matrix from the library. ,right The edge weights and node transition probabilities in the graph are dynamically adjusted; this process transforms the general graph alignment model into a dynamic weighted multimodal graph alignment model for this specific sample. Step 3, Path Optimization and Score Generation: ... The improved version is input into the dynamic model described above. The algorithm employs a partial-order alignment recursive logic based on the adaptation graph topology, which differs from traditional linear sequence alignment. Its score matrix... The update formula is: ; in, Represents all points to nodes in a directed graph structure. The set of predecessor nodes; This refers to the edge weights or node matching scores after dynamic adjustment in step two. A penalty is applied to the preset vacancy, typically a negative constant, such as -10, or dynamically configured based on the pathogenic background type determined in step two; for example, the absolute value is reduced in an MSI background. Furthermore... Specifically refers to all nodes that directly point to each other in a directed graph structure. The set of predecessor nodes, i.e., the existence of directed edges. This formula ensures that the alignment path can perform an optimal search along nonlinear shortcuts in the graph structure, such as structural variation edges; the system utilizes the above-mentioned improvements. The algorithm obtains the score matrix. The largest optimal alignment path is identified, and a pathogenicity confidence score is calculated based on the score of that path. The calculation formula is as follows: ; in, The cumulative score is the score for the optimal alignment path. To determine the read length of the input sequencing data, and These are the preset average score coefficient and standard deviation coefficient per unit length in the context of a random sequence, respectively, set for example based on the parameters of the Gumbel distribution; this score Essentially a Standardized values ​​are used to quantify the extent to which the current alignment path is superior to random background noise, thereby characterizing its fit with the etiological deduction logic. Step 4, Self-evolutionary Update of the Map: Based on Identify the list of structural variations, extract the high-confidence but unrecorded structural variations, and use them as new topological nodes or edges to update the topology. In the process, a corrected second topology reference map is generated; This embodiment achieves a paradigm shift from searching for differences from the reference genome to searching for variants that conform to pathogenic logic by introducing prior knowledge of etiology and dynamically adjusting the alignment parameters. That is, the system no longer presets a single standard answer, but regards the optimal path in the graph under specific etiologies, such as oxidative damage weighting, as the theoretical variant path deduced by the etiology model. This design enables the algorithm to tolerate variants that conform to specific etiological logic when faced with complex structural variants commonly found in gastric cancer samples, while maintaining suppression of sequencing noise, thereby significantly improving the detection rate and accuracy of pathogenic mutations in low-purity samples. Example 2:

[0018] Step one includes: S11. Collect raw sequencing read data of the target sample using high-throughput sequencing equipment; S12. Construct a multidimensional topological reference graph of gastric cancer and configure it as a directed graph structure that integrates known gastric cancer structural variation breakpoint information. S13. Construct a mutation feature vector library to store base mismatch penalty value vectors and vacancy penalty vectors for different etiological scenarios such as Helicobacter pylori infection, EB virus infection, chemotherapy drug induction, and microsatellite instability.

[0019] This embodiment specifies the data acquisition and basic library construction process in step one; S11, Raw sequencing read data acquisition: via Or high-throughput sequencing equipment such as the MGI platform to collect raw sequencing read data of the target sample; S12. Construction of a multidimensional topological reference map for gastric cancer: In this embodiment, The system is configured as a directed graph structure that incorporates known breakpoint information on gastric cancer structural variations. Specifically, it integrates gastric cancer-specific inversions, translocations, and large fragment insertion / deletion breakpoints from databases such as TCGA and ICGC, and establishes shortcut edges in the graph. Simultaneously, the system loads annotation files from the reference genome, such as BED or GFF format, to annotate microsatellite homopolymer regions in the genome. The coordinates of high-density island areas and known viral integration hotspots are mapped to graph node attributes; Each node Includes attribute set ,in, An enumerated variable is used, with specific values ​​including HOMOPOLYMER, CPG_ISLAND, VIRUS_HOTSPOT, and DEFAULT; this attribute is used to respond to the weight adjustment instruction in step three to identify whether it belongs to a specific coordinate region that needs weight adjustment. S13. Construction of the mutation feature vector library: The system stores feature vectors for different etiological scenarios; Base mismatch penalty value vector: For specific oxidative damage caused by Helicobacter pylori H. pylori infection, a vector is stored to reduce the penalty for specific base mismatches such as C>A. Vacancy penalty vector: For microsatellite instability (MSI) scenarios, a vector is stored to reduce the insertion or missing penalty in the homopolymer region; By introducing targeted topological structures and feature vectors during the initialization phase, this embodiment lays the material foundation for subsequent accurate alignment, ensuring that the algorithm can identify complex variation paths that do not exist in standard linear genomes. Example 3:

[0020] Step two, determining the pathogenic background type, includes: S21, Adopt Frequency analysis algorithms process raw sequencing read data to generate characteristic frequency distribution data; S22. Calculate the similarity between the characteristic frequency distribution data and each preset type in the etiological fingerprint, and determine the type with the highest similarity as the pathogenic background type; the pathogenic background type includes at least microsatellite instability, virus-positive type, chromosome instability, and chemotherapy drug-induced type; S23. In response to the determined pathogenic background type, generate a weight adjustment instruction. The weight adjustment instruction is used to define a reduction or increase of the comparison penalty in a specific coordinate region of the gastric cancer multidimensional topological reference map.

[0021] This embodiment describes in detail the method for determining the pathogenic background type in step two; S21, Frequency analysis: The system adopts The frequency analysis algorithm processes the raw sequencing read data to generate characteristic frequency distribution data; the generated characteristic frequency distribution vector is centered, that is, the frequency of each dimension is subtracted from the genome-wide average frequency of that dimension, so that the characteristic vector contains bias information in the positive and negative directions, in order to eliminate the influence of the background baseline. S22. Similarity Calculation and Type Determination: To quantify the association between samples and known etiological patterns, this embodiment introduces an etiological similarity coefficient. The calculation formula is as follows: ; in, The total number of dimensions of the feature frequency distribution vector, i.e. The number of feature types; : Specific trinucleotide mutation patterns in the target sample, etc. The observation frequency of this feature, whose value is derived from S21. Analysis results; Such as MSI feature maps, etc. The standard frequencies of corresponding features in a pre-defined etiological fingerprint spectrum are derived from a pre-defined database. : Range of values , used to characterize the degree of matching between a sample and a specific etiological type; The system calculates the similarity between the characteristic frequency distribution data and at least several preset types, including microsatellite instability, viral positive, and chromosomal instability, and then... The highest type was identified as the pathogenic background type. ; S23. Generate weight adjustment instructions: in response to a given... The system generates a weight adjustment instruction; this instruction is used to define weights in microsatellite duplication zones or virus integration hotspots, etc. The penalty points are reduced or increased in specific coordinate areas compared to the penalty points. By performing a panoramic inference of etiological background before comparison, this embodiment can provide macro-level guidance for subsequent detailed comparison, avoid blind search, greatly reduce computational space and improve the specificity of comparison. Example 4:

[0022] Step two, the dynamic adjustment process, includes: S24. If the pathogenic background type is determined to be microsatellite instability, then retrieve the corresponding feature vector, reduce the insertion and deletion penalty of the homopolymer region in the gastric cancer multidimensional topological reference map, and generate the first dynamic weight matrix. S25. If the pathogenic background type is determined to have oxidative damage characteristics, the corresponding feature vector is retrieved, and the mismatch penalty of specific base transversion is reduced in the gastric cancer multidimensional topological reference map to generate a second dynamic weight matrix. S26. Assign the values ​​of the generated first dynamic weight matrix or second dynamic weight matrix to the corresponding edges of the gastric cancer multidimensional topological reference graph to complete the parameter configuration of the dynamic weighted multimodal graph comparison model. This embodiment further clarifies the dynamic adjustment process in step two, demonstrating how etiological knowledge can be transformed into parameters of a mathematical model. S24. Dynamic adjustment of microsatellite unstable MSI: If Once identified as a microsatellite unstable type, the system retrieves the corresponding feature vector and generates the first dynamic weight matrix. In this scenario, to accurately identify slip mutations caused by DNA mismatch repair defects, this embodiment introduces an adaptive vacancy penalty function. The calculation formula is as follows: ; in, The system's default general empty space opening penalty is a constant. Relaxation coefficient, range of values This is used to control the extent of the penalty point reduction, and in this embodiment, it is dynamically set according to the MSI level; specifically, The etiological similarity coefficient calculated in step two They are positively correlated, and the calculation formula is: ; in, The preset response gain constant is 1.2 in this embodiment; this formula ensures that when the sample does not exhibit MSI characteristics, The relaxation coefficient is 0, meaning the penalty score is not reduced; when the MSI characteristic is extremely strong, the penalty score reduction is limited to within 90%. The location indication function, whose judgment logic is related to the node attributes defined in Example 2, determines the location when comparing positions. Corresponding node attributes The value is 1 when it is marked as HOMOPOLYMER, and 0 otherwise. The purpose of this formula is to significantly reduce the insertion / deletion penalty in the homopolymer region, thereby allowing the algorithm to tolerate length variations in this region and generate the first dynamic weight matrix. S25. Dynamic adjustment of oxidative damage characteristics: If The system is identified as having oxidative damage characteristics, such as an environment infected with Helicobacter pylori, and generates a second dynamic weight matrix. In this scenario, for a specific C>A transversion, the system reduces its mismatch penalty. The specific reduction is achieved through a matrix element update formula: Let the mismatch penalty of C>A in the original penalty matrix be... , usually a negative value, the updated penalty score The calculation is as follows: ; in, This is the oxidative damage correction factor, with a default value of 0.5; since the penalty is negative, it is multiplied by... This reduces the absolute value of the error, which in turn increases the tolerance for this type of error in the algorithm, i.e., reduces the penalty, in order to match the DNA chemical modification features caused by oxidative damage. S26. Parameter Configuration: Assign values ​​to the generated first or second dynamic weight matrix. The corresponding edges are used to complete the parameter configuration of the dynamic weighted multimodal graph comparison model; This embodiment makes the alignment algorithm flexible by adjusting the penalty score mathematically. While maintaining a high penalty score for random sequencing errors, it specifically lowers the identification threshold for variants that conform to specific etiological logic, thereby effectively solving the technical problem of difficult detection of low-abundance pathogenic mutations under high background noise. Example 5:

[0023] Step three also includes a classification evaluation of the optimal alignment path: S31. Preset a first confidence threshold and a second confidence threshold, wherein the first confidence threshold is less than the second confidence threshold; S32. Compare the pathogenicity confidence score with the first confidence threshold and the second confidence threshold, and execute the following judgment logic: If the pathogenicity confidence score is less than the first confidence threshold, the optimal alignment path is determined to be sequencing noise or background interference and is marked as invalid alignment. If the pathogenicity confidence score is greater than or equal to the first confidence threshold and less than the second confidence threshold, the optimal alignment path is determined to be a suspected variant and marked as a region to be verified. If the pathogenicity confidence score is greater than or equal to the second confidence threshold, the optimal alignment path is determined to be a high-confidence pathogenic mutation and marked as a valid variant; The method also includes a processing step for regions marked as to be verified: S36. Extract the local sequence of the region to be verified, and use the Bayesian inference model in combination with the pathogenic background type to perform a secondary evaluation to obtain the corrected pathogenicity probability. S37. If the correction probability of pathogenicity is greater than the preset confirmation threshold, the region to be verified will be upgraded and marked as a valid mutation; if the correction probability of pathogenicity is less than or equal to the preset confirmation threshold, the region to be verified will be downgraded and marked as an invalid alignment.

[0024] This embodiment describes in detail the classification and evaluation of the optimal alignment path and the secondary confirmation mechanism of the region to be verified in step three; S31-S32, Threshold-based primary classification Prior to this step, the system pre-determines the threshold through ROC curve analysis. Specifically, it is pre-trained using a standard dataset with known pathogenic mutations and a healthy dataset with known negative mutations, selecting a threshold where the false positive rate (FPR) is below 0.1%. The value is used as the second confidence threshold. Select when FPR is less than 5% The value is used as the first confidence threshold. The system presets a first confidence threshold. Second confidence threshold ,in As a specific example of this embodiment, based on Based on the training set data of the sequencing platform at a sequencing depth of 30X, the specific values ​​of the above thresholds were determined as follows: First confidence threshold Second confidence threshold ; Pathogenicity confidence score Compare with these two thresholds: like If the result is determined to be sequencing noise or background interference, it is marked as an invalid alignment. like These were identified as high-confidence pathogenic mutations and marked as valid variants. like This was identified as a suspected variant and marked as a region to be verified. S36. Bayesian secondary evaluation of the region to be verified For the regions to be verified that are in the grayscale zone, this embodiment introduces a Bayesian modified pathogenicity model to calculate the modified pathogenicity probability. The calculation formula is as follows: ; in, : Likelihood probability based on score transformation; since the mutation feature vector library stores the penalty values ​​in logarithmic form. The system transforms it into a probability space using the Boltzmann distribution formula: ; in, The scaling parameter is used to calibrate the dynamic range of the penalty scoring system. In this embodiment... ; Let be the partition function, and its calculation formula is: In actual calculations, the sum of the score indices of all candidate alignment paths within a local area can be used as an approximation; this step realizes the mathematical mapping from physical alignment scores to statistical probabilities. The prior probability of a pathogenic mutation occurring at this location is derived from statistical frequencies in databases such as COSMIC. The marginal probability of local data is calculated using the following formula: ; in, This represents the assumption of a non-pathogenic random background; and This represents the prior probability that a local region belongs to a non-pathogenic background. S37, Secondary Grading: If If the value exceeds the preset confirmation threshold, the region to be verified will be upgraded and marked as a valid mutation; otherwise, it will be downgraded to an invalid alignment. This embodiment employs a two-tiered evaluation mechanism of initial screening and refinement; by utilizing Bayesian inference combined with prior knowledge of pathogenic background, it can effectively uncover latent mutations with weak signals but consistent with etiological logic, thereby minimizing the false negative rate. Example 6:

[0025] Step three also includes: S33. For paths marked as valid mutations, parse the topological structure data; S34. Identify single nucleotide polymorphisms, insertions, deletions, and fusion gene breakpoints in the optimal alignment path; S35. Summarize all identified variant types and construct a structural variant list; the structural variant list includes the genomic location coordinates of the variant, the variant type, and the corresponding etiological association label.

[0026] This embodiment details the process of analyzing effective variants; S33-S35, Construction of the Structural Variation List For paths marked as valid mutations, the system parses their topological structure in the graph model; Single nucleotide polymorphism : Identified as a base substitution on a graph node; Insertion Missing : Identifies nodes skipped or additional nodes inserted in the graph path; Fusion gene breakpoints: identified as anomalous edges in the graph that cross different chromosomes or long-distance genome coordinates; The system will summarize all identified variant types and construct a list of structural variants. This list includes not only the genomic location coordinates and variant type of the variant, but also the corresponding etiological association tags, such as the C>A mutation associated with H. pylori. By outputting a list of structural variations tagged with etiological information, this embodiment not only provides information on where the variations are located, but also provides clues as to what might cause them, thus offering a richer dimension for clinical diagnosis. Example 7:

[0027] Step four includes: S41. Traverse the list of structural variations and filter out novel structural variations that do not exist in the initial structure of the gastric cancer multidimensional topological reference map. S42. Extract the sequence information of the novel structural variation as a new node, and extract the connection relationship of the novel structural variation as a new edge; S43. Embed the new nodes and new edges into the gastric cancer multidimensional topological reference graph to form a modified second topological reference graph, so that the subsequent comparison process can directly use the new edges as a shortcut path. The embedding process of S43 specifically includes: S431. Perform cluster analysis on novel structural variations, count their frequency of occurrence in a pre-set historical population sample database, and generate a frequency ranking table. S432. If the frequency of occurrence of novel structural variations is higher than the preset frequency threshold, they are updated as general atlas nodes to the base layer of the gastric cancer multidimensional topological reference graph as permanent reference paths. S433. If the frequency of occurrence of novel structural variations is less than or equal to a preset frequency threshold, they are stored as specific map nodes in a temporary variation layer and configured to be activated and loaded only when samples of the same pathogenic background type are detected.

[0028] This embodiment elaborates on the system's self-evolutionary update mechanism, especially the embedding strategy for new nodes; S41. Novelty Screening: Traverse the list of structural variations and filter out those not present in the original structure. Novel structural variations in the initial structure; S42. Information Extraction: Extract sequence information of novel structural variations as new nodes, and extract their connection relationships as new edges; S43. Hierarchical Embedding Strategy: To balance the universality and specificity of the map, this embodiment adopts a hierarchical embedding mechanism: S431. Frequency Statistics: Perform cluster analysis on novel structural variations and count their frequency of occurrence in a pre-defined historical population sample database. Generate a frequency sorting table; S432, General Layer Update: If If the frequency exceeds a preset threshold, for example, if it appears in more than 5% of gastric cancer samples, it indicates that the variant is prevalent; the system updates it as a general atlas node. In the base layer, it serves as a permanent reference path; this means that all subsequent sample alignments will refer to this path. S433, Specific Layer Update: If If the frequency is less than or equal to a preset frequency threshold, it indicates that the mutation is rare or individual-specific; the system stores it as a specific map node in the temporary mutation layer. This temporary variant layer is configured to be activated and loaded into the alignment map only when a new sample of the same pathogenic background type as the sample from which the variant originates is detected, such as being EBV positive. The layered embedding strategy in this embodiment reflects the intelligence of the system; it not only continuously expands the ability to directly capture high-frequency variations by updating the base layer, but also avoids the map from expanding infinitely due to the introduction of too many rare variations through the temporary variation layer, thus ensuring a balance between comparison efficiency and map size. Example 8:

[0029] Step four also includes the report generation step: S44. Based on the list of structural variations, generate a visual gene variation report containing variation sites and pathogenicity risk levels; S45. If the pathogenic background type is virus-positive, then the virus sequence integration site information will be output in the report. S46. If the pathogenic background type is microsatellite instability, then the report will output a prompt message indicating related gene repair defects.

[0030] This embodiment describes the steps for generating the final report; S44-S46 Intelligent Report Generation Generate a visual gene variation report based on the list of structural variations; If the pathogenic background type is a virus-positive type such as EBV+, the report will automatically link and output the integration site information of the virus sequence in the host genome; If the pathogenic background type is microsatellite instability (MSI), the report will output related information on defects in gene repair such as MLH1 and MSH2 to assist in clinical medication decisions such as the use of immune checkpoint inhibitors. The report generated in this embodiment goes beyond the scope of traditional variant lists, directly providing etiological explanations and accompanying diagnostic information closely related to clinical treatment, greatly enhancing the clinical application value of the test results.

[0031] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences, characterized in that, The specific steps include: Step 1: Obtain gene sequencing data of the target sample, load the preset gastric cancer multidimensional topology reference graph, and load the preset mutation feature vector library; wherein, the gastric cancer multidimensional topology reference graph is a non-linear network data structure containing bifurcations and loops, nodes represent gene sequence fragments, and edges represent connection relationships or mutation jump paths; the mutation feature vector library contains position-specific penalty matrices corresponding to different pathogenic factors; Step 2: Perform feature scanning on gene sequencing data to generate feature frequency distribution, and match it with the preset etiological fingerprint map to determine the pathogenic background type of the target sample; based on the pathogenic background type, call the matching penalty matrix from the mutation feature vector library to dynamically adjust the edge weights and node transfer probabilities in the gastric cancer multidimensional topological reference graph, and construct a dynamic weighted multimodal graph comparison model. Step 3: Input the gene sequencing data into the dynamic weighted multimodal graph alignment model, use the path optimization algorithm to obtain the optimal alignment path, and calculate the degree of fit between the optimal alignment path and the theoretical variation path based on etiological inference to generate a pathogenicity confidence score. Step 4: Identify the list of structural variations based on the pathogenicity confidence score, and extract the unrecorded structural variations with high confidence from the list of structural variations. Use these as new topological nodes or edges to update the gastric cancer multidimensional topological reference graph in reverse, so as to generate the corrected second topological reference graph.

2. The bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences according to claim 1, characterized in that: Step one includes: S11. Collect raw sequencing read data of the target sample using high-throughput sequencing equipment; S12. Construct a multidimensional topological reference graph of gastric cancer and configure it as a directed graph structure that integrates known gastric cancer structural variation breakpoint information. S13. Construct a mutation feature vector library to store base mismatch penalty value vectors and vacancy penalty vectors for different etiological scenarios such as Helicobacter pylori infection, EB virus infection, chemotherapy drug induction, and microsatellite instability.

3. The bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences according to claim 2, characterized in that: Step two, determining the pathogenic background type, includes: S21, Adopt Frequency analysis algorithms process raw sequencing read data to generate characteristic frequency distribution data; S22. Calculate the similarity between the characteristic frequency distribution data and each preset type in the etiological fingerprint, and determine the type with the highest similarity as the pathogenic background type; the pathogenic background type includes at least microsatellite instability, virus-positive type, chromosome instability, and chemotherapy drug-induced type; S23. In response to the determined pathogenic background type, generate a weight adjustment instruction. The weight adjustment instruction is used to define a reduction or increase of the comparison penalty in a specific coordinate region of the gastric cancer multidimensional topological reference map.

4. The bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences according to claim 3, characterized in that: Step two, the dynamic adjustment process, includes: S24. If the pathogenic background type is determined to be microsatellite instability, then retrieve the corresponding feature vector, reduce the insertion and deletion penalty of the homopolymer region in the gastric cancer multidimensional topological reference map, and generate the first dynamic weight matrix. S25. If the pathogenic background type is determined to have oxidative damage characteristics, then retrieve the corresponding feature vector, reduce the mismatch penalty of specific base transversion in the gastric cancer multidimensional topological reference map, and generate the second dynamic weight matrix. S26. Assign the values ​​of the generated first dynamic weight matrix or second dynamic weight matrix to the corresponding edges of the gastric cancer multidimensional topological reference graph to complete the parameter configuration of the dynamic weighted multimodal graph comparison model.

5. The bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences according to claim 1, characterized in that: Step three also includes a classification evaluation of the optimal alignment path: S31. Preset a first confidence threshold and a second confidence threshold, wherein the first confidence threshold is less than the second confidence threshold; S32. Compare the pathogenicity confidence score with the first confidence threshold and the second confidence threshold, and execute the following judgment logic: If the pathogenicity confidence score is less than the first confidence threshold, the optimal alignment path is determined to be sequencing noise or background interference and is marked as invalid alignment. If the pathogenicity confidence score is greater than or equal to the first confidence threshold and less than the second confidence threshold, the optimal alignment path is determined to be a suspected variant and marked as a region to be verified. If the pathogenicity confidence score is greater than or equal to the second confidence threshold, the optimal alignment path is determined to be a high-confidence pathogenic mutation and marked as a valid variant.

6. The bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences according to claim 5, characterized in that: Step three also includes: S33. For paths marked as valid mutations, parse the topological structure data; S34. Identify single nucleotide polymorphisms, insertions, deletions, and fusion gene breakpoints in the optimal alignment path; S35. Summarize all identified variant types and construct a structural variant list; the structural variant list includes the genomic location coordinates of the variant, the variant type, and the corresponding etiological association label.

7. The bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences according to claim 6, characterized in that: Step four includes: S41. Traverse the list of structural variations and filter out novel structural variations that do not exist in the initial structure of the gastric cancer multidimensional topological reference map. S42. Extract the sequence information of the novel structural variation as a new node, and extract the connection relationship of the novel structural variation as a new edge; S43. Embed the new nodes and edges into the gastric cancer multidimensional topological reference graph to form a modified second topological reference graph, so that the subsequent comparison process can directly use the new edges as a shortcut path.

8. The bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences according to claim 7, characterized in that: Step four also includes the report generation step: S44. Based on the list of structural variations, generate a visual gene variation report containing variation sites and pathogenicity risk levels; S45. If the pathogenic background type is virus-positive, then the virus sequence integration site information will be output in the report. S46. If the pathogenic background type is microsatellite instability, then the report will output a prompt message indicating related gene repair defects.

9. The bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences according to claim 5, characterized in that: The method also includes a processing step for regions marked as to be verified: S36. Extract the local sequence of the region to be verified, and use the Bayesian inference model in combination with the pathogenic background type to perform a secondary evaluation to obtain the corrected pathogenicity probability. S37. If the correction probability of pathogenicity is greater than the preset confirmation threshold, the region to be verified will be upgraded and marked as a valid mutation; if the correction probability of pathogenicity is less than or equal to the preset confirmation threshold, the region to be verified will be downgraded and marked as an invalid alignment.

10. The bioinformatics comparison and analysis method for gastric cancer pathogenic gene sequences according to claim 7, characterized in that: The embedding process of S43 specifically includes: S431. Perform cluster analysis on novel structural variations, count their frequency of occurrence in a pre-set historical population sample database, and generate a frequency ranking table. S432. If the frequency of occurrence of novel structural variations is higher than the preset frequency threshold, they are updated as general atlas nodes to the base layer of the gastric cancer multidimensional topological reference graph as permanent reference paths. S433. If the frequency of occurrence of novel structural variations is less than or equal to a preset frequency threshold, they are stored as specific map nodes in a temporary variation layer and configured to be activated and loaded only when samples of the same pathogenic background type are detected.

Citation Information

Patent Citations

  • Big model technology-based biological information analysis system

    CN120260674A

  • Microsatellite instability characterization

    US20190108312A1