Gene sequence screening method and system for tumor markers
By iteratively dividing gene sequences and performing similarity difference analysis, the optimal length is adaptively determined, which solves the problem of inaccurate gene alignment and improves the accuracy of tumor marker screening.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE SECOND AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIVERSITY
- Filing Date
- 2025-07-01
- Publication Date
- 2026-05-28
Smart Images

Figure CN2025106387_28052026_PF_FP_ABST
Abstract
Description
A method and system for screening gene sequences of tumor markers Technical Field
[0001] This invention relates to the field of bioinformatics technology, specifically to a method and system for screening gene sequences of tumor markers. Background Technology
[0002] Tumor detection can be performed by extracting gene products from human tissues and body fluids. These gene products can then be compared with gene data of tumor markers in known gene banks. By comparing the concentration of different markers in the patient's genes, data support can be provided to relevant medical personnel.
[0003] When comparing and matching the gene sequences of patients to be tested with those of tumor markers, hash values can be used for labeling and matching, enabling rapid gene alignment. However, during hash value retrieval, due to the diversity of gene sequences, even minor differences in gene sequences during transcription, such as transcriptional deletions or errors leading to changes in a small number of bases, can result in significant differences in the hash values of the gene sequences. This can lead to mismatches between gene sequences, affecting the gene alignment results. Therefore, gene sequences can be segmented for matching. However, setting the data size of the segments too large or too small can affect the gene alignment results, leading to less accurate gene sequence screening results. Summary of the Invention
[0004] To address the technical problem that existing technologies suffer from inappropriate gene sequence segmentation, leading to inaccurate gene alignment results and consequently inaccurate gene sequence screening results, the present invention aims to provide a gene sequence screening method and system for tumor markers. The specific technical solution adopted is as follows:
[0005] In a first aspect, the present invention provides a method for screening gene sequences of tumor markers, comprising:
[0006] The gene sequences to be analyzed from patients with gastrointestinal tract disease and each original gene sequence of each tumor marker in the database are obtained. The original gene sequences are iteratively divided into gene data blocks of each original gene sequence at each preset length using different lengths.
[0007] Based on the matching of data similarity between each original gene sequence and other original gene sequences of the same type and length, as well as the differences between gene data blocks, the segmentation evaluation value of each original gene sequence at each length is obtained.
[0008] The optimal length is obtained by considering the differences in the partition evaluation values of all original gene sequences at each length and the partition evaluation values of all original gene sequences at adjacent lengths.
[0009] The gene sequence to be analyzed is divided using the optimal length. Based on the division result, the gene sequence to be analyzed is hash-matched with the division result corresponding to tumor markers in the database to obtain the gene sequence screening result of tumor markers for patients with gastrointestinal tract disease.
[0010] Preferably, the step of obtaining the segmentation evaluation value for each original gene sequence of each length based on the matching of data similarity between each original gene sequence and other original gene sequences of the same type of tumor marker at the same length, as well as the differences between gene data blocks, specifically includes:
[0011] For any tumor marker and any length;
[0012] Based on the similarity between each gene data block in each original gene sequence and each gene data block in every other original gene sequence, determine the global similarity feature value of each gene data block in each original gene sequence;
[0013] Based on the global similarity feature value, each gene data block in the original gene sequence is filtered to determine the differential data blocks;
[0014] Based on the maximum similarity between each differential data block in each original gene sequence and the corresponding matching data block in other original gene sequences, the partitioning evaluation value of each original gene sequence at any given length is obtained.
[0015] Preferably, determining the global similarity feature value of each gene data block in each original gene sequence based on the similarity between each gene data block in each original gene sequence and each gene data block in every other original gene sequence specifically includes:
[0016] Any gene data block in any original gene sequence is designated as the target data block, and the gene data block at the same position as the target data block in each other original gene sequence (excluding the aforementioned original gene sequence) is designated as the reference data block of the target data block; the target data block and the reference data block belong to the gene data corresponding to the same type of tumor marker.
[0017] Based on the DTW distance between the target data block and each reference data block, the similarity feature factor between the target data block and each reference data block is determined;
[0018] Based on the number of reference data blocks corresponding to the similarity feature factor being greater than a preset similarity threshold, the global similarity feature value of the target data block is determined, and the global similarity feature value is a normalized value.
[0019] Preferably, the step of filtering each gene data block in the original gene sequence based on the global similarity feature value to determine the differential data blocks specifically includes:
[0020] When the global similarity feature value of the target data block is less than the preset difference threshold, the target data block is a difference data block.
[0021] Preferably, the step of obtaining the partitioning evaluation value of each original gene sequence of any length based on the maximum similarity between each differential data block in each original gene sequence and the corresponding matching data block in other original gene sequences specifically includes:
[0022] In any original gene sequence of any length, the partitioning feature factor of each differential data block is determined based on the negative correlation coefficient of the maximum value of the similarity feature factor between each differential data block in the original gene sequence and the corresponding reference data block; the partitioning evaluation value of any original gene sequence of any length is determined by combining the partitioning feature factors of all differential data blocks in the original gene sequence.
[0023] Preferably, the step of obtaining the optimal length based on the differences in the partitioning evaluation values of all original gene sequences at each length and the partitioning evaluation values of all original gene sequences at adjacent lengths specifically includes:
[0024] In ascending order of length, the optimization degree of each length is obtained based on the difference in the evaluation values of all original gene sequences between each length and its adjacent lengths; the minimum length corresponding to the optimization degree being greater than the preset optimization threshold is taken as the optimal length.
[0025] Preferably, the step of determining the preference level of each length based on the difference in evaluation values of all original gene sequences between each length and its adjacent lengths specifically includes:
[0026] The mean value of the partition evaluation values of all original gene sequences at each length is calculated to obtain the partition effect characterization value for each length; in the order of increasing length, the partition effect slope value of each length is determined based on the rate of change of the partition effect characterization value between each length and the adjacent preceding length.
[0027] For any given length, a feature evaluation coefficient is determined based on the negative correlation coefficient of the difference in the slope value of the partitioning effect between the given length and the adjacent preceding length. Based on the feature evaluation coefficient and the partitioning effect characterization value of the given length, the degree of preference of the given length is obtained. Both the feature evaluation coefficient and the partitioning effect characterization value are positively correlated with the degree of preference.
[0028] Preferably, the step of performing hash matching between the gene sequence to be analyzed and the corresponding segmentation results of tumor markers in the database based on the segmentation results to obtain the gene sequence screening results of tumor markers for gastrointestinal patients specifically includes:
[0029] For any tumor marker, the difference data block corresponding to the minimum value of the global similarity feature under the optimal length is used as the seed data block. Based on the seed data block and each data block in the partitioning result of the gene sequence to be analyzed, hash matching is performed to obtain the matching result between the gene sequence to be analyzed and the tumor marker.
[0030] Based on the matching results of each tumor marker, the concentration percentage of each tumor marker in the gene sequence to be analyzed is obtained, and the gene sequence screening results of the tumor markers are determined according to the order of concentration from large to small.
[0031] Preferably, the matching result between the gene sequence to be analyzed and any one of the tumor markers is as follows:
[0032] When the proportion of matching data blocks in the partitioning results of the seed data block and the gene sequence to be analyzed is greater than the preset proportion threshold, the original gene sequence of the seed data block is matched with the gene sequence to be analyzed to obtain the matching result of the tumor marker.
[0033] Secondly, the present invention provides a gene sequence screening system for tumor markers, comprising:
[0034] The data extraction module is used to obtain the gene sequences to be analyzed from patients with gastrointestinal tract disease and each original gene sequence of each tumor marker in the database. The original gene sequences are iteratively divided into gene data blocks of each original gene sequence at each preset length using different lengths.
[0035] The segmentation processing module is used to obtain the segmentation evaluation value of each original gene sequence at each length based on the matching of data similarity between each original gene sequence and other original gene sequences of the same type of tumor marker at the same length, as well as the differences between gene data blocks.
[0036] The feature evaluation module is used to determine the optimal length based on the difference in the partition evaluation values of all original gene sequences at each length and the partition evaluation values of all original gene sequences at adjacent lengths.
[0037] The gene screening module is used to divide the gene sequence to be analyzed using the optimal length, and to perform hash matching between the gene sequence to be analyzed and the division results corresponding to tumor markers in the database to obtain the gene sequence screening results of tumor markers for patients with gastrointestinal tract diseases.
[0038] The embodiments of the present invention have at least the following beneficial effects:
[0039] This invention first obtains the gene sequences to be analyzed from patients undergoing gene alignment and the original gene sequences from a database. Then, iteratively partitions the gene information in the database into blocks of varying lengths, providing a data foundation for subsequent feature analysis based on these different lengths. Next, it analyzes the similarity and differences of genes in the database for the same type of tumor markers across different block lengths, quantifying the partitioning effect of each original gene sequence at each length (i.e., the partitioning evaluation value). Similarity reflects more general data expression across different gene sequences, while differences reflect more specific data expression. Furthermore, the partitioning evaluation results for all original gene sequences at different lengths are combined with the evaluation results of adjacent lengths to analyze the variation in differences between partitions of different lengths, in order to select the optimal choice and achieve an adaptive process for determining the optimal length. Finally, the gene sequences to be analyzed are partitioned using the optimal length and matched with the corresponding partitioning results of the optimal length in the database. This yields a more optimized block result, making the gene alignment matching results more accurate and valuable, ultimately leading to better gene screening results. Attached Figure Description
[0040] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 is a flowchart of the steps of a gene sequence screening method for tumor markers provided by the present invention;
[0042] Figure 2 is a flowchart of the steps for obtaining the segmentation evaluation value of the original gene sequence provided by the present invention;
[0043] Figure 3 is a flowchart of the steps of the method for obtaining the optimal length provided by the present invention;
[0044] Figure 4 is a local distribution coordinate diagram of the segmentation evaluation values of the original gene sequence provided by the present invention;
[0045] Figure 5 is a system block diagram of a gene sequence screening system for tumor markers provided by the present invention.
[0046] Figure 6 is a schematic diagram of the structure of a computer device provided by the present invention. Detailed Implementation
[0047] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a gene sequence screening method and system for tumor markers proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0049] The following detailed description, in conjunction with the accompanying drawings, illustrates the specific scheme of the gene sequence screening method and system for tumor markers provided by this invention.
[0050] Please refer to Figure 1, which shows a flowchart of a gene sequence screening method for tumor markers according to an embodiment of the present invention. The method includes the following steps:
[0051] Step S100: Obtain the gene sequence to be analyzed from the patient with gastrointestinal tract disease and each original gene sequence of each tumor marker in the database. Iteratively divide the original gene sequence into gene data blocks of each original gene sequence at each preset length using different lengths.
[0052] When the human body develops a tumor, the concentration of corresponding proteins or substances increases. These substances can be called tumor markers. Tumor marker synthesis generally involves protein expression through gene sequences. By comparing and screening the genes corresponding to tumor markers, the expression of tumor marker gene sequences can be preliminarily determined. When performing gene alignment screening, it is advisable to divide the gene sequence into blocks for comparison. However, if the block size is too large, it will not serve its purpose, and the impact of mutations or errors on the hash value will still exist. If the block size is too small, it is easy to identify multiple identical gene information in the gene sequence, which will also affect the accuracy of gene alignment. The longer the data block, the higher the probability of errors when matching the patient's gene sequence; conversely, the shorter the gene sequence, the lower its uniqueness within the entire database.
[0053] Based on this, this embodiment first obtains the protein expression of the patients to be compared with the gene. At the same time, it performs different length block processing on the gene information of each type of tumor marker in the gene database, which provides a data basis for subsequent adaptive analysis of the effect of block division of different lengths.
[0054] Specifically, the gene sequences to be analyzed from patients with gastrointestinal tract disease are first obtained, and multiple different gene expressions corresponding to each type of tumor marker are obtained from the database and recorded as the original gene sequences. That is, the same tumor marker may have multiple different gene expression forms in the database, and thus the same tumor marker corresponds to multiple original gene sequences.
[0055] In this embodiment, the gene sequences in the gene library are iteratively divided into blocks of different lengths. The iteration range for the length is (0, 1000), and the step size between each two adjacent length iterations is 5. It is understood that each length value is an integer, and the length increases sequentially according to the iteration step size. In other embodiments, the minimum length can also be the length of the smallest gene unit (base) in the gene sequence of the gene library. This length is used as the initial length for iterative increases, performing block division operations of different lengths for the original gene sequence.
[0056] It should be noted that, within a given length range, each original gene sequence of each type of tumor marker was segmented into blocks to obtain gene data blocks. This provides a data foundation for subsequent analysis of the segmentation of different original gene sequences of the same tumor marker under the same length segmentation results.
[0057] The extraction process for genetic information from patients with gastrointestinal issues is as follows: Approximately 10g of fecal sample is obtained from the colonoscopy and preoperative samples using a fecal specimen kit, and a fecal occult blood test is performed immediately. Reagent preparation: 90 ml of isopropanol is added to buffer GFA; 135 ml and 160 ml of anhydrous ethanol are added to protein removal buffer RD and washing buffer PWD, respectively; the preparation time is written and labeled. 0.2g of fecal sample is added to a 2 ml centrifuge tube A (200 μL of liquid sample is aspirated), followed by 500 μL, 100 μL, and 0.25g of buffers SA and SC, and a grinding bead, respectively; after vortexing and mixing, the tube is placed in a 70°C metal bath for lysis for 15 minutes. After lysis, the tube is centrifuged at 12000 rpm for 1 minute, and the supernatant is transferred to a new 2 ml centrifuge tube B. Add 200 μL of buffer SH to centrifuge tube B, vortex to mix, incubate at 4°C for 10 minutes, centrifuge at 12000 rpm for 2 minutes, transfer the supernatant to 2 mL centrifuge tube C, then add 500 μL of buffer GFR and 10 μL of magnetic bead suspension G, vortex to mix for 2 minutes. Place centrifuge tube C on a magnetic rack and let it stand for 30 seconds until the magnetic beads are completely adhered to the inner wall, then carefully aspirate all liquid. Remove centrifuge tube C and add 700 μL of protein removal solution RD, vortex to mix for 3 minutes, place centrifuge tube on a magnetic rack and let it stand for 30 seconds until the magnetic beads are completely adhered to the inner wall, then carefully aspirate all liquid. e. Remove centrifuge tube C and add 700 μL of protein removal solution PWD, vortex to mix for 3 minutes, place centrifuge tube on a magnetic rack and let it stand for 30 seconds until the magnetic beads are completely adhered to the inner wall, then carefully aspirate all liquid. Do not remove centrifuge tube C this time; allow it to air dry at room temperature for 5 minutes. Remove centrifuge tube C, add 80 μL of elution buffer TB, vortex to mix for 1 minute, and incubate in a 56°C metal bath for 5 minutes. Place centrifuge tube C on a magnetic rack and let it stand for 2 minutes until the magnetic beads are completely adsorbed onto the inner wall. Then carefully aspirate the DNA solution into a new centrifuge tube D.
[0058] Step S200: Based on the matching of data similarity between each original gene sequence and other original gene sequences of the same type of tumor marker at the same length, as well as the differences between gene data blocks, the segmentation evaluation value of each original gene sequence at each length is obtained.
[0059] When screening gene sequences for tumor markers, the gene information of tumor markers in known gene banks can be compared with the gene information of patients with gastrointestinal problems. By combining the hash value comparison with the base information, the matching accuracy of the gene information in the patient to be tested can be determined. After dividing the original gene sequence into segments of different lengths, a dynamic time warping algorithm can be used to perform data matching analysis on each corresponding gene data block, quantifying the evaluation of the segmentation effect at different lengths.
[0060] Meanwhile, considering that gene information comparison requires comparing the patient's gene information to be matched with all gene information of the same type of tumor marker one by one, and that the original gene sequences corresponding to the same type of tumor marker may have unique expression patterns, different lengths of block operations are used to screen out the most differential gene data blocks in each original gene sequence as gene expression blocks with unique or special characteristics. Using the expression of unique characteristics for subsequent gene comparison can effectively improve the accuracy of block gene comparison.
[0061] Based on this, in the results of segmenting the gene sequence by different lengths, the matching effect is analyzed by examining the matching similarity and matching difference between gene data blocks at corresponding positions in the segmentation results of the same category, so as to quantify the quality of using different lengths to segment the gene sequence. In this embodiment, as shown in Figure 2, the method for obtaining the segmentation evaluation value can be implemented by steps S201 to S203.
[0062] It should be noted that the evaluation of the segmentation operation is based on the segmentation results for each different length value. At the same time, the gene expression of special features or unique features is analyzed and determined for the gene information of the same tumor marker. Therefore, in this embodiment, the process of quantitatively analyzing the segmentation evaluation value is described using all the original gene sequences corresponding to any tumor marker, and the segmentation results under any length are used as examples for illustration.
[0063] Step S201: Determine the global similarity feature value of each gene data block in each original gene sequence based on the similarity between each gene data block in each original gene sequence and each gene data block in every other original gene sequence.
[0064] For any tumor marker and any segmentation length, the positional distribution of gene data blocks in each original gene sequence is the same. When comparing gene features, the matching between gene data blocks at the same position should also be considered. Therefore, feature analysis is first performed on the similarity of gene data blocks in the current segmentation result.
[0065] Taking the original gene sequence of any tumor marker under any length of partitioning result as an example, we can further illustrate this by designating any gene data block in any original gene sequence as the target data block, and designating the gene data block at the same position as the target data block in each other original gene sequence (excluding the aforementioned original gene sequence) as the reference data block of the target data block. The target data block and the reference data block belong to the gene data corresponding to the same type of tumor marker. For example, if the i-th gene data block in the a-th original gene sequence is taken as the target data block, then the i-th gene data block in each other original gene sequence belonging to the same partitioning result and the same tumor marker as the a-th original gene sequence is the reference data block of the target data block.
[0066] Then, based on the DTW distance between the target data block and each reference data block, the similarity feature factor between the target data block and each reference data block is determined. The calculation method of the DTW distance is a well-known technique and will not be described in detail here. The DTW distance between the target data block and each reference data block reflects the difference in the corresponding matching between the target data block and the reference data blocks. By performing a negative correlation mapping on the calculated DTW distance, the similarity between the target data block and each reference data block can be obtained.
[0067] In this embodiment, the DTW distance between the target data block and each reference data block is negatively correlated and normalized to obtain a similarity feature factor between the target data block and each reference data block. It should be understood that the negative correlation normalization method is a well-known technique, and implementers can choose the calculation method according to the specific implementation scenario. As a specific example, the similarity feature factor can be expressed as... X a,i X represents the i-th gene data block in the a-th original gene sequence, which is also the target data block; b,i Let L(X) represent the i-th gene data block in the b-th original gene sequence, which is also a reference data block for the target data block, and a ≠ b; a,i ,X b,i ) represents the similarity feature factor between the target data block and a reference data block, max() represents the function to find the maximum value, min() represents the function to find the minimum value, max(DTW) represents the maximum DTW distance between the target data block and the reference data block, and min(DTW) represents the minimum DTW distance between the target data block and the reference data block. a,i ,X b,i ) represents the target data block X a,i With reference data block X b,i The DTW distance between them.
[0068] The smaller the DTW distance value, the larger the similarity feature factor value between the two gene data blocks, indicating a greater degree of similarity and matching between them. Using the dynamic time warping algorithm for difference analysis can achieve optimal matching between two different original gene sequences. This effectively avoids the problem of relatively small differences when only one unit feature is added or removed, resulting in large differences in the calculated results, thus improving the accuracy of the data analysis results to a certain extent.
[0069] Finally, based on the number of reference data blocks corresponding to when the similarity feature factor is greater than a preset similarity threshold, the global similarity feature value of the target data block is determined, and the global similarity feature value is a normalized value.
[0070] The larger the value of the similarity feature factor between the target data block and the reference data block, the greater the similarity between the two corresponding gene data blocks. By counting the number of the target data block and all reference data blocks with a large degree of similarity, we can reflect the similarity performance characteristics of the target data block in the database. The higher the similarity ratio, the more consistent the target data block is, which means that the target data block is not a gene expression block with special or unique characteristics.
[0071] Specifically, when the similarity feature factor between the target data block and the reference data block is greater than a preset similarity threshold, the number of reference data blocks is counted, indicating a high degree of gene similarity between the target data block and the reference data block. In this embodiment, the similarity threshold is set to 0.8. Since the similarity feature factor is a normalized value, the similarity threshold ranges from (0,1). The closer the similarity threshold is to 1, the more stringent the measurement of gene similarity in the target data block; the closer the similarity threshold is to 0, the more lenient the measurement of gene similarity in the target data block.
[0072] In this embodiment, the number of reference data blocks that meet the threshold requirements for all similarity feature factors corresponding to the target data block is counted. The ratio between this number and the total number of all reference data blocks corresponding to the target data block is used as the global similarity feature value of the target data block, reflecting the proportion of gene data blocks with a high degree of similarity features to the target data block.
[0073] Step S202: Based on the global similarity feature value, each gene data block in the original gene sequence is screened to determine the differential data blocks.
[0074] The higher the global similarity value of the target data block, the greater the similarity between the gene information of the target data block and the gene information at the same position in other original gene sequences, and the more likely the target data block is to represent general gene expression. Conversely, the lower the global similarity value of the target data block, the less similar the gene information of the target data block is to the gene information at the same position in other original gene sequences, and the more likely the target data block is to represent gene expression with specific characteristics.
[0075] Using the same method as the global similarity feature value of the target data block, the global similarity feature value of each gene data block in each original gene sequence can be obtained. Then, the gene data blocks can be screened by the global similarity feature value to identify gene expression parts that may have special features, that is, differential data blocks.
[0076] Specifically, taking the target data block as an example, when the global similarity feature value of the target data block is less than a preset difference threshold, the target data block is a difference data block. The difference threshold is set to 0.3, which can be set by the implementer according to the specific implementation scenario. That is, the global similarity feature value of each gene data block is judged separately, and when the threshold requirement is met, the corresponding gene data block is a difference data block. A difference data block represents the gene expression portion in the corresponding original gene sequence that has special or differential features.
[0077] Step S203: Based on the maximum similarity between each differential data block in each original gene sequence and the corresponding matching data block in other original gene sequences, obtain the partitioning evaluation value of each original gene sequence under any length.
[0078] Analyzing the similarity between the expression regions of all characteristic genes in an original gene sequence and other original gene sequences can effectively reflect the segmentation results of the current length block partitioning. It should be noted that, under the same tumor marker, each gene data block in each original gene sequence has a corresponding reference data block, which represents the gene data blocks at the same positions in other original gene sequences; in other words, each differential data block also has a corresponding reference data block.
[0079] Specifically, in any original gene sequence of any length, the partitioning feature factor of each differential data block is determined based on the negative correlation coefficient of the maximum value of the similarity feature factor between each differential data block in the original gene sequence and the corresponding reference data block; the partitioning evaluation value of any original gene sequence of any length is determined by combining the partitioning feature factors of all differential data blocks in the original gene sequence.
[0080] As a concrete example, this embodiment uses the partitioning result corresponding to any length value as an example for illustration, and also uses any original gene sequence as an example for illustration. Then, under the value of length d, the partitioning evaluation value of the a-th original gene sequence can be expressed as: W a (d) represents the partition evaluation value of the a-th original gene sequence under the partition result of length d, X' a,c This represents the c-th differential data block of a original gene sequences, N. a (d) represents the total number of differentially expressed data blocks contained in the a-th original gene sequence under the partitioning result of length d, L(X' a,c ) max This represents the maximum value of the similarity feature factor between the c-th differential data block in the a-th original gene sequence and the corresponding reference data block under the partitioning result of length d.
[0081] The negative correlation coefficient 1-L(X') is used to calculate the maximum similarity between the difference data block and the corresponding reference data block. a,c ) max This indicates the effectiveness of data matching at the current length, avoiding situations where a single gene data block exhibits significant differences. In subsequent iterations, if most data within a gene data block differs greatly from other gene data blocks, this block will appear exceptional. However, exceptional gene data blocks may exist in multiple overlapping gene data block intervals. Therefore, the similarity between gene data blocks needs to be used as a weight to adjust and optimize the overall gene sequence partitioning. The partitioning evaluation value of the original gene sequence characterizes the partitioning performance of the original gene sequence using the current length.
[0082] Step S300: Based on the differences in the partition evaluation values of all original gene sequences at each length and the partition evaluation values of all original gene sequences at adjacent lengths, the optimal length is obtained.
[0083] By analyzing the gene similarity expression and special feature expression among different original gene sequences corresponding to the same tumor marker in the segmentation results for each length, the segmentation evaluation value can be used to quantify the segmentation effect of different original gene sequences being segmented at different lengths. By comprehensively evaluating the segmentation effect of all original gene sequences under the same length value, the effect of segmentation using the corresponding length can be globally characterized. The result of the global characterization can be used to select the length with the best segmentation effect.
[0084] In this embodiment, the method for obtaining the optimal length is shown in Figure 3, which can be implemented by steps S301 to S302.
[0085] Step S301: In ascending order of length, the optimization degree of each length is obtained based on the difference in the evaluation values of all original gene sequences between each length and its adjacent lengths.
[0086] For each length value, each original gene sequence corresponds to a partition evaluation value. Figure 4 shows the coordinate distribution of partition evaluation values for some original gene sequences under some length values. That is, a partial data is selected for illustrative representation. In Figure 4, x represents the horizontal axis, which is the length value, and y represents the vertical axis, which is the partition evaluation value of different original gene sequences in the corresponding length partition results.
[0087] In this embodiment, the mean value of the partition evaluation values of all original gene sequences at each length is calculated to obtain the partition effect characterization value for each length. That is, the partition effect characterization value comprehensively reflects the performance of the partitioning effect of the corresponding length.
[0088] Following an increasing length order, the slope value of the partitioning effect for each length is determined based on the rate of change of the partitioning effect representation value between each length and its preceding adjacent length. Each length's value corresponds to a partitioning effect representation value, and the method for obtaining the partitioning effect slope value for each length can be analogous to the slope value of a data point in a two-dimensional coordinate system. The calculation method is a well-known technique, and in this embodiment, it is simply explained as a specific example.
[0089] Specifically, different length values are used as the x-axis, and the corresponding segmentation effect value is used as the y-axis to construct data points for each length. Then, curve fitting is performed on the data points of each length to obtain the fitted segmentation effect curve. The slope of the data points at each length can be calculated on the segmentation effect curve, and then used as the segmentation effect slope value for each length. The methods for fitting the data points and calculating the slope of the data points on the curve are well-known techniques and will not be described in detail here.
[0090] For any given length, a feature evaluation coefficient is determined based on the negative correlation coefficient of the difference in the slope value of the partitioning effect between the given length and the adjacent preceding length. Based on the feature evaluation coefficient and the partitioning effect characterization value of the given length, the degree of preference of the given length is obtained. Both the feature evaluation coefficient and the partitioning effect characterization value are positively correlated with the degree of preference.
[0091] As a concrete example, the optimality Q corresponding to the t-th length t It can be represented as Among them, Q t This indicates the degree of preference for the t-th length. K represents the partitioning performance value for the t-th length sequence, which is the mean of the partitioning evaluation values for all original gene sequences at the t-th length. t K represents the slope value of the partition effect of the t-th length. t-1 Let represent the slope value of the partition effect of the (t-1)th length, exp[] represents the exponential function with the natural constant e as the base, and Norm represents the linear normalization function.
[0092] exp[-(K t -K t-1 [)] represents the feature evaluation coefficient of the t-th length, reflecting the difference in the partitioning efficiency value between the t-th length and its adjacent preceding length. K t -K t-1 The smaller the value of , the faster the rate of change of the partitioning effect between the (t-1)th length and the tth length changes from a rapid change to a slow change. This means that the improvement of the effect of the length value is close to the inflection point of the length. The larger the value of the corresponding feature evaluation coefficient, the greater the degree of optimization of the current length, and the better the effect of using this length to partition the gene sequence.
[0093] Meanwhile, the segmentation performance value comprehensively represents the evaluation of the segmentation results of all original gene sequences at the current length. The larger the value, the greater the optimization degree of the corresponding current length, and the better the effect of gene sequence segmentation operation using this length. That is, the optimization degree represents the accuracy of feature matching when using the corresponding length to segment different original gene sequences.
[0094] It should be noted that when analyzing the changes in the partitioning effect during the iterative block division process, it is necessary to analyze the partitioning results under at least several different length values. That is, lengths whose optimality cannot be calculated in this embodiment are not considered.
[0095] Step S302: The minimum length corresponding to the degree of preference being greater than the preset preference threshold is taken as the optimal length.
[0096] The higher the optimization level of different lengths, the better the feature comparison effect of the corresponding length of gene sequence division. Therefore, when the optimization level is greater than the preset optimization threshold, the division effect of the corresponding length is better. Considering that the longer the block division length, the more prone it is to errors in the gene comparison process using hash values, the minimum length that meets the threshold requirement is selected as the optimal length. In this embodiment, the optimization threshold is set to 0.7, which can be set by the implementer according to the specific implementation scenario, and the range of the optimization threshold is (0,1).
[0097] Step S400: The gene sequence to be analyzed is divided using the optimal length. Based on the division result, the gene sequence to be analyzed is hash-matched with the division result corresponding to tumor markers in the database to obtain the gene sequence screening result of tumor markers for gastrointestinal patients.
[0098] Gene sequences are extremely large and complex, making direct full-sequence alignment very time-consuming and computationally resource-intensive. Segmenting the sequence and generating hash values can significantly accelerate the search and alignment process. Hash values generated by hash functions have a fixed and unique length, making them effective indexes for fast lookups and alignments. Segmentation and hash matching allow for parallel processing of different sequence fragments or data blocks, thereby improving overall processing efficiency and speed. This is particularly important for processing large-scale genomic data. In this embodiment, by iterating through different segmentation lengths and performing feature comparison, an adaptive segmentation length that best represents the original gene expression in the current database is determined. Using this segmentation length for feature comparison of the patient's gene information to be compared yields more accurate hash matching results.
[0099] Specifically, by using the optimal length to divide the patient's gene sequence to be analyzed into blocks, the corresponding gene partitioning results for the patient can be obtained, including each data block of the gene sequence to be analyzed. Similarly, by using the optimal length to divide each original gene sequence in the database into blocks, each gene data block of the original gene sequence can be obtained, which can be obtained from the above data analysis process.
[0100] Furthermore, the hash value of each data block in the gene sequence to be analyzed, and the hash value of each gene data block in the original gene sequence are calculated separately. Using these hash values, feature comparison and matching operations can be performed on each data block in the gene sequence to be analyzed and gene data blocks at the same position in each original gene sequence to obtain matching results. Feature comparison and matching using hash values is a well-known technique; that is, if the hash values are the same, it is a match, and if they are different, it is a mismatch. This will not be discussed further here.
[0101] Finally, the screening results of tumor markers for patients with gastrointestinal tract diseases can be displayed through the matching results. For example, the number of matching gene sequences can be arranged in descending order for each tumor marker. This sorting result can provide relevant doctors with some big data support.
[0102] More preferably, the data processing of directly performing hash value matching is also quite large. It can perform preliminary screening on gene expression blocks with special characteristics under the current partitioning results in the database, and determine whether the obtained gene sequences need to be expanded for extended matching.
[0103] Specifically, for any tumor marker, the difference data block corresponding to the minimum value of the global similarity feature under the optimal length is used as the seed data block. Thus, the seed data block represents the most unique performance among all gene data blocks of the current tumor marker.
[0104] Hash matching is performed on each data block in the partitioning result of the seed data block and the gene sequence to be analyzed to obtain the matching result between the gene sequence to be analyzed and any one of the tumor markers. Similarly, data blocks with the same hash value are matched, and data blocks with different hash values are not matched.
[0105] The specific method for obtaining the matching results is as follows: When the proportion of matching data blocks in the partitioning results of the seed data block and the gene sequence to be analyzed is greater than a preset proportion threshold, the original gene sequence containing the seed data block is matched with the gene sequence to be analyzed to obtain the matching result of the tumor marker. The proportion threshold is set to 50%, meaning that when the proportion of matching data blocks is greater than 50%, it indicates that more than half of the specific gene expressions are contained in the current patient gene information to be compared. To avoid errors in the comparison, an extended matching operation can be performed, that is, a longer gene sequence is used for matching to determine the matching result. When the proportion of matching data blocks in the partitioning results of the seed data block and the sequence to be analyzed is less than the preset proportion threshold, the proportion of that data block can be directly used as the matching result.
[0106] Based on the matching results of each tumor marker, the concentration percentage of each tumor marker in the gene sequence to be analyzed is obtained, and the gene sequence screening results of the tumor markers are determined according to the order of concentration from largest to smallest. The concentration percentage can be represented by the percentage of matching data blocks.
[0107] As shown in Figure 5, this embodiment of the invention also provides a gene sequence screening system for tumor markers, the system comprising:
[0108] The data extraction module is used to obtain the gene sequences to be analyzed from patients with gastrointestinal tract disease and each original gene sequence of each tumor marker in the database. The original gene sequences are iteratively divided into gene data blocks of each original gene sequence at each preset length using different lengths.
[0109] The segmentation processing module is used to obtain the segmentation evaluation value of each original gene sequence at each length based on the matching of data similarity between each original gene sequence and other original gene sequences of the same type of tumor marker at the same length, as well as the differences between gene data blocks.
[0110] The feature evaluation module is used to determine the optimal length based on the difference in the partition evaluation values of all original gene sequences at each length and the partition evaluation values of all original gene sequences at adjacent lengths.
[0111] The gene screening module is used to divide the gene sequence to be analyzed using the optimal length, and to perform hash matching between the gene sequence to be analyzed and the division results corresponding to tumor markers in the database to obtain the gene sequence screening results of tumor markers for patients with gastrointestinal tract diseases.
[0112] It should be noted that the system provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the gene sequence screening system for tumor markers and the gene sequence screening method for tumor markers provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0113] This application also provides a computer device. Please refer to FIG6, which shows a schematic diagram of a computer device structure provided in an embodiment of the present invention. The computer device includes a memory 601, a processor 602, and a computer program 603 stored in the memory 601 and running on the processor 602. When the processor 602 executes the computer program 603, the computer device can perform any of the gene sequence screening methods for tumor markers described above.
[0114] This application also provides a computer program product that, when run on a computer device, enables the computer device to execute any of the aforementioned gene sequence screening methods for tumor markers.
[0115] This application also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer device, the computer device can execute any of the aforementioned methods for screening gene sequences of tumor markers.
[0116] In the embodiments provided in this application, it should be understood that the computer device, computer program product and computer-readable storage medium provided are all used to perform the corresponding methods provided above, and therefore the beneficial effects they can achieve can be referred to the beneficial effects of the methods provided above, which will not be repeated here.
[0117] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for screening gene sequences of tumor markers, characterized in that, The method includes the following steps: The gene sequences to be analyzed from patients with gastrointestinal tract disease and each original gene sequence of each tumor marker in the database are obtained. The original gene sequences are iteratively divided into gene data blocks of each original gene sequence at each preset length using different lengths. Based on the matching of data similarity between each original gene sequence and other original gene sequences of the same type and length, as well as the differences between gene data blocks, the segmentation evaluation value of each original gene sequence at each length is obtained. The optimal length is obtained by considering the differences in the partition evaluation values of all original gene sequences at each length and the partition evaluation values of all original gene sequences at adjacent lengths. The gene sequence to be analyzed is divided using the optimal length. Based on the division result, the gene sequence to be analyzed is hash-matched with the division result corresponding to tumor markers in the database to obtain the gene sequence screening result of tumor markers for patients with gastrointestinal tract disease.
2. The method for screening gene sequences of tumor markers according to claim 1, characterized in that, The method involves matching the data similarity between each original gene sequence of the same type and length and other original gene sequences, as well as the differences between gene data blocks, to obtain the segmentation evaluation value for each original gene sequence of each length. Specifically, this includes: For any tumor marker and any length; Based on the similarity between each gene data block in each original gene sequence and each gene data block in every other original gene sequence, determine the global similarity feature value of each gene data block in each original gene sequence; Based on the global similarity feature value, each gene data block in the original gene sequence is filtered to determine the differential data blocks; Based on the maximum similarity between each differential data block in each original gene sequence and the corresponding matching data block in other original gene sequences, the partitioning evaluation value of each original gene sequence at any given length is obtained.
3. The method for screening gene sequences of tumor markers according to claim 2, characterized in that, The step of determining the global similarity feature value of each gene data block in each original gene sequence based on the similarity between each gene data block in each original gene sequence and each gene data block in every other original gene sequence specifically includes: Any gene data block in any original gene sequence is designated as the target data block, and the gene data block at the same position as the target data block in each other original gene sequence (excluding the aforementioned original gene sequence) is designated as the reference data block of the target data block; the target data block and the reference data block belong to the gene data corresponding to the same type of tumor marker. Based on the DTW distance between the target data block and each reference data block, the similarity feature factor between the target data block and each reference data block is determined; Based on the number of reference data blocks corresponding to the similarity feature factor being greater than a preset similarity threshold, the global similarity feature value of the target data block is determined, and the global similarity feature value is a normalized value.
4. The method for screening gene sequences of tumor markers according to claim 3, characterized in that, The step of filtering each gene data block in the original gene sequence based on the global similarity feature value to determine the differential data blocks specifically includes: When the global similarity feature value of the target data block is less than the preset difference threshold, the target data block is a difference data block.
5. The method for screening gene sequences of tumor markers according to claim 3, characterized in that, The step of obtaining a partitioning evaluation value for each original gene sequence of any length based on the maximum similarity between each differential data block in each original gene sequence and the corresponding matching data blocks in other original gene sequences specifically includes: In any original gene sequence of any length, the partitioning feature factor of each differential data block is determined based on the negative correlation coefficient of the maximum value of the similarity feature factor between each differential data block in the original gene sequence and the corresponding reference data block; the partitioning evaluation value of any original gene sequence of any length is determined by combining the partitioning feature factors of all differential data blocks in the original gene sequence.
6. The method for screening gene sequences of tumor markers according to claim 1, characterized in that, The optimal length is obtained by considering the differences in the partitioning evaluation values of all original gene sequences at each length and the partitioning evaluation values of all original gene sequences at adjacent lengths. Specifically, this includes: In ascending order of length, the optimization degree of each length is obtained based on the difference in the evaluation values of all original gene sequences between each length and its adjacent lengths; the minimum length corresponding to the optimization degree being greater than the preset optimization threshold is taken as the optimal length.
7. The method for screening gene sequences of tumor markers according to claim 6, characterized in that, The optimization degree of each length is obtained based on the difference in evaluation values of all original gene sequences between each length and its adjacent lengths, specifically including: The mean value of the partition evaluation values of all original gene sequences at each length is calculated to obtain the partition effect characterization value for each length; in the order of increasing length, the partition effect slope value of each length is determined based on the rate of change of the partition effect characterization value between each length and the adjacent preceding length. For any given length, a feature evaluation coefficient is determined based on the negative correlation coefficient of the difference in the slope value of the partitioning effect between the given length and the adjacent preceding length. Based on the feature evaluation coefficient and the partitioning effect characterization value of the given length, the degree of preference of the given length is obtained. Both the feature evaluation coefficient and the partitioning effect characterization value are positively correlated with the degree of preference.
8. The method for screening gene sequences of tumor markers according to claim 2, characterized in that, The step involves hash matching the gene sequences to be analyzed with the corresponding segmentation results of tumor markers in the database based on the segmentation results, to obtain the gene sequence screening results for tumor markers in gastrointestinal patients. Specifically, this includes: For any tumor marker, the difference data block corresponding to the minimum value of the global similarity feature under the optimal length is used as the seed data block. Based on the seed data block and each data block in the partitioning result of the gene sequence to be analyzed, hash matching is performed to obtain the matching result between the gene sequence to be analyzed and the tumor marker. Based on the matching results of each tumor marker, the concentration percentage of each tumor marker in the gene sequence to be analyzed is obtained, and the gene sequence screening results of the tumor markers are determined according to the order of concentration from large to small.
9. The method for screening gene sequences of tumor markers according to claim 8, characterized in that, The specific matching results between the gene sequence to be analyzed and any of the tumor markers are as follows: When the proportion of matching data blocks in the partitioning results of the seed data block and the gene sequence to be analyzed is greater than the preset proportion threshold, the original gene sequence of the seed data block is matched with the gene sequence to be analyzed to obtain the matching result of the tumor marker.
10. A gene sequence screening system for tumor markers, characterized in that, The system includes: The data extraction module is used to obtain the gene sequences to be analyzed from patients with gastrointestinal tract disease and each original gene sequence of each tumor marker in the database. The original gene sequences are iteratively divided into gene data blocks of each original gene sequence at each preset length using different lengths. The segmentation processing module is used to obtain the segmentation evaluation value of each original gene sequence at each length based on the matching of data similarity between each original gene sequence and other original gene sequences of the same type of tumor marker at the same length, as well as the differences between gene data blocks. The feature evaluation module is used to determine the optimal length based on the difference in the partition evaluation values of all original gene sequences at each length and the partition evaluation values of all original gene sequences at adjacent lengths. The gene screening module is used to divide the gene sequence to be analyzed using the optimal length, and to perform hash matching between the gene sequence to be analyzed and the division results corresponding to tumor markers in the database to obtain the gene sequence screening results of tumor markers for patients with gastrointestinal tract diseases.