A pattern recognition method and device for medical gene sequencing data
By introducing technologies such as adaptive sliding windows, machine learning-driven adaptive scoring matrix and Gaussian hybrid model in HPV virus genome sequence analysis, the problems of insufficient data filtering accuracy and insufficient processing capabilities of complex variant modes are solved, and higher analysis accuracy and complex pattern recognition capabilities are achieved.
Patent Information
- Application Number
- CN202510162794.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-14
AI Technical Summary
In the analysis of HPV virus genome sequences, the problems of insufficient data filtering accuracy, insufficient ability to deal with complex variant patterns, and many false positive results.
Adaptive sliding window mechanism is used to perform sequence preprocessing, and a machine learning-driven adaptive scoring matrix is built for sequence alignment, and analyzing variant regions are identified and analyzed in combination with Gaussian hybrid model and improved clustering algorithm.
Improves the accuracy of sequence data filtering, enhances the ability to process complex variance patterns, and reduces false positive results.
Smart Images

Figure CN119626346B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a method and device for pattern recognition of medical gene sequencing data. Background Art
[0002] In the field of bioinformatics data processing technology, human papillomavirus (HPV) genome sequence analysis and mutation detection system is an important technical direction. In particular, the data processing of the E2 gene sequence in the HPV genome is of great technical significance for determining the integration status of the viral sequence. In the prior art, E2 gene sequence analysis mainly relies on traditional sequence alignment algorithms for data processing. These processing methods compare the target sequence to be analyzed with the standard reference sequence to identify the difference sites in the sequence, including various types of mutations such as gene deletion, point mutation, and frameshift mutation.
[0003] However, traditional sequence data processing methods have many technical defects when dealing with HPV viral genome sequence analysis: the existing technology generally uses the Smith-Waterman algorithm with fixed parameters for sequence alignment, and its scoring matrix lacks the ability to dynamically adapt to sequence specificity and evolutionary conservation, resulting in insufficient accuracy when processing highly variable regions. At the same time, traditional methods often use fixed-size sliding windows and unified quality thresholds when performing sequence preprocessing, which cannot adapt to the GC content characteristics of different regions, resulting in insufficient accuracy in data filtering. In the detection of variant sites, existing technologies mostly rely on simple sequence difference statistics, lack comprehensive consideration of local background characteristics and evolutionary information of the sequence, and are prone to false positive results. In addition, in the process of variant classification, traditional methods often process adjacent variants independently, ignoring the possible correlation between variants, and are unable to accurately identify and analyze complex mutation patterns.
[0004] Therefore, there is an urgent need for a technical solution that can improve the quality of sequence data filtering and enhance the ability to handle complex variation patterns. Summary of the invention
[0005] In order to solve the deficiencies of the prior art, the present application discloses a pattern recognition method and device for medical gene sequencing data. The present application solves the technical problems of insufficient accuracy of data filtering in the prior art, and is implemented through the following technical solutions.
[0006] A pattern recognition method for medical gene sequencing data comprises the following steps: preprocessing a target gene sequence obtained by sequencing, including sequence reading, quality assessment, pairing splicing, and length screening, to generate a preprocessed sequence file; analyzing a reference sequence library based on the preprocessed sequence file, constructing an adaptive scoring matrix, and dynamically adjusting site scores by using neural network training; using the adaptive scoring matrix, scanning sequence differences with a sliding window and combining with Gaussian mixture model analysis to generate a variant region information file; reading the variant region information file, constructing a base conversion matrix for sequence comparison, and generating a variant classification result by analyzing the comparison path; standardizing a variant region feature vector based on the variant classification result, and generating a mutation pattern type by using an improved clustering algorithm.
[0007] The target gene sequence obtained by sequencing is preprocessed, including sequence reading, quality assessment, pairing splicing, length screening, and the steps of generating a preprocessed sequence file include: reading the gene sequence file output by the sequencer, and establishing a corresponding relationship array between the base sequence and the quality value; using an adaptive sliding window to assess the sequence quality, and marking the area with a quality value lower than a preset threshold as a low-quality area; using pairing information to splice the paired sequences, setting the minimum overlap length and the allowed mismatch rate; performing length screening on the spliced sequence, and using a dynamic programming algorithm to remove the connector sequences at both ends of the sequence; writing the preprocessed sequence into a temporary file, and converting it into a target format.
[0008] The steps of analyzing the reference sequence library based on the preprocessed sequence file, constructing an adaptive scoring matrix, and dynamically adjusting the site score by using neural network training include: reading the preprocessed sequence file, analyzing the gene sequence in the reference sequence library, counting the base occurrence frequency of each site, and generating a frequency matrix; calculating the information entropy of each site based on the frequency matrix, and determining the conservation score of the site according to the information entropy; dividing the sequences in the reference sequence library into a training set and a validation set according to a preset ratio, and performing a multiple sequence alignment on the training set; extracting the feature vector of each site in the multiple sequence alignment result, and inputting it into the neural network for training; dynamically adjusting the site-specific score based on the output of the trained neural network; determining the conservation level according to the site-specific score, and adopting a differentiated scoring strategy for regions with different conservation levels.
[0009] The steps of generating a variant region information file by using an adaptive scoring matrix, scanning sequence differences with a sliding window and combining with a Gaussian mixture model analysis include: based on a differentiated scoring strategy, scanning sequence differences with a sliding window, marking regions with a mismatch rate exceeding a preset threshold as potential variant regions; extracting the context of the preceding and following sequences for the potential variant regions, and calculating local background features; using the local background features, using a Gaussian mixture model to determine the content threshold interval, and integrating sequence alignment information to evaluate conservation; calculating the secondary structure stability of the sequence for the potential variant region, and using the free energy change value as a feature; and merging and analyzing multiple closely adjacent variant sites based on the features to generate a variant region information file.
[0010] The steps of reading the variant region information file, constructing a base conversion matrix for sequence alignment, and generating a variant classification result by analyzing the alignment path include: reading the variant region information file, constructing a base conversion matrix, and determining the variant type by sequence alignment; based on the base conversion matrix, aligning the reference sequence and the target sequence of the variant region to obtain the optimal alignment path; dividing the variant region into three types of substitution, insertion or deletion according to the optimal alignment path; calculating the feature vectors of the before and after sequences of the divided variant region, including base composition and structural characteristics; analyzing the sequence fragments between the variant regions, determining the associated mutations, and generating the variant classification result.
[0011] The steps of normalizing the feature vector of the variant region based on the variant classification result and generating the mutation pattern type through the improved clustering algorithm include: reading the variant classification result, constructing a feature vector for each variant region, including the features of the variant type code, position value, and length value; normalizing the feature vector so that the mean and standard deviation of each dimension meet preset conditions; grouping based on the standardized feature vector using the improved clustering algorithm, and introducing a distance weight based on sequence similarity; and determining the convergence condition of the clustering grouping result through iterative optimization to determine the mutation pattern type.
[0012] The steps of calculating the information entropy of each site based on the frequency matrix and determining the conservation score of the site according to the information entropy include: calculating the logarithmic value of the base occurrence frequency of each site, and accumulating the product of the logarithmic value and the base frequency of the site; taking the negative value of the cumulative sum to obtain the information entropy value of the site; comparing the information entropy value of the site with a preset reference value to determine the conservation score of the site.
[0013] The step of dynamically adjusting the site-specific score based on the trained neural network output includes: inputting the feature vector of the site into the trained neural network to obtain the predicted probability distribution vector of the occurrence of various nucleotides at the site; calculating the logarithmic value of the predicted probability of each nucleotide in the predicted probability distribution vector, and summing the product of the logarithmic value and the predicted probability to take the negative value; searching for homologous sequences in an evolutionary database and performing multiple comparisons on the obtained homologous sequences; counting the occurrence frequency of each nucleotide in the multiple comparison results, calculating the sum of the logarithmic products of the frequencies, and obtaining an evolutionary conservation score; and performing a weighted combination of the predicted probability distribution calculation result and the evolutionary conservation score to obtain an adjusted value of the site-specific score.
[0014] The steps of determining the conservation level according to the site-specific score and adopting a differentiated scoring strategy for regions with different conservation levels include: comparing the evolutionary conservation score of the site with a preset first threshold, and comparing the information entropy with a preset second threshold; dividing the site into a highly conserved region, a variable region or a moderately conserved region according to the comparison result; setting a first set of benchmark values for matching scores, mismatch deductions and gap penalties for highly conserved regions; setting a second set of benchmark values for matching scores, mismatch deductions and gap penalties for variable regions; adopting a weighted average of the first and second sets of benchmark values for moderately conserved regions; calculating the probability distribution of each type of substitution, and correcting the benchmark values of each conservatism level region according to the substitution probability distribution to obtain a final score.
[0015] A pattern recognition device for medical gene sequencing data, characterized in that it includes: a processor, a memory, and a system bus; wherein the processor and the memory are connected via the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, which, when executed by the processor, enable the processor to execute the pattern recognition method for medical gene sequencing data described in any one of the above embodiments.
[0016] In a pattern recognition method and device for medical gene sequencing data disclosed above, the embodiment of the present application introduces an adaptive sliding window mechanism in the sequence preprocessing stage, which can improve the accuracy of quality control. Secondly, in terms of the sequence alignment algorithm, a machine learning-driven adaptive scoring matrix is constructed, and dynamic adjustment of site-specific scores is achieved through a three-layer neural network structure. In terms of variant site detection and classification, the embodiment of the present application establishes a multi-level analysis framework to achieve the positioning of potential variant regions. By analyzing the distance and co-occurrence characteristics between variant regions, the identification of complex mutation patterns is achieved. In the pattern clustering analysis stage, a distance weight mechanism based on sequence similarity is introduced, which significantly improves the accuracy of clustering and enhances the ability to handle complex variant patterns. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 A schematic diagram of a flow chart of a pattern recognition method for medical gene sequencing data disclosed in an embodiment of the present application;
[0019] Figure 2 A schematic diagram of the conservation analysis of the E2 gene sequence disclosed in the examples of the present application;
[0020] Figure 3 A sequencing quality distribution comparison diagram disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0021] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.
[0022] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor represent the necessary logical order between them. It should also be understood that in the embodiments of the present disclosure, "multiple" can refer to two or more, and "at least one" can refer to one, two or more. It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, in the absence of explicit limitation or contrary revelation given in the context, it can generally be understood as one or more. In addition, the term "and / or" in the present disclosure is only a description of the association relationship of the associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present disclosure generally indicates that the associated objects before and after are an "or" relationship. It should also be understood that the description of each embodiment in the present disclosure emphasizes the differences between the embodiments, and the same or similar parts can refer to each other. For the sake of brevity, they will not be repeated one by one.
[0023] At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present disclosure and its application or use. The techniques, methods and devices known to ordinary technicians in the relevant fields may not be discussed in detail, but where appropriate, the techniques, methods and devices should be considered part of the specification. It should be noted that similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0024] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0025] Figure 1 A flow chart of a pattern recognition method for medical gene sequencing data disclosed in an embodiment of the present application.
[0026] It is important to understand that in the field of gene sequencing data analysis, the most commonly used sequence alignment and variation detection tools include BLAST (Basic Local Alignment Search Tool), BWA (Burrows-Wheeler Aligner) and GATK (Genome Analysis Toolkit). BLAST uses a seed extension strategy for sequence alignment and uses a fixed-length seed sequence (usually 11bp) as an anchor point. Although it has advantages in computational efficiency, its scoring matrix uses a fixed BLOSUM or PAM matrix and cannot dynamically adjust parameters according to sequence specificity, which leads to false positives when processing highly variable regions. BWA mainly relies on the Burrows-Wheeler transformation for sequence indexing, and uses a unified quality threshold (usually Q20 or Q30) for filtering during sequence preprocessing, failing to consider the impact of local sequence features on quality assessment. Although GATK introduced a Bayesian model-based variation detection algorithm during variation detection, its variation site location mainly relied on simple mismatch rate statistics, and it often adopted independent assumptions when dealing with adjacent variations, making it unable to effectively identify and analyze complex associated variation patterns.
[0027] In contrast, the embodiments of the present application have carried out targeted technical improvements and parameter optimization in multiple key links. In the sequence preprocessing stage, unlike the fixed quality threshold strategy of BWA, the embodiments of the present application introduced an adaptive window mechanism based on local GC content. The window size is dynamically adjusted within the range of 30-45bp, and the Phred algorithm is combined to perform targeted corrections to bases with low quality values, which significantly improves the accuracy of data preprocessing. At the level of sequence alignment algorithms, it breaks through the limitations of the BLAST fixed scoring matrix and constructs an adaptive scoring system based on deep learning. The system integrates 23-dimensional feature information including nucleotide type, local GC content, secondary structure prediction, etc. through a three-layer neural network (128-64-32 node structure), and realizes accurate scoring of different conservative regions. Especially when dealing with highly variable regions, the system can be based on the information entropy of the site. and evolutionary conservation score Match scores and mismatch penalties are automatically adjusted to provide more accurate sequence alignments than a fixed scoring matrix.
[0028] In the variation detection and classification link, the embodiment of the present application overcomes the shortcomings of tools such as GATK in processing complex variation patterns. By introducing a scanning mechanism with a 23bp window size and a 7bp step length, combined with the dynamic adjustment of the GC content threshold by the Gaussian mixture model, high-precision positioning of potential variation regions is achieved. In particular, when processing adjacent variations, a 3×4 base conversion matrix and an associated variation analysis mechanism based on a 34bp distance threshold are designed. When the co-occurrence frequency between variations exceeds 0.567 and the sequence feature similarity is higher than 0.823, the system can automatically identify it as an associated variation and process it uniformly. This method significantly improves the ability to recognize complex mutation patterns. In the final pattern clustering analysis, a 23-dimensional feature vector including 3-dimensional one-hot encoding of variation types, position information, length features, etc. is constructed, and a distance weight mechanism based on sequence similarity is introduced (the weight increases by 1.2 times when the similarity is greater than 0.845), which achieves more accurate classification and identification of variation patterns. These technological innovations not only overcome the limitations of existing tools in processing complex variation patterns, but also disclose a faster and more reliable framework for analyzing gene sequencing data.
[0029] like Figure 1As shown, at step S101, the target gene sequence obtained by sequencing is preprocessed, including sequence reading, quality assessment, pairing splicing, and length screening to generate a preprocessed sequence file. This includes: reading the gene sequence file output by the sequencer, and establishing a corresponding relationship array between the base sequence and the quality value; using an adaptive sliding window to assess the sequence quality, marking the area with a quality value lower than a preset threshold as a low-quality area; using pairing information to splice the paired sequences, setting the minimum overlap length and the allowed mismatch rate; performing length screening on the spliced sequence, and using a dynamic programming algorithm to remove the connector sequences at both ends of the sequence; writing the preprocessed sequence into a temporary file and converting it into a target format.
[0030] In one embodiment, data preprocessing is performed on the target gene sequence obtained by sequencing. The FASTQ format file output by the sequencer is read into the memory, and a corresponding relationship array between the base sequence and the quality value is established. For each sequence, the three fields of its ID identification, base sequence, and quality string are extracted and stored in a string array of length N, respectively, where N is the total number of sequences. The sequence quality is evaluated by the adaptive sliding window method, and the window size is dynamically adjusted within the range of 30-45bp according to the local GC content, and the step length is correspondingly floating within the range of 3-7bp. When the average quality value in the window is less than 23, the region is marked as a low-quality region. For sequences with a sequencing read length less than 67bp or a low-quality region accounting for more than 12% of the total length, it is filtered. The quality value of each base is calculated by the Phred algorithm, and the bases with a quality value less than 18 are corrected. Using double-end pairing information, paired sequences are spliced, and the minimum overlap length is set to 31bp during splicing, and the allowed mismatch rate is 0.8-0.9%. The spliced sequence is length-screened to retain sequence fragments with a length within the range of 89-457bp. A dynamic programming algorithm was used to remove the adapter sequences at both ends of the sequence, and to remove regions in the sequence where more than 4 Ns appeared consecutively. The preprocessed sequences and their quality information were written to a temporary file in binary format, and the file header recorded statistical information such as the number of sequences and length interval. The processed sequences were converted to FASTA format, and a unique identifier was assigned to each sequence.
[0031] At step S102, the reference sequence library is analyzed based on the preprocessed sequence file, an adaptive scoring matrix is constructed, and the site score is dynamically adjusted using neural network training. This includes: reading the preprocessed sequence file, analyzing the gene sequence in the reference sequence library, counting the base occurrence frequency of each site, and generating a frequency matrix; calculating the information entropy of each site based on the frequency matrix, and determining the conservative score of the site according to the information entropy; dividing the sequences in the reference sequence library into a training set and a validation set according to a preset ratio, and performing a multiple sequence alignment on the training set; extracting the feature vector of each site in the multiple sequence alignment result, and inputting it into the neural network for training; dynamically adjusting the site-specific score based on the trained neural network output; determining the conservative level according to the site-specific score, and adopting a differentiated scoring strategy for regions with different conservative levels.
[0032] Specifically, an improved sequence alignment algorithm model was constructed. By analyzing 1000 known E2 gene sequences in the reference sequence library, the base occurrence frequency of each site was counted to generate a 4×L frequency matrix, where L is the sequence length. For each site i, its information entropy was calculated. ,in is the frequency of occurrence of base j at site i. The conservation score of the site is determined according to the size of the information entropy. Based on the Smith-Waterman local alignment algorithm, a machine learning-driven adaptive scoring matrix is constructed. By analyzing the E2 gene sequence population in the reference sequence library, a deep learning method is used to calculate the dynamic conservation weight for each site. For highly conserved regions, the alignment parameters are dynamically adjusted according to sequence statistics, with the match score range between 2.0-2.5, the mismatch penalty between -1.5 and -2.0, and the gap penalty between -1.8 and -2.3; for variable regions, the system calculates the corresponding parameter interval through the sequence evolution model. The dynamic programming score matrix is stored in a sparse matrix, and only non-zero elements with scores greater than the threshold are recorded. The compressed row storage format is used, and each non-zero element records its column index and score value. By maintaining the index range of the active calculation band, the invalid calculation in the matrix filling process is reduced. The minimum similarity threshold of the sequence alignment is set to 87.3%, and the maximum allowable gap ratio is 3.2%. During the alignment process, an improved backtracking algorithm was used to optimize the search space by setting branch and bound conditions. When the local alignment score was lower than the threshold of 13.7, direct pruning was performed. For sequences longer than 234 bp, a segmented alignment strategy was adopted to divide the sequence into multiple overlapping subsequences for alignment. The subsequence length was 156 bp, and the overlapping length of adjacent subsequences was 47 bp. The pre-processed temporary files were read by memory mapping to achieve efficient data access.
[0033] Furthermore, the information entropy of each site is calculated based on the frequency matrix, and the conservation score of the site is determined according to the information entropy, including: calculating the logarithm of the base occurrence frequency of each site, and summing the product of the logarithm and the base frequency of the site; taking the negative value of the cumulative sum to obtain the information entropy value of the site; comparing the information entropy value of the site with a preset reference value to determine the conservation score of the site.
[0034] Based on the output of the trained neural network, the site-specific score is dynamically adjusted, including: inputting the feature vector of the site into the trained neural network to obtain the predicted probability distribution vector of the occurrence of various nucleotides at the site; calculating the logarithmic value of the predicted probability of each nucleotide in the predicted probability distribution vector, and summing the product of the logarithmic value and the predicted probability to take the negative value; searching for homologous sequences in the evolutionary database and performing multiple comparisons on the obtained homologous sequences; counting the occurrence frequency of each nucleotide in the multiple comparison results, calculating the sum of the logarithmic product of the frequencies, and obtaining an evolutionary conservation score; and performing a weighted combination of the predicted probability distribution calculation result and the evolutionary conservation score to obtain an adjusted value of the site-specific score.
[0035] The conservation level is determined according to the site-specific score, and a differentiated scoring strategy is adopted for regions with different conservation levels, including: comparing the evolutionary conservation score of the site with a preset first threshold, and comparing the information entropy with a preset second threshold; dividing the site into a highly conserved region, a variable region or a moderately conserved region according to the comparison results; setting a first set of benchmark values for matching scores, mismatch deductions and gap penalties for highly conserved regions; setting a second set of benchmark values for matching scores, mismatch deductions and gap penalties for variable regions; adopting a weighted average of the first and second sets of benchmark values for moderately conserved regions; calculating the probability distribution of each type of substitution, and correcting the benchmark values of each conservation level region according to the substitution probability distribution to obtain the final score.
[0036] In one embodiment, the scoring matrix of the Smith-Waterman local alignment algorithm is first constructed. The E2 gene sequences in the reference sequence library are randomly divided into a training set and a validation set at a ratio of 78.9%, and the sequences in the training set are subjected to multiple sequence alignment to generate an initial position-specific scoring matrix. The site-specific score is obtained by calculating the nucleotide frequency distribution at each position.
[0037] ,in represents the frequency of nucleotide j at position i, Indicates the background frequency of the nucleotide. For any two sequence alignment positions (i, j), the initial score is calculated by The nucleotide composition feature vector within the 29bp window of the local sequence context is integrated. For each site, features including nucleotide type, local GC content, and secondary structure prediction score are extracted to construct a 23-dimensional feature vector. These feature vectors are input into a feedforward neural network with three hidden layers, with 128, 64, and 32 neurons in the hidden layers, respectively, and a ReLU activation function is used. The network is trained by minimizing the cross entropy loss function on the validation set, and the Adam optimizer is used for parameter update. The initial learning rate is set to 0.00234, and the training is stopped when the relative change rate of the validation set loss value for 5 consecutive epochs is less than 0.00127.
[0038] After obtaining the trained neural network, the site-specific score is dynamically adjusted using the network output. For each site i, its feature vector is input into the trained neural network to obtain the probability distribution vector p of the output layer. i Based on this probability distribution, the information entropy of the site is calculated ,in Indicates the predicted probability of the kth nucleotide appearing at site i. Combine information entropy with sequence evolution information, search for homologous sequences in the evolution database through BLAST, and set the E-value threshold to 1e-12. Perform multiple alignments on the obtained homologous sequences and calculate the evolutionary conservation score ,in is the frequency of the kth nucleotide in the multiple alignments. Based on the information entropy and evolutionary conservation score, the conservation level of the site is calculated: Below the threshold of 0.537 and H i When it is lower than 0.682, the site is marked as a highly conserved region; Above 0.892 or H i When the value was higher than 1.237, the site was marked as a variable region; the remaining sites were marked as moderately conserved regions.
[0039] Differentiated scoring strategies are used for regions with different conservation levels. For highly conserved regions, the match score benchmark value is set to 2.37, the mismatch penalty benchmark value is -1.83, and the gap penalty benchmark value is -2.13. For variable regions, the benchmark values are adjusted to 1.73, -1.27, and -1.62, respectively. The benchmark value for moderately conserved regions is the weighted average of the two. On this basis, the benchmark value is fine-tuned based on the evolutionary pattern information of the site. By analyzing the homologous sequences in the evolutionary database, the probability distribution of the substitution pattern for each site is calculated. , which represents the probability that nucleotide j mutates to k. For any alignment position (i, j), the actual score is calculated as follows: If it is a matching site, score = ; If it is a mismatch site, score = ; If it is an empty space, penalty = ,in is the adjustment coefficient, represents the probability of nucleotide j being missing.
[0040] At step S103, an adaptive scoring matrix is used to scan sequence differences using a sliding window and combined with Gaussian mixture model analysis to generate a variant region information file. This includes: based on a differentiated scoring strategy, a sliding window is used to scan sequence differences, and regions with mismatch rates exceeding a preset threshold are marked as potential variant regions; the context of the previous and next sequences of the potential variant regions is extracted to calculate local background features; the local background features are used to determine the content threshold interval using a Gaussian mixture model, and the sequence alignment information is integrated to evaluate conservation; the secondary structure stability of the sequence is calculated for the potential variant region, and the free energy change value is used as a feature; based on the features, multiple closely adjacent variant sites are merged and analyzed to generate a variant region information file.
[0041] For the location of sequence variation sites. Based on the alignment results, the sliding window method was used to scan the sequence differences, with the window size set to 23bp and the step length set to 7bp. When the mismatch rate within the window exceeded 6.8%, the region was marked as a potential variation region. For each potential variation region, the sequence context of 50bp before and after was extracted. The local background features of this 101bp length sequence fragment were calculated, including GC content distribution, repetitive sequence density, etc. The Gaussian mixture model was used to dynamically determine the GC content threshold interval based on the local sequence background, and the sequence alignment information between species was integrated to evaluate evolutionary conservation. The sequence was divided with a step length of 3bp, the number of occurrences of all possible triplets was counted, and the complexity score was obtained after normalization. Regions with a complexity greater than 0.783 and meeting the dynamic GC content threshold were determined as reliable variation regions. The information of each determined variation region was stored in the variation feature array, and each array element contained the starting position, end position, reference sequence fragment, and target sequence fragment of the variation region. For each variation region detected, the secondary structure stability of its local sequence was calculated, and the free energy change value was used as one of the important features of the variation site. For multiple closely adjacent variant sites, if the distance between them is less than 12bp, they are merged into one variant region for analysis. All identified variant region information is written into the intermediate file to prepare for subsequent classification and annotation.
[0042] At step S104, the variant region information file is read, a base conversion matrix is constructed for sequence alignment, and a variant classification result is generated by analyzing the alignment path. This includes: reading the variant region information file, constructing a base conversion matrix, and determining the variant type by sequence alignment; based on the base conversion matrix, aligning the reference sequence and the target sequence of the variant region to obtain the optimal alignment path; dividing the variant region into three types of substitution, insertion or deletion according to the optimal alignment path; calculating the feature vectors of the before and after sequences of the divided variant region, including base composition and structural characteristics; analyzing the sequence fragments between the variant regions, determining the associated variation, and generating the variant classification result.
[0043] In one embodiment, the identified mutations are classified and annotated. The mutation region information file generated in the previous step is read to construct a 3×4 base conversion matrix, where the rows represent the original bases A / T / C / G and the columns represent the mutated bases. For each mutation region, the specific mutation type is determined by sequence alignment: first, a local alignment matrix is constructed, and the matching score is set to 1.0, the mismatch deduction is -0.7, the gap start penalty is -1.2, and the gap extension penalty is -0.4. Based on this scoring scheme, the reference sequence and the target sequence of the mutation region are accurately aligned to obtain the optimal alignment path. According to the matching, mismatching and gap conditions in the alignment path, the mutation region is divided into three basic types: substitution, insertion or deletion. For each mutation region, record its mutation type, starting position, mutation length, affected base sequence and other information. At the same time, the characteristic vector of the 21bp sequence before and after the mutation region is calculated, including information such as base composition, secondary structure characteristics and local complexity. When the distance between two variant regions is less than 34 bp, the sequence fragments between them are extracted to analyze the conservation and functional characteristics of the sequences. If the co-occurrence frequency exceeds 0.567 and the sequence feature similarity is higher than 0.823, these variants are marked as associated variants and their feature information is merged. The classification and annotation results are saved in a structured format, including multiple data tables such as basic variant information, sequence features, and association relationships.
[0044] At step S105, the feature vector of the variant region is standardized based on the variant classification result, and the mutation pattern type is generated by the improved clustering algorithm. This includes: reading the variant classification result, constructing a feature vector for each variant region, including the features of the variant type code, position value, and length value; standardizing the feature vector so that the mean and standard deviation of each dimension meet the preset conditions; using the improved clustering algorithm to group based on the standardized feature vector, introducing the distance weight based on sequence similarity; determining the convergence condition of the clustering grouping result through iterative optimization to determine the mutation pattern type.
[0045] For the clustering analysis process of establishing the variation pattern. The variation classification results of the previous step can be read to construct a 23-dimensional feature vector for each variation region: the variation type is converted into a 3-dimensional one-hot encoding, the normalized relative position value occupies 1 dimension, the logarithm of the variation length occupies 1 dimension, the local GC content occupies 1 dimension, the sequence complexity occupies 1 dimension, and the base composition of the 8bp context before and after is represented by 16 dimensions. All feature vectors are standardized so that the mean of each dimension is 0 and the standard deviation is 1. The improved K-means clustering algorithm is used for grouping, and the initial number of cluster centers is set to 4. In the clustering process, the distance weight based on sequence similarity is introduced, that is, when the sequence context similarity of two variation regions is higher than 0.845, their distance weight in the feature space is increased by 1.2 times. Through the iterative optimization process, when the ratio of the distance between classes to the distance within the class is less than 0.0023, and the change in the position of the cluster center for three consecutive iterations is less than 0.0067, the clustering is considered to have reached convergence. According to the clustering results, the sequence features and associated mutation information of the variant regions in each class were integrated to generate a feature description summary of the class. The typical variant sequences in each cluster were compared with the known patterns in the reference database, and the local sequence alignment algorithm was used with a matching threshold of 0.8-0.85 to determine the type of mutation pattern represented by each cluster.
[0046] Figure 2 The following is a schematic diagram of the conservation analysis of the E2 gene sequence disclosed in the embodiment of the present application. As shown in the figure, the schematic diagram of the conservation analysis of the E2 gene sequence includes three interrelated graphic components: upper, middle and lower. Among them, the horizontal axis uses the sequence position value, ranging from 0 to 450.
[0047] In the first graphic component above, the vertical axis represents the base frequency, with a value range of 0-2.00. The four bases A, T, C, and G are marked in red, green, blue, and light blue respectively. The data shows that in the three sequence intervals of 50-150, 180-280, and 300-400, the frequency of A base is stable at 1.00-1.25, while the frequencies of other bases are relatively low; in the remaining intervals, the frequencies of the four bases fluctuate significantly, with the maximum fluctuation reaching 0.75.
[0048] In the second graphic component in the middle, the ordinate represents the local eigenvalue, which ranges from 0 to 2.0. The black curve shows that the local eigenvalue is stable at around 0.5 in the highly conserved region, while it rises to above 1.5 in the variable region. The background area is divided into four colors: gray, green, orange, and pink, corresponding to the sequence intervals of 0-100, 100-200, 200-300, and 300-400, respectively. The sequencing data show that when the eigenvalue is less than 0.5, the sequence conservation of the corresponding region reaches more than 87.3%.
[0049] In the third graphic component below, the ordinate represents the conservation score, which ranges from 0 to 1.0. The purple solid line represents the site-specific score, and the degree of conservation is divided into three intervals by the red dotted line (0.537) and the green dotted line (0.892). The data show that in the sequence intervals of 100-150, 200-250, 300-350, the conservation score is lower than 0.6, while in other intervals it is generally higher than 0.8. When the conservation score is higher than 0.892, the site is marked as a highly conserved region; when the conservation score is lower than 0.537, the site is marked as a variable region; when the score is between the two, it is marked as a moderately conserved region.
[0050] Figure 3 A sequencing quality distribution comparison diagram disclosed in the embodiment of the present application. The diagram contains three sub-diagrams: a sequencing quality interval distribution box plot at the top, a GC content and quality score correlation diagram at the bottom left, and a dynamic step density distribution diagram at the bottom right.
[0051] As shown in the box plot, the horizontal axis is the sequencing read length interval (bp), and the sequence is divided into five sections according to 89-156bp, 156-234bp, 234-347bp, 347-457bp, and 457bp+; the vertical axis is the Phred quality score. Among them, the blue box line data group corresponds to the Smith-Waterman traditional method, the green box line data group corresponds to the adaptive sliding window method of the embodiment of the present application, and the red dotted line is located at the Phred quality score of 23.7. Through the adaptive sliding window method of the embodiment of the present application, the window size is dynamically adjusted within the range of 30-45bp, and the step size is correspondingly floated within the range of 3-7bp, so that the Phred quality score of each read length interval is significantly higher than the quality threshold line 23.7. The specific values are as follows: the median quality score of the 89-156bp range increased from 20.5 to 24.7, the 156-234bp range increased from 21.8 to 26.5, the 234-347bp range increased from 23.2 to 27.8, the 347-457bp range increased from 25.6 to 29.3, and the range above 457bp increased from 27.9 to 31.2.
[0052] As shown in the GC content correlation diagram on the lower left, the horizontal axis is the GC content interval (%), ranging from 30-60%; the left vertical axis is the quality score (bp). The blue line is the baseline, and the green line and its light-colored envelope area represent the quality distribution of the adaptive method of the embodiment of the present application. When the GC content is in the range of 35-40%, the quality score is at a low point, with the lowest point corresponding to a score of 33.5 at a GC content of 42%; then it increases with the increase of GC content, and the quality score reaches 37.5 at 55%.
[0053] As shown in the step length distribution density diagram on the lower right, the horizontal axis is the dynamic step length range (bp) and the vertical axis is the frequency density. The curve is a single-peak distribution with a peak at 5bp, corresponding to a frequency density of 0.8. The frequency density values in the 3-4bp and 6-7bp intervals are relatively low, showing a symmetrical downward trend.
[0054] Furthermore, an embodiment of the present application also discloses a pattern recognition device for medical gene sequencing data, comprising: a processor, a memory, and a system bus; the processor and the memory are connected via the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, which, when executed by the processor, enable the processor to execute any of the above methods.
[0055] In summary, the embodiment of the present application introduces an adaptive sliding window mechanism in the sequence preprocessing stage. The window size is dynamically adjusted within the range of 30-45bp according to the local GC content, and the step size fluctuates within the range of 3-7bp accordingly, which significantly improves the accuracy of quality control. Secondly, in terms of the sequence alignment algorithm, a machine learning-driven adaptive scoring matrix is constructed, and the dynamic adjustment of site-specific scores is achieved through a three-layer neural network structure (the number of hidden layer neurons is 128, 64, and 32, respectively). The network implements a differentiated scoring strategy for different conservative regions by integrating 23-dimensional features such as nucleotide type, local GC content, and secondary structure prediction score. Specifically, the system is based on the information entropy (H) of the site. i ) and evolutionary conservation score (C i ) Dynamically adjust the scoring parameters. i Below 0.537 and H i When the score was lower than 0.682, the site was identified as a highly conserved region, with a higher match score (2.37) and a stricter mismatch penalty (-1.83). i Above 0.892 or H i If the value is higher than 1.237), a more flexible parameter setting is adopted.
[0056] In terms of variant site detection and classification, the embodiment of the present application has established a multi-level analysis framework. The sequence differences are scanned by the sliding window method (window size 23bp, step length 7bp), and the GC content threshold interval is dynamically determined in combination with the Gaussian mixture model, so as to achieve accurate positioning of potential variant regions. In the process of variant classification, a 3×4 base conversion matrix and an associated variant analysis mechanism are introduced, and accurate identification of complex mutation patterns is achieved by analyzing the distance (threshold 34bp) and co-occurrence features (frequency threshold 0.567) between variant regions. Finally, in the pattern clustering analysis link, a 23-dimensional feature vector system is designed, which includes 3-dimensional one-hot encoding, position information, length features and sequence context information of the variant type, and accurate classification of the variant pattern is achieved through the improved K-means algorithm, in which a distance weight mechanism based on sequence similarity is introduced. When the sequence context similarity is higher than 0.845, the feature space distance weight increases by 1.2 times, which significantly improves the accuracy of clustering. These technical improvements effectively solve the limitations of traditional methods in dealing with complex variant patterns, and achieve more accurate and rapid analysis and identification of medical gene sequencing data.
[0057] Furthermore, an embodiment of the present application also discloses a computer program product, which, when executed on a terminal device, enables the terminal device to execute any of the above methods.
[0058] It can be known from the description of the above implementation mode that those skilled in the art can clearly understand that all or part of the steps in the above-mentioned embodiment method can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a storage medium such as ROM / RAM, a disk, an optical disk, etc., including several instructions for enabling a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.
[0059] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0060] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0061] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A pattern recognition method for medical gene sequencing data, characterized in that: The following steps are involved: The target gene sequence obtained by sequencing is preprocessed, including sequence reading, quality assessment, pairing splicing, length screening, and generating a preprocessed sequence file; wherein, the quality assessment includes: in the sequence preprocessing stage, based on the adaptive window mechanism of local GC content, the window size is dynamically adjusted within the range of 30-45bp, and the bases with low quality values are corrected in a targeted manner in combination with the Phred algorithm; Analyze the reference sequence library based on the preprocessed sequence file, build an adaptive scoring matrix, and use neural network training to dynamically adjust the site scores, including: Read the preprocessed sequence file, analyze the gene sequence in the reference sequence library, count the base occurrence frequency of each site, and generate a frequency matrix; The information entropy of each site is calculated based on the frequency matrix, and the conservation score of the site is determined according to the information entropy; Sequences in the reference sequence library are divided into a training set and a validation set according to a preset ratio, and multiple sequence alignment is performed on the training set; Extract the feature vector of each site in the multiple sequence alignment results and input it into the neural network for training; Based on the trained neural network output, the site-specific score is dynamically adjusted; The conservation level is determined based on the site-specific score, and a differentiated scoring strategy is used for regions with different conservation levels; The steps of calculating the information entropy of each site based on the frequency matrix and determining the conservation score of the site according to the information entropy include: Calculate the logarithm of the base frequency at each site, and add up the product of the logarithm and the base frequency at the site; Take the negative value of the cumulative sum to get the information entropy value of the site; Compare the site information entropy value with the preset reference value to determine the conservation score of the site; Using an adaptive scoring matrix, a sliding window was used to scan sequence differences and combined with Gaussian mixture model analysis to generate a variant region information file, including: introducing a scanning mechanism with a 23bp window size and a 7bp step length, and combining the Gaussian mixture model to dynamically adjust the GC content threshold; Read the variant region information file, construct a base conversion matrix for sequence alignment, and generate variant classification results by analyzing the alignment path; Based on the mutation classification results, the feature vector of the mutation region is standardized, and the mutation pattern type is generated through the improved clustering algorithm, including: Read the mutation classification results and construct a feature vector for each mutation region, including the features of mutation type code, position value, and length value; Standardize the feature vector so that the mean and standard deviation of each dimension meet the preset conditions; Based on the standardized feature vectors, an improved clustering algorithm is used for grouping, and a distance weight based on sequence similarity is introduced; The convergence condition is determined through iterative optimization of the clustering grouping results to determine the mutation pattern type; Among them, the mutation classification results of the previous step are read, and a 23-dimensional feature vector is constructed for each variant region: the mutation type is converted into a 3-dimensional one-hot encoding, the normalized relative position value occupies 1 dimension, the logarithm of the mutation length occupies 1 dimension, the local GC content occupies 1 dimension, the sequence complexity occupies 1 dimension, and the base composition of the 8bp context before and after is represented by 16 dimensions; all feature vectors are standardized so that the mean of each dimension is 0 and the standard deviation is 1; the improved K-means clustering algorithm is used for grouping, and the initial number of cluster centers is set to 4. In the clustering process, a distance weight based on sequence similarity is introduced, that is, when the sequence context of two variant regions is When the similarity is higher than 0.845, their distance weight in the feature space increases by 1.2 times; through the iterative optimization process, when the ratio of the inter-class distance to the intra-class distance changes less than 0.0023, and the change in the cluster center position for three consecutive iterations is less than 0.0067, the clustering is considered to have reached convergence; according to the clustering results, the sequence characteristics and associated variation information of the variant regions in each class are integrated to generate a feature description summary of the class, and the typical variant sequences in each cluster are compared with the known patterns in the reference database. The local sequence alignment algorithm is used to set the matching threshold to 0.8-0.85 to determine the type of mutation pattern represented by each cluster.
2. The pattern recognition method for medical gene sequencing data according to claim 1, characterized in that: in, The target gene sequence obtained by sequencing is preprocessed, including sequence reading, quality assessment, pairing splicing, length screening, and the steps of generating a preprocessed sequence file include: Read the gene sequence file output by the sequencer and establish a corresponding relationship array between the base sequence and the quality value; The sequence quality is evaluated using an adaptive sliding window, and regions with quality values below a preset threshold are marked as low-quality regions; Use the pairing information to splice the paired sequences, set the minimum overlap length and allowable mismatch rate; The length of the spliced sequences was screened, and the adapter sequences at both ends of the sequences were removed using a dynamic programming algorithm; Write the preprocessed sequence to a temporary file and convert it to the target format.
3. The pattern recognition method for medical gene sequencing data according to claim 1, characterized in that: in, The steps of using an adaptive scoring matrix, scanning sequence differences with a sliding window and combining Gaussian mixture model analysis to generate a variant region information file include: Based on the differential scoring strategy, a sliding window is used to scan sequence differences, and regions with mismatch rates exceeding a preset threshold are marked as potential variant regions; Extract the previous and next sequence contexts for the potential variable region and calculate the local background features; Using local background features, a Gaussian mixture model was used to determine the content threshold interval, and sequence alignment information was integrated to evaluate conservation; The secondary structure stability of the sequence was calculated for the potential variable region, and the free energy change value was used as a feature; Based on the features, multiple closely adjacent variant sites are merged and analyzed to generate a variant region information file.
4. The pattern recognition method for medical gene sequencing data according to claim 1, characterized in that: in, The steps of reading the variant region information file, constructing a base conversion matrix for sequence alignment, and generating variant classification results by analyzing the alignment path include: Read the variant region information file, construct the base conversion matrix, and determine the variant type through sequence alignment; Based on the base conversion matrix, the reference sequence and target sequence of the variable region are aligned to obtain the optimal alignment path; The variant regions were divided into three types: substitution, insertion or deletion according to the optimal alignment path; Calculate the feature vectors of the preceding and following sequences for the divided variant regions, including base composition and structural features; Analyze the sequence fragments between the variant regions, determine the associated variants, and generate variant classification results.
5. The pattern recognition method for medical gene sequencing data according to claim 1, characterized in that: in, The step of dynamically adjusting the site-specific score based on the trained neural network output comprises: The feature vector of the site is input into the trained neural network to obtain the predicted probability distribution vector of the occurrence of various nucleotides at the site; Calculate the logarithm of the predicted probability of each nucleotide in the predicted probability distribution vector, and sum the product of the logarithm and the predicted probability and take the negative value; Search homologous sequences in evolutionary databases and perform multiple comparisons on the homologous sequences obtained; The frequency of occurrence of each nucleotide in the multiple alignment results was counted, and the sum of the logarithmic products of the frequencies was calculated to obtain the evolutionary conservation score; The predicted probability distribution calculation results are weighted combined with the evolutionary conservation score to obtain the adjusted value of the site-specific score.
6. The pattern recognition method for medical gene sequencing data according to claim 1, characterized in that: in, The steps of determining the conservation level according to the site-specific score and adopting a differentiated scoring strategy for regions with different conservation levels include: Comparing the evolutionary conservation score of the site with a preset first threshold, and comparing the information entropy with a preset second threshold; Based on the comparison results, the sites were classified as highly conserved regions, variable regions, or moderately conserved regions; Setting the first set of benchmark values for match scores, mismatch penalties, and gap penalties for highly conserved regions; Setting a second set of benchmark values for matching scores, mismatch penalty scores, and gap penalty scores for the variable region; For moderately conservative areas, a weighted average of the first and second group benchmark values was used; The probability distribution of each type of replacement is calculated, and based on the replacement probability distribution, the benchmark value of each conservative level area is corrected to obtain the final score value.
7. A pattern recognition device for medical gene sequencing data, characterized in that: include: A processor, a memory, and a system bus; wherein the processor and the memory are connected via the system bus; The memory is used to store one or more programs, and the one or more programs include instructions, which, when executed by the processor, enable the processor to execute the pattern recognition method for medical gene sequencing data according to any one of claims 1-6.
Citation Information
Patent Citations
Method for predicting positions of human protein sub-cellular fractions
CN106778070A
Genome copy number variation detection integration algorithm
CN113270141A