Genotype completion method and system based on multi-reference gene template row-column collaborative attention mechanism, terminal and storage medium

Through the genotype completion method of the collaborative attention mechanism of the multi-reference gene template genotype template, the problem of accuracy and slow speed of low-frequency allele completion is solved, efficient genotype completion is achieved, and the reliability and accuracy of genome-wide correlation analysis is improved.

CN120299523APending Publication Date: 2025-07-11HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510379527.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing deep learning-based genotype completion method has poor accuracy and slow speed in low-frequency allele complementation, which affects the reliability and accuracy of subsequent genome-wide correlation analysis.

Method used

A genotype completion method based on the ranks and synergistic attention mechanism of multi-reference gene templates is adopted. Through multi-sequence sorting, position embedding, ranks and attention mechanism and gating mechanism, combined with linkage imbalanced biological prior knowledge, model structure and algorithm strategies are optimized to improve the completion performance of low-frequency alleles.

Benefits of technology

The completion accuracy and completion speed of low-frequency alleles are significantly improved, and the reliability and accuracy of genome-wide correlation analysis is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299523A_ABST
    Figure CN120299523A_ABST
Patent Text Reader

Abstract

The invention discloses a genotype completion method and system based on a multi-reference gene template row-column collaborative attention mechanism, a terminal and a storage medium. The method comprises the following steps: acquiring a gene sequence and a corresponding position sequence; on the basis of a query sequence obtained after the gene sequence mask processing, sorting reference sequences in the multi-reference gene template based on multi-sequence alignment according to multiple biological attributes to obtain an MSA matrix, and performing mapping processing on the MSA matrix to obtain a target MSA matrix; according to the target MSA matrix and the position sequence, training a genotype completion model based on a multi-reference gene template row-column collaborative attention mechanism to obtain a target genotype completion model; and completing the model by using the target genotype to obtain the genotype of the deletion site in the deletion sequence to be detected. By optimizing the model structure and improving the algorithm strategy, the completion accuracy and completion speed of the low-frequency alleles are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics technology, and particularly to a genotype completion method, system, terminal and computer-readable storage medium based on a multi-reference gene template row-column collaborative attention mechanism. Background Art

[0002] With the in-depth study of genomics, the identification and analysis of single nucleotide polymorphisms and insertions / deletions play an important role in fields such as genetics, medical research, and agricultural breeding. However, traditional sequence alignment and variant detection methods face many challenges when dealing with large-scale genomic data, especially in the accurate identification of variant sites and the accurate completion of deletion sites. In recent years, sequence completion methods based on deep learning have gradually become a research hotspot. However, some current genotype completion methods based on deep learning have the disadvantages of poor accuracy and slow completion speed for low-frequency alleles. Therefore, the current genotype completion methods still have deficiencies in the completion performance of low-frequency alleles, and this deficiency will affect the reliability and accuracy of subsequent tasks such as genome-wide association analysis.

[0003] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0004] The main purpose of the present invention is to provide a genotype completion method, system, terminal and computer-readable storage medium based on a multi-reference gene template row-column collaborative attention mechanism, aiming to solve the problems of poor accuracy and slow speed in the completion of low-frequency alleles by existing genotype completion methods based on deep learning.

[0005] To achieve the above-mentioned invention purpose, the present invention provides a genotype completion method based on a multi-reference gene template row-column collaborative attention mechanism. The genotype completion method based on a multi-reference gene template row-column collaborative attention mechanism includes:

[0006] Obtain a VCF (Variant Call Format) file, process the VCF file to obtain a gene sequence, and obtain a position sequence corresponding to the gene sequence;

[0007] Perform a masking process on the gene sequence to obtain a query sequence. Based on the query sequence, sort the reference sequences in the multi-reference gene template based on multiple biological attributes through multiple sequence alignment to obtain an MSA matrix, and perform a mapping process on the MSA matrix to obtain a target MSA matrix;

[0008] Construct a genotype completion model based on the row-column collaborative attention mechanism of multiple reference gene templates, and train the genotype completion model according to the target MSA matrix and the position sequence to obtain the target genotype completion model;

[0009] Obtain the sequence with missing sites to be tested, and input the sequence with missing sites to be tested into the target genotype completion model, and the target genotype completion model outputs the genotypes of the missing sites in the sequence with missing sites to be tested.

[0010] Optionally, the steps of obtaining the VCF file, processing the VCF file to obtain a gene sequence, and obtaining the corresponding position sequence of the gene sequence specifically include:

[0011] Obtain a VCF file, and parse the VCF file to obtain an initial gene sequence;

[0012] Perform normalization processing on the initial gene sequence by means of multi-condition filtering to obtain a gene sequence S', where the multi-condition filtering includes gene frequency screening and mutation type filtering, and the gene sequence S' is:

[0013] S'=(S1,S2,...,S i ,...,S m )∈R 1×m ;

[0014] where S i represents the element at the i-th site in the gene sequence S', R represents a one-dimensional matrix, and m represents the maximum sequence length of the input genotype completion model;

[0015] Obtain the corresponding position sequence P' of the gene sequence, and the position sequence P' is:

[0016] P'=(P1,P2,...,P i ,...,P m )∈R 1×m ;

[0017] where P i represents the position of the element S i at the i-th site on the chromosome.

[0018] Optionally, the steps of masking the gene sequence to obtain a query sequence, sorting the reference sequences in the multiple reference gene templates based on multiple sequence alignment according to multiple biological attributes based on the query sequence to obtain an MSA matrix, and performing mapping processing on the MSA matrix to obtain a target MSA matrix specifically include:

[0019] Randomly select some elements from the gene sequence according to the set ratio for masking to obtain the query sequence W;

[0020] W = (S1, S2,..., S i = [MASK],..., S m ) ∈ R 1×m ;

[0021] Among them, [MASK] represents the masking token;

[0022] Calculate the first similarity between the query sequence and each reference sequence in the multi-reference gene template respectively;

[0023] According to the first similarity between the query sequence and each reference sequence, sort the multiple reference sequences in descending order of the first similarity for the first time;

[0024] Concatenate prefix information to the query sequence and each reference sequence, where the prefix information includes multiple biological attributes, and the multiple biological attributes include kinship, species, and region;

[0025] Calculate the second similarity between the prefix information of the query sequence and the prefix information of each reference sequence respectively;

[0026] According to the second similarity between the prefix information of the query sequence and the prefix information of each reference sequence, sort the multiple reference sequences after the first sorting in descending order of the second similarity, where the priority of the first sorting is higher than that of the second sorting;

[0027] Select the first set number of reference sequences after the second sorting in the multi-reference gene template, and randomly select some elements of the first set number of reference sequences according to the set ratio for masking to obtain the first set number of masked reference sequences;

[0028] Combine the query sequence and the first set number of masked reference sequences to obtain the MSA matrix;

[0029] Map each element in the MSA matrix to a natural number through the mapping vocabulary to obtain the target MSA matrix;

[0030] Among them, the target MSA matrix includes the target query sequence and the target reference sequence.

[0031] Optionally, to construct a genotype completion model based on the row-column co-attention mechanism of the multi-reference gene template, and train the genotype completion model according to the target MSA matrix and the position sequence to obtain the target genotype completion model, specifically including:

[0032] Construct a genotype completion model based on a multi-reference gene template row-column collaborative attention mechanism, where the genotype completion model includes an embedding layer, a multi-sequence attention module, and a fully connected layer;

[0033] Input the target MSA matrix and the position sequence into the embedding layer, and the embedding layer performs initialization processing on the feature representation of the target MSA matrix to obtain an embedding representation matrix;

[0034] Adjust the embedding representation matrix using the similarity weight between the target query sequence and each target reference sequence to obtain an adjusted embedding representation matrix:

[0035] MSA s = MSA'·similarity_weight;

[0036] where MSA s represents the adjusted embedding representation matrix, MSA' represents the embedding representation matrix, and similarity_weight represents the similarity weight;

[0037] Input the adjusted embedding representation matrix into the multi-sequence attention module. The multi-sequence attention module combines the Transformer QKV mechanism, the gating mechanism, and the biological prior knowledge of linkage disequilibrium to perform row-column attention calculation, gating mechanism to adjust weights, and residual connection and normalization on the adjusted embedding representation matrix to obtain an output vector;

[0038] Input the output vector into the fully connected layer. The fully connected layer maps the output vector to a probability vector of each element in the corresponding mapping vocabulary, and takes the element corresponding to the maximum probability in the probability vector as the genotype of the missing site;

[0039] Compare the genotype of the missing site with the gene sequence to obtain a comparison result, and train the genotype completion model according to the comparison result to obtain a target genotype completion model.

[0040] Optionally, the step of inputting the target MSA matrix and the position sequence into the embedding layer, and the embedding layer performing initialization processing on the feature representation of the target MSA matrix to obtain an embedding representation matrix specifically includes:

[0041] Input the target MSA matrix and the position sequence into the embedding layer, and the embedding layer maps each natural number of the input sequence in the target MSA matrix to a high-dimensional vector to obtain a first embedding vector;

[0042] According to the position sequence, perform sine-cosine encoding on each natural number of the input sequence to obtain a second embedding vector;

[0043] Add the first embedding vector and the second embedding vector to obtain the embedding representation matrix corresponding to the target MSA matrix:

[0044] E = E token + E pos ;

[0045] where E represents the embedding representation matrix, E token represents the first embedding vector, and E pos represents the second embedding vector;

[0046] where the input sequence includes the target query sequence and the target reference sequence in the target MSA matrix.

[0047] Optionally, input the adjusted embedding representation matrix into the multi-sequence attention module. The multi-sequence attention module combines the Transformer QKV mechanism, the gating mechanism, and the biological prior knowledge of linkage disequilibrium to perform row-column attention calculation, gating mechanism to adjust weights, and residual connection and normalization on the adjusted embedding representation matrix to obtain an output vector, specifically including:

[0048] Input the adjusted embedding representation matrix into the multi-sequence attention module. The multi-sequence attention module combines the Transformer QKV mechanism, the gating mechanism, and the biological prior knowledge of linkage disequilibrium to perform row-column attention calculation on the adjusted embedding representation matrix, where the row-column attention calculation includes row attention calculation and column attention calculation:

[0049] The row attention calculation includes calculating the row attention score Attention row :

[0050]

[0051] Q = MSA s · W Q ;

[0052] K = MSA s · W K ;

[0053] V = MSA s · W V ;

[0054] Gate row = σ(MSA s );

[0055] Among them, Sigmoid(·) represents the activation function, Q represents the query vector, K represents the key vector, T represents the transpose, and bias LD represents the prior value of the linkage disequilibrium matrix, and d k represents the dimension of the key vector, V represents the value vector, and Gate row represents the gating weight of the column attention, and W Q represents the weight matrix corresponding to the query vector, and W K represents the weight matrix corresponding to the key vector, and W V represents the weight matrix corresponding to the value vector, and σ represents the Sigmoid function;

[0056] The column attention calculation includes calculating the column attention score Attention col :

[0057]

[0058] Gate col =σ(MSA s );

[0059] Among them, Gate col represents the gating weight of the column attention;

[0060] For low-frequency sites, the gating mechanism is used to increase the gating weight corresponding to the low-frequency sites;

[0061] Through residual connection, the input of each layer in the multi-sequence attention module is added to the output;

[0062] Through normalization, the output of each layer in the multi-sequence attention module is standardized to obtain the output vector.

[0063] Optionally, after obtaining the missing sequence to be measured and inputting the missing sequence to be measured into the target genotype completion model, and the target genotype completion model outputs the genotype of the missing site in the missing sequence to be measured, it further includes:

[0064] According to the genotype of the missing site in the missing sequence to be measured, the missing site in the missing sequence to be measured is completed.

[0065] To achieve the above-mentioned invention purpose, the present invention also provides a genotype completion system based on a multi-reference gene template row-column collaborative attention mechanism. The genotype completion system based on the multi-reference gene template row-column collaborative attention mechanism includes:

[0066] Sequence acquisition module: used to acquire a VCF file, process the VCF file to obtain a gene sequence, and acquire the position sequence corresponding to the gene sequence;

[0067] A target MSA matrix construction module is used to perform mask processing on the gene sequence to obtain a query sequence, and based on the query sequence, the reference sequences in the multiple reference gene templates are sorted based on multiple sequence alignment according to multiple biological attributes to obtain an MSA matrix, and the MSA matrix is mapped to obtain a target MSA matrix;

[0068] Genotype completion model training module: used to construct a genotype completion model based on a multi-reference gene template row-column collaborative attention mechanism, and train the genotype completion model according to the target MSA matrix and the position sequence to obtain a target genotype completion model;

[0069] Genotype completion module: used to obtain the missing sequence to be tested and input the missing sequence to be tested into the target genotype completion model, and the target genotype completion model outputs the genotype of the missing site in the missing sequence to be tested.

[0070] To achieve the above-mentioned purpose of the invention, the present invention also provides a terminal, which includes: a memory, a processor, and a genotype completion program based on a multi-reference gene template row-column collaborative attention mechanism stored in the memory and executable on the processor, wherein the genotype completion program based on a multi-reference gene template row-column collaborative attention mechanism implements the steps of the genotype completion method based on a multi-reference gene template row-column collaborative attention mechanism as described above when executed by the processor.

[0071] To achieve the above-mentioned purpose of the invention, the present invention also provides a computer-readable storage medium, which stores a genotype completion program based on a multi-reference gene template row-column collaborative attention mechanism. When the genotype completion program based on a multi-reference gene template row-column collaborative attention mechanism is executed by a processor, the steps of the genotype completion method based on a multi-reference gene template row-column collaborative attention mechanism as described above are implemented.

[0072] In the present invention, a VCF file is obtained, processed to obtain a gene sequence, and a position sequence corresponding to the gene sequence is obtained; the gene sequence is masked to obtain a query sequence, and based on the query sequence, the reference sequences in a multi-reference gene template are sorted based on multiple sequence alignment according to multiple biological attributes to obtain an MSA (Multiple Sequence Alignment) matrix, and the MSA matrix is mapped to obtain a target MSA matrix; a genotype completion model based on a row-column collaborative attention mechanism of a multi-reference gene template is constructed, and the genotype completion model is trained according to the target MSA matrix and the position sequence to obtain a target genotype completion model; a to-be-tested missing sequence is obtained and input into the target genotype completion model, and the target genotype completion model outputs the genotypes of the missing sites in the to-be-tested missing sequence. The present invention improves the multi-sequence sorting method, sorts gene sequence fragments based on multiple biological attributes for multi-sequence sorting and uses it as the model input; improves the position embedding method, and when the model is input, the position of the gene locus on the chromosome is used as the position embedding; uses row and column attention mechanisms to learn the connections between single fragment sites and between the same position sites of multiple similar fragments; in the process of row-column attention calculation, biological prior knowledge of linkage disequilibrium is added, and a gating mechanism is added after row-column attention calculation to use the gating mechanism to increase the gating weight of low-frequency sites, thereby increasing the attention score of low-frequency sites. By optimizing the model structure and improving the algorithm strategy, the present invention effectively improves the completion accuracy and completion speed of low-frequency alleles, significantly improves the completion performance of low-frequency alleles, realizes genotype completion on the whole genome sequence, and is beneficial to improving the reliability and accuracy of subsequent tasks such as whole genome association analysis. Description of the Drawings

[0073] Figure 1 is a flowchart of a preferred embodiment of the genotype completion method based on a row-column collaborative attention mechanism of a multi-reference gene template of the present invention;

[0074] Figure 2 is another flowchart of a preferred embodiment of the genotype completion method based on a row-column collaborative attention mechanism of a multi-reference gene template of the present invention;

[0075] Figure 3 is a schematic diagram of constructing an MSA matrix of the present invention;

[0076] Figure 4 is a structural diagram of a preferred embodiment of the genotype completion system based on a row-column collaborative attention mechanism of a multi-reference gene template of the present invention;

[0077] Figure 5It is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed implementation manners

[0078] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further elaborates on the present invention by way of examples with reference to the accompanying drawings. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0079] With the in-depth study of genomics, the identification and analysis of single nucleotide polymorphisms (SNPs, Single Nucleotide Polymorphism) and insertions and deletions (INDELs, Insertion-Deletion) play important roles in fields such as genetics, medical research, and agricultural breeding. However, traditional sequence alignment and variant detection methods face many challenges when dealing with large-scale genomic data, especially in the accurate identification of variant sites and the accurate filling of deletion sites. In recent years, sequence completion methods based on deep learning have gradually become a research hotspot. Among them, the Transformer architecture has received extensive attention due to its powerful sequence modeling ability, providing new possibilities for the analysis of genomic data.

[0080] In the field of deep learning, BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language representation model based on the Transformer architecture) captures context information in a sequence through a bidirectional encoder, thus achieving a significant performance improvement in natural language processing tasks. A key feature of the BERT model is its attention mechanism, which can dynamically focus on information at different positions in the sequence, thereby improving the model's representation ability. Therefore, BERT is applied to the field of genomics, and the BERT model is applied to the sequence completion task to learn the long-range dependencies in the sequence and effectively improve the accuracy of completion. However, although BERT shows potential in the sequence completion task, its most basic completion method still has limitations when dealing with SNPs / INDELs. A significant problem is that BERT often ignores the correlation between variant sites / deletion sites, namely linkage disequilibrium (LD). Linkage disequilibrium refers to the phenomenon that certain alleles tend to be inherited together, which is of great significance in genetic analysis. In variant detection and gene completion, considering LD information can improve the accuracy of variant site prediction and deletion site completion, that is, if LD information can be added during the attention calculation process, the accuracy of variant site prediction and deletion site completion can be improved. This idea provides an important direction for further improving the Transformer-based model.

[0081] The current genotype completion methods are mainly divided into two categories: the first category is the gene completion method based on the hidden Markov model or similar traditional computing models, which mainly models the linkage disequilibrium of gene sequences and infers the missing genotypes through the observed genotype fragments. Its advantages are high accuracy and flexibility, but it has high requirements for computing resources, and the effect depends on the quality of multiple reference gene templates and the degree of population matching. It is widely used in genetic research. The second category is the gene completion method based on deep learning models, but some current deep learning-based gene completion methods have poor completion performance on low-frequency alleles, which affects subsequent whole-genome association analysis and other work. In actual research, many genomic data come from low-coverage sequencing, and existing methods have poor robustness to low-coverage data. Due to the large number of genotype missing and error sites in low-coverage data, traditional methods are difficult to fully utilize the information of these data, which in turn affects the completion effect.

[0082] In order to solve the above technical problems, the present invention provides a genotype completion method based on a multi-reference gene template row-column collaborative attention mechanism, obtains a VCF file, processes the VCF file to obtain a gene sequence, and obtains a position sequence corresponding to the gene sequence; masks the gene sequence to obtain a query sequence, based on the query sequence, sorts the reference sequences in the multi-reference gene template based on a multiple sequence alignment according to multiple biological attributes to obtain an MSA matrix, and maps the MSA matrix to obtain a target MSA matrix; constructs a genotype completion model based on a multi-reference gene template row-column collaborative attention mechanism, trains the genotype completion model according to the target MSA matrix and the position sequence to obtain a target genotype completion model; obtains a missing sequence to be tested, and inputs the missing sequence to be tested into the target genotype completion model, and the target genotype completion model outputs the genotype of the missing site in the missing sequence to be tested. The present invention improves the multi-sequence sorting method, sorts the gene sequence fragments according to multiple biological attributes and uses them as model input; improves the position embedding method, uses the position of the gene site on the chromosome as the position embedding when the model is input; uses the row and column attention mechanism to learn the connection between single fragment sites and multiple similar fragments at the same position; in the row and column attention calculation process, adds the biological prior knowledge of linkage disequilibrium, and adds a gating mechanism after the row and column attention calculation, and uses the gating mechanism to improve the gating weight of low-frequency sites, thereby improving the attention score of low-frequency sites. The present invention effectively improves the completion accuracy and completion speed of low-frequency alleles by optimizing the model structure and improving the algorithm strategy, significantly improves the completion performance of low-frequency alleles, realizes genotype completion on the whole genome sequence, and is conducive to improving the reliability and accuracy of subsequent tasks such as whole genome association analysis.

[0083] The following further elaborates on the application content by describing the embodiments in conjunction with the accompanying drawings.

[0084] A preferred embodiment of the genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism of the present invention is as Figure 1 and Figure 2 shown, and specifically includes:

[0085] S1. Obtain a VCF file, process the VCF file to obtain a gene sequence, and obtain a position sequence corresponding to the gene sequence.

[0086] In an implementation manner of this embodiment, the obtaining of the VCF file, processing the VCF file to obtain a gene sequence, and obtaining a position sequence corresponding to the gene sequence specifically includes:

[0087] Obtain a VCF file and parse the VCF file to obtain an initial gene sequence;

[0088] Perform normalization processing on the initial gene sequence by means of multi-condition filtering to obtain a gene sequence S', where the multi-condition filtering includes gene frequency screening and mutation type filtering, and the gene sequence S' is:

[0089] S' = (S1, S2,..., S i ,..., S m ) ∈ R 1×m ;

[0090] where S i represents the element at the i-th site in the gene sequence S', R represents a one-dimensional matrix, and m represents the maximum sequence length of the input genotype completion model;

[0091] Obtain the position sequence P' corresponding to the gene sequence, and the position sequence P' is:

[0092] P' = (P1, P2,..., P i ,..., P m ) ∈ R 1×m ;

[0093] where P i represents the position of the element S i on the chromosome at the i-th site.

[0094] Specifically, first construct a pre-training dataset. Since the training data is chromosomal mutation sequences with lengths in the millions or even higher, it is necessary to truncate the chromosomal mutation sequences into equal-length segments during the construction of the pre-training dataset to obtain gene sequences.

[0095] S2. Mask the gene sequence to obtain a query sequence. Based on the query sequence, sort the reference sequences in the multi-reference gene template according to multiple biological attributes through multiple sequence alignment to obtain an MSA matrix, and perform a mapping process on the MSA matrix to obtain a target MSA matrix.

[0096] In an implementation manner of this embodiment, the steps of masking the gene sequence to obtain a query sequence, sorting the reference sequences in the multi-reference gene template according to multiple biological attributes through multiple sequence alignment to obtain an MSA matrix, and performing a mapping process on the MSA matrix to obtain a target MSA matrix specifically include:

[0097] Randomly select a part of the elements in the gene sequence for masking according to a set ratio to obtain a query sequence W;

[0098] W = (S1, S2,..., S i = [MASK],..., S m ) ∈ R 1×m ;

[0099] where [MASK] represents a mask flag;

[0100] Calculate the first similarity between the query sequence and each reference sequence in the multi-reference gene template respectively;

[0101] According to the first similarity between the query sequence and each reference sequence, sort the multiple reference sequences in descending order of the first similarity for the first time;

[0102] Concatenate prefix information to the query sequence and each reference sequence, where the prefix information includes multiple biological attributes, and the multiple biological attributes include kinship, species, and region;

[0103] Calculate the second similarity between the prefix information of the query sequence and the prefix information of each reference sequence respectively;

[0104] According to the second similarity between the prefix information of the query sequence and the prefix information of each reference sequence, sort the multiple reference sequences sorted for the first time in descending order of the second similarity. Among them, the priority of the first sorting is higher than that of the second sorting;

[0105] Select the first set number of reference sequences after the second sorting in the multi-reference gene template, and randomly select a part of the elements of the first set number of reference sequences for masking according to a set ratio to obtain the first set number of masked reference sequences;

[0106] Combine the query sequence and the pre-set value bar mask reference sequence to obtain an MSA matrix;

[0107] Map each element in the MSA matrix to a natural number through a mapping vocabulary to obtain a target MSA matrix;

[0108] Among them, the target MSA matrix includes a target query sequence and a target reference sequence.

[0109] Specifically, in order to train the model to have the ability to fill in missing sites, the input data will be masked before entering the embedding layer. To simulate the situation of missing sites in gene sequences, a part of the elements (tokens) in the sequence are randomly selected for masking at a ratio of 40%, and the original value is replaced with the marker [MASK]. During the construction of the pre-training dataset, a multiple sequence alignment operation is required, as Figure 3 shown. The specific steps include: First, randomly cover a part of the elements in the gene sequence at a ratio of 40%; then, calculate the similarity between the covered gene sequence (i.e., the query sequence) and each reference sequence in the multi-reference gene template (note: the covered element sites are not considered during the calculation), and sort the reference sequences of the multi-reference gene template from high to low according to the calculated similarity between the query sequence and each reference sequence; subsequently, splice the query sequence and the sorted multi-reference gene template with prefix information (including kinship, species or race, region or country); further sort according to the similarity of the prefix information while ensuring that the similarity sorting relationship of the reference sequences is not disrupted, that is, the first sorting priority is higher than the second-level sorting; finally, select the top n reference sequences with the highest similarity in the multi-reference gene template and randomly mask these reference sequences at a ratio of 40% to finally construct a multiple sequence alignment (MSA) matrix that meets the model input requirements. The multiple sequence alignment matrix obtained through this sorting method has a high similarity with the query sequence, and also has a high similarity in prefix information such as race information and kinship. Therefore, the chromosomal mutation information will also show a high similarity. This design can help the model better learn the mutation information, thereby improving the success rate of filling in missing sites.

[0110] Then, map each element in the MSA matrix to a natural number through a mapping vocabulary to obtain a target MSA matrix, that is, map each element in the target query sequence and the masked reference sequence in the target MSA matrix to a natural number to obtain a target query sequence and a target reference sequence. Specifically, perform vocabulary mapping on the elements (tokens) of the query sequence and the masked reference sequence, and convert them into natural numbers, that is, sequence data. The mapping vocabulary used is V, and part of the mapping vocabulary V consists of the following elements:

[0111] V = {[MASK], [PAD], 0, 1};

[0112] Among them, 0 represents that the mutation at this position has not occurred, 1 represents that the mutation at this position has occurred, [PAD] is used to pad the sequence with insufficient length and has no actual meaning, and [MASK] represents the missing mutation information at a certain site. In addition, the second half of the mapping vocabulary V also contains content related to whether there is a kinship, different superfamilies, and different homecomings. The size of the vocabulary V is denoted as v. Different elements tokens in the query sequence and the masked reference sequence will be mapped to natural number token_ids through the vocabulary V. Therefore, before inputting into the model, each element of the query sequence and the masked reference sequence in the MSA matrix is replaced by the corresponding token_id through the mapping of the vocabulary V, obtaining a matrix containing multiple sequences, that is, the target MSA matrix.

[0113] S3. Construct a genotype completion model based on the row-column collaborative attention mechanism of multiple reference gene templates, and train the genotype completion model according to the target MSA matrix and the position sequence to obtain the target genotype completion model.

[0114] In an implementation manner of this embodiment, the constructing a genotype completion model based on the row-column collaborative attention mechanism of multiple reference gene templates and training the genotype completion model according to the target MSA matrix and the position sequence to obtain the target genotype completion model specifically includes:

[0115] Construct a genotype completion model based on the row-column collaborative attention mechanism of multiple reference gene templates, where the genotype completion model includes an embedding layer, a multi-sequence attention module, and a fully connected layer;

[0116] Input the target MSA matrix and the position sequence into the embedding layer, and the embedding layer performs initialization processing on the feature representation of the target MSA matrix to obtain an embedding representation matrix;

[0117] Adjust the embedding representation matrix by using the similarity weight between the target query sequence and each target reference sequence to obtain an adjusted embedding representation matrix:

[0118] MSA s = MSA'·similarity_weight;

[0119] Among them, MSA s represents the adjusted embedding representation matrix, MSA' represents the embedding representation matrix, similarity_weight represents the similarity weight, which is a weight calculated based on the similarity between the target query sequence and the target reference sequence and is used to guide the model to pay more attention to sequences with high similarity;

[0120] Input the adjusted embedded representation matrix into the multi-sequence attention module. The multi-sequence attention module combines the Transformer QKV mechanism, the gating mechanism, and the biological prior knowledge of linkage disequilibrium to perform row-column attention calculation, adjust weights through the gating mechanism, and perform residual connection and normalization on the adjusted embedded representation matrix to obtain an output vector;

[0121] Input the output vector into the fully connected layer. The fully connected layer maps the output vector to a probability vector for each element in the corresponding mapping vocabulary, and takes the element corresponding to the maximum probability in the probability vector as the genotype of the missing site;

[0122] Compare the genotype of the missing site with the gene sequence to obtain a comparison result, and train the genotype completion model according to the comparison result to obtain a target genotype completion model.

[0123] Specifically, in the main model part, the entire framework completes feature extraction and representation learning of the multiple sequence alignment matrix (referring to the target MSA matrix) through the attention mechanism and the gating mechanism. Through the design based on the multi-sequence attention mechanism and the gating mechanism, the model realizes the accurate modeling of the complex relationship between the query sequence and the reference sequence (referring to the target query sequence and the target reference sequence). In particular, the learning ability of the mutation sites in the multiple sequence alignment matrix has been significantly improved, thereby significantly improving the completion performance of low-frequency alleles.

[0124] The entire model includes 12 layers of multi-sequence attention modules, and each layer includes three core steps: row-column attention calculation, weight adjustment through the gating mechanism, and residual connection and normalization. In the multi-sequence attention calculation, the model captures deeper context relationships between the query sequence and the reference sequence layer by layer, and optimizes the attention distribution by combining similarity weights. The gating mechanism dynamically adjusts the information interaction intensity of different sequences to further enhance the attention to key sequences. The design of residual connection and normalization ensures that the output of each layer can be stably fused with the input, thereby maintaining the stability of the gradient and accelerating the training convergence. The output of each layer will be used as the input of the next layer. Through layer-by-layer stacking and feature fusion, the model can gradually learn more complex sequence relationships and context features. After 12 layers of stacking, the final output of the model is a query sequence representation containing context information. In this representation, each token combines the mutation information from the reference sequence and the correlation features between sequences. This output can not only effectively capture the similarity between the query sequence and the reference sequence, but also provide high-quality feature support for subsequent tasks (such as predicting masked tokens or analyzing mutation functions). Taking the prediction of masked tokens as an example: the output of the multi-sequence attention module will finally pass through a fully connected layer. For each output vector o ∈ R 1×h It will transform it into a probability vector of each value of this mutation site (each vector) in the vocabulary. For a certain mutation site, the prediction result of the model is the token with the maximum probability in the probability vector.

[0125] In an implementation manner of this embodiment, when inputting the target MSA matrix and the position sequence into the embedding layer, the embedding layer performs initialization processing on the feature representation of the target MSA matrix to obtain an embedding representation matrix, which specifically includes:

[0126] Input the target MSA matrix and the position sequence into the embedding layer. The embedding layer maps each natural number (token_id) of the input sequence in the target MSA matrix to a high-dimensional vector to obtain a first embedding vector;

[0127] According to the position sequence, perform sine-cosine encoding on each natural number (token_id) of the input sequence to obtain a second embedding vector;

[0128] Add the first embedding vector and the second embedding vector to obtain the embedding representation matrix corresponding to the target MSA matrix:

[0129] E = E token + E pos ;

[0130] where E represents the embedding representation matrix, E token represents the first embedding vector, and E pos represents the second embedding vector;

[0131] The position embedding adopts sine-cosine encoding, which is expressed as:

[0132]

[0133] where i is the absolute position, k is the latitude index, and h is the vector dimension;

[0134] where the input sequence includes the target query sequence and the target reference sequence in the target MSA matrix.

[0135] Specifically, the role of the embedding layer of the model is to map the input token_id sequence (in the target MSA matrix, each row represents a token_id sequence, including the target query sequence and the target reference sequence) to a vector space of a fixed dimension, so as to obtain vector representations. These vector representations can gradually capture semantic information, syntactic information, and task-related features during the subsequent attention mechanism learning process. This vectorization process facilitates the model to more efficiently capture the correlations between different tokens during the attention calculation, thereby improving the effect and accuracy of the attention calculation. Based on the target multiple sequence alignment matrix, the model initializes the feature representation of the input sequences (i.e., the target query sequence and the target reference sequence, which are token_id sequences) through the embedding layer. The embedding layer mainly includes two parts: token embedding and position embedding. Token embedding maps each token_id in the sequence to a high-dimensional vector to capture its semantic features; position embedding provides position information for each token_id, enabling the model to perceive the order of token_ids in the sequence. Through the combination of these two parts, the input sequence is transformed into a representation matrix containing rich feature information (i.e., the embedding representation matrix), which serves as the input for the subsequent attention mechanism.

[0136] In an implementation manner of this embodiment, when inputting the adjusted embedding representation matrix into the multiple sequence attention module, the multiple sequence attention module combines the Transformer QKV mechanism, the gating mechanism, and the biological prior knowledge of linkage disequilibrium to perform row-column attention calculation, gating mechanism to adjust weights, and residual connection and normalization on the adjusted embedding representation matrix, and obtain an output vector, specifically including:

[0137] Input the adjusted embedding representation matrix into the multiple sequence attention module. The multiple sequence attention module combines the Transformer QKV mechanism, the gating mechanism, and the biological prior knowledge of linkage disequilibrium to perform row-column attention calculation on the adjusted embedding representation matrix. Among them, the row-column attention calculation includes row attention calculation and column attention calculation:

[0138] The row attention calculation includes calculating the row attention score Attention row :

[0139]

[0140] Q = MSA s ·W Q ;

[0141] K = MSA s ·W K ;

[0142] V = MSA s ·W V ;

[0143] Gate row = σ(MSA s );

[0144] Among them, Sigmoid(·) represents the activation function, Q represents the query vector, K represents the key vector, T represents the transpose, bias LD represents the prior value of the linkage disequilibrium matrix, which is used to measure the linkage relationship between sequence segments; d k represents the dimension of the key vector, which is used for scaling to prevent gradient explosion; V represents the value vector, Gate row represents the gating weight of the row attention, W Q represents the weight matrix corresponding to the query vector, W K represents the weight matrix corresponding to the key vector, W V represents the weight matrix corresponding to the value vector, σ represents the Sigmoid function;

[0145] The column attention calculation includes calculating the column attention score Attention col :

[0146]

[0147] Gate col == σ(MSA s );

[0148] Among them, Gate col represents the gating weight of the column attention;

[0149] For low-frequency sites, the gating mechanism is used to increase the gating weight corresponding to the low-frequency sites;

[0150] Through residual connection, the input of each layer in the multi-sequence attention module is added to the output;

[0151] Through normalization, the output of each layer in the multi-sequence attention module is standardized to obtain the output vector.

[0152] Specifically, the core of the model is to capture the deep association between the query sequence and the reference sequences through the context information interaction among multiple sequences. The input data includes the embedded multiple sequence alignment matrix (i.e., the embedded representation matrix) and the similarity weight (similarity_weight) between each reference sequence and the query sequence, and the latter is used to guide the attention distribution. During the attention calculation process, the model adopts the standard Transformer QKV mechanism to map the input matrix into the query vector Q, the key vector K, and the value vector V respectively. By combining the similarity weights between sequences, the model can preferentially focus on the reference sequences with higher similarity to the query sequence, thereby more effectively capturing the context relationship between sequences and highlighting the characteristics of mutation information. In addition, in the multiple sequence attention module, a gating mechanism is introduced in each layer to control the information flow weight between different sequences. Specifically, the model generates a gating weight value for each sequence to adjust the information flux of that sequence. Through this mechanism, the model can enhance the feature interaction of the reference sequences highly relevant to the query sequence while suppressing the interference of low-similarity sequences to the model. This design enables the model to focus more on the information of key sequences, thereby improving the model's learning ability for mutation information, especially showing higher accuracy in the prediction task of masked sites and significantly enhancing the completion performance of low-frequency alleles. It should be noted that here the multiple sequence alignment matrix actually refers to the target MSA matrix, the query sequence actually refers to the target query sequence, and the reference sequences refer to the target reference sequences, all of which are token_id sequences.

[0153] S4. Obtain the missing sequence to be tested and input the missing sequence to be tested into the target genotype completion model, and the target genotype completion model outputs the genotypes of the missing sites in the missing sequence to be tested.

[0154] In one implementation manner of this embodiment, after obtaining the missing sequence to be tested and inputting the missing sequence to be tested into the target genotype completion model, and the target genotype completion model outputs the genotypes of the missing sites in the missing sequence to be tested, it further includes:

[0155] Complete the missing sites of the missing sequence to be tested according to the genotypes of the missing sites in the missing sequence to be tested.

[0156] The present invention significantly improves the completion performance of low-frequency alleles while reducing computational consumption and memory occupancy. By optimizing the model structure and introducing new algorithm strategies, the present invention comprehensively enhances the ability to capture low-frequency variant information, thereby achieving the performance improvement of genotype completion and enhancing the accuracy and reliability of downstream analysis tasks.

[0157] In another implementation manner of this embodiment, as Figure 2As shown, first, the MSA matrix (referring to the target MSA matrix) and position information are converted into token embeddings and position embeddings. Then, by calculating the similarity between the target query sequence and each target reference sequence in the MSA matrix, similarity weights are generated and assigned to the MSA matrix. Next, after the similarity weight assignment process, the MSA matrix performs row and column attention calculations, gating weight adjustment weights, and residual connection and normalization. Among them, the row attention includes local bias (linkage disequilibrium LD), while the column attention has no bias. These operations are repeated 12 times to enhance the model's ability to capture sequence features. Finally, the output vector is transformed into a probability distribution through a fully connected layer, and the element corresponding to the maximum probability is used as the genotype of the missing site.

[0158] In another implementation manner of this embodiment, as Figure 3 shown, first, the query sequence and the multi-reference gene template are segmented into fixed-length segments and masked. Then, the multi-reference gene template sequence (i.e., the reference sequence) is aligned with the query sequence. The multi-reference gene template consists of multiple reference sequences. After alignment, the sequences are ranked according to similarity, and the number of occurrences of bases at each position is counted to form a sequence count matrix. Then, according to the information in the sequence count matrix, the sequences are ranked by prefix similarity, and the sequences ranked Top(n) are selected as the final MSA result. Among them, P represents paternal / non-paternal, S represents superpopulation, and C represents country / race. The whole process includes steps such as masking, similarity ranking, sequence counting, matrix transformation, and prefix similarity ranking to construct a high-quality MSA.

[0159] In addition, based on the above genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism, the present invention also provides a genotype completion system based on the multi-reference gene template row-column collaborative attention mechanism. Among them, a preferred embodiment of the genotype completion system based on the multi-reference gene template row-column collaborative attention mechanism is as Figure 4 shown, specifically including:

[0160] Sequence acquisition module 01: used to acquire a VCF file, process the VCF file to obtain a gene sequence, and acquire the position sequence corresponding to the gene sequence;

[0161] Target MSA matrix construction module 02: used to mask the gene sequence to obtain a query sequence, based on the query sequence, sort the reference sequences in the multi-reference gene template by multiple sequence alignment according to multiple biological attributes to obtain an MSA matrix, and perform mapping processing on the MSA matrix to obtain a target MSA matrix;

[0162] Genotype Completion Model Training Module 03: It is used to construct a genotype completion model based on the row-column collaborative attention mechanism of multiple reference gene templates, and train the genotype completion model according to the target MSA matrix and the position sequence to obtain the target genotype completion model;

[0163] Genotype Completion Module 04: It is used to obtain the to-be-detected missing sequence and input the to-be-detected missing sequence into the target genotype completion model, and the target genotype completion model outputs the genotypes of the missing sites in the to-be-detected missing sequence.

[0164] In addition, based on the above genotype completion method and system based on the row-column collaborative attention mechanism of multiple reference gene templates, the present invention also correspondingly provides a terminal. Among them, a preferred embodiment of the terminal is as Figure 5 shown, and specifically includes a processor 10, a memory 20 and a display 30. Figure 5 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0165] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as the hard disk or memory of the terminal. The memory 20 may also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, and a Flash Card equipped on the terminal. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store the application software installed on the terminal and various types of data, such as storing the program code of the terminal. The memory 20 may also be used to temporarily store the data that has been output or will be output. In one embodiment, a genotype completion program 40 based on the row-column collaborative attention mechanism of multiple reference gene templates is stored on the memory 20, and the genotype completion program 40 based on the row-column collaborative attention mechanism of multiple reference gene templates can be executed by the processor 10, so as to implement the steps of the genotype completion method based on the row-column collaborative attention mechanism of multiple reference gene templates in this application.

[0166] The processor 10 may be a Central Processing Unit (CPU), a microprocessor or other data processing chips in some embodiments, and is used to run the program code stored in the memory 20 or process data, such as executing the genotype completion program 40 based on the row-column collaborative attention mechanism of multiple reference gene templates, etc.

[0167] The display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. in some embodiments. The display 30 is used to display information on the terminal and to display a visual user interface.

[0168] In one embodiment, when the processor 10 executes the genotype completion program 40 based on the multi-reference gene template row-column collaborative attention mechanism in the memory 20, the steps of the genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism as described above are implemented.

[0169] The present invention also correspondingly provides a computer-readable storage medium. Among them, the computer-readable storage medium stores a genotype completion program based on the multi-reference gene template row-column collaborative attention mechanism. When the genotype completion program based on the multi-reference gene template row-column collaborative attention mechanism is executed by the processor, the steps of the genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism as described above are implemented.

[0170] It should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal including that element.

[0171] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A genotype completion method based on a multi-reference gene template row-column collaborative attention mechanism, characterized in that The genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism includes: Obtain a VCF file, process the VCF file to obtain a gene sequence, and obtain the corresponding position sequence of the gene sequence; Perform a masking process on the gene sequence to obtain a query sequence. Based on the query sequence, sort the reference sequences in the multi-reference gene template through multiple sequence alignment according to multiple biological attributes to obtain an MSA matrix, and perform a mapping process on the MSA matrix to obtain a target MSA matrix; Construct a genotype completion model based on the multi-reference gene template row-column collaborative attention mechanism. According to the target MSA matrix and the position sequence, train the genotype completion model to obtain a target genotype completion model; Obtain a to-be-detected missing sequence and input the to-be-detected missing sequence into the target genotype completion model. The target genotype completion model outputs the genotypes of the missing sites in the to-be-detected missing sequence.

2. The genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism according to claim 1, wherein The obtaining of the VCF file, processing the VCF file to obtain a gene sequence, and obtaining the corresponding position sequence of the gene sequence specifically includes: Obtain a VCF file and parse the VCF file to obtain an initial gene sequence; Perform a standardization process on the initial gene sequence through multi-condition filtering to obtain a gene sequence S’, where the multi-condition filtering includes gene frequency screening and mutation type filtering, and the gene sequence S’ is: S’ = (S1, S2,..., S i ,..., S m ) ∈ R 1×m ; Among them, S i represents the element at the i-th site in the gene sequence S', R represents a one-dimensional matrix, and m represents the maximum sequence length of the input genotype completion model; Obtain the corresponding position sequence P’ of the gene sequence, and the position sequence P’ is: P’ = (P1, P2,..., P i ,..., P m ) ∈ R 1×m ; Among them, P i represents the element S at the i-th locus i on the chromosome.

3. The genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism according to claim 2, wherein The performing of the masking process on the gene sequence to obtain a query sequence, sorting the reference sequences in the multi-reference gene template through multiple sequence alignment according to multiple biological attributes based on the query sequence to obtain an MSA matrix, and performing a mapping process on the MSA matrix to obtain a target MSA matrix specifically includes: Randomly select a part of the elements in the gene sequence for masking according to a set ratio to obtain a query sequence W; W = (S1, S2,..., S i = [MASK],..., S m ) ∈ R 1×m ; Where, [MASK] represents a masking mark; Calculate the first similarity between the query sequence and each reference sequence in the multi-reference gene template respectively; Perform a first sorting on multiple reference sequences in descending order of the first similarity according to the first similarity between the query sequence and each reference sequence; Concatenate prefix information to the query sequence and each reference sequence, where the prefix information includes multiple biological attributes, and the multiple biological attributes include kinship, species, and region; Calculate the second similarity between the prefix information of the query sequence and the prefix information of each reference sequence respectively; Perform a second sorting on the multiple reference sequences after the first sorting in descending order of the second similarity according to the second similarity between the prefix information of the query sequence and the prefix information of each reference sequence, where the priority of the first sorting is higher than that of the second sorting; Select the first set number of reference sequences after the second sorting in the multi-reference gene template, and randomly select partial elements of the first set number of reference sequences according to a set ratio for masking processing to obtain the first set number of masked reference sequences; Combine the query sequence and the first set number of masked reference sequences to obtain an MSA matrix; Map each element in the MSA matrix to a natural number through a mapping vocabulary to obtain a target MSA matrix; Wherein, the target MSA matrix includes a target query sequence and target reference sequences.

4. The genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism according to claim 3, wherein The constructing a genotype completion model based on the row-column collaborative attention mechanism of the multi-reference gene template, and training the genotype completion model according to the target MSA matrix and the position sequence to obtain a target genotype completion model specifically includes: Construct a genotype completion model based on the row-column collaborative attention mechanism of the multi-reference gene template, wherein the genotype completion model includes an embedding layer, a multi-sequence attention module and a fully connected layer; Input the target MSA matrix and the position sequence into the embedding layer, and the embedding layer performs initialization processing on the feature representation of the target MSA matrix to obtain an embedding representation matrix; Adjust the embedding representation matrix by using the similarity weights between the target query sequence and each target reference sequence to obtain an adjusted embedding representation matrix: MSA s = MSA' · similarity_weight; Among them, MSA s represents the adjusted embedded representation matrix, MSA' represents the embedded representation matrix, and similarity_weight represents the similarity weight; Input the adjusted embedding representation matrix into the multi-sequence attention module, and the multi-sequence attention module combines the Transformer QKV mechanism, the gating mechanism and the biological prior knowledge of linkage disequilibrium to perform row-column attention calculation, gating mechanism weight adjustment, and residual connection and normalization on the adjusted embedding representation matrix to obtain an output vector; Input the output vector into the fully connected layer, and the fully connected layer maps the output vector to a probability vector of each element in the corresponding mapping vocabulary, and uses the element corresponding to the maximum probability in the probability vector as the genotype of the missing site; Compare the genotype of the missing site with the gene sequence to obtain a comparison result, and train the genotype completion model according to the comparison result to obtain a target genotype completion model.

5. The genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism according to claim 4, wherein The inputting the target MSA matrix and the position sequence into the embedding layer, and the embedding layer performs initialization processing on the feature representation of the target MSA matrix to obtain an embedding representation matrix specifically includes: Input the target MSA matrix and the position sequence into the embedding layer, and the embedding layer maps each natural number of the input sequence in the target MSA matrix to a high-dimensional vector to obtain a first embedding vector; Perform sine-cosine encoding on each natural number of the input sequence according to the position sequence to obtain a second embedding vector; Add the first embedding vector and the second embedding vector to obtain the embedding representation matrix corresponding to the target MSA matrix: E = E token + E pos ; Among them, E represents the embedding representation matrix, and E token represents the first embedding vector, and E pos represents the second embedding vector; Wherein, the input sequence includes the target query sequence and target reference sequences in the target MSA matrix.

6. The genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism according to claim 4, wherein Input the adjusted embedded representation matrix into the multi-sequence attention module. The multi-sequence attention module combines the Transformer QKV mechanism, the gating mechanism, and the biological prior knowledge of linkage disequilibrium to perform row and column attention calculations, adjust weights through the gating mechanism, and perform residual connection and normalization on the adjusted embedded representation matrix to obtain an output vector, specifically including: Input the adjusted embedded representation matrix into the multi-sequence attention module. The multi-sequence attention module combines the Transformer QKV mechanism, the gating mechanism, and the biological prior knowledge of linkage disequilibrium to perform row and column attention calculations on the adjusted embedded representation matrix. Among them, the row and column attention calculations include row attention calculation and column attention calculation: The row attention calculation includes calculating the row attention score Attention row : Q = MSA s ·W Q ; K = MSA s ·W K ; V = MSA s ·W V ; Gate row = σ(MSA s ); Among them, Sigmoid(·) represents the activation function, Q represents the query vector, K represents the key vector, T represents the transpose, and bias LD represents the prior value of the linkage disequilibrium matrix, and d k represents the dimension of the key vector, V represents the value vector, and Gate row represents the gating weight of the row attention, and W Q represents the weight matrix corresponding to the query vector, and W K represents the weight matrix corresponding to the key vector, and W V represents the weight matrix corresponding to the value vector, and σ represents the Sigmoid function; The column attention calculation includes calculating the column attention score Attention col : Gate col = σ(MSA s ); Among them, Gate col represents the gating weights for column attention; For low-frequency sites, use the gating mechanism to increase the gating weight corresponding to the low-frequency sites; Through residual connection, add the input of each layer in the multi-sequence attention module to the output; Through normalization, standardize the output of each layer in the multi-sequence attention module to obtain an output vector.

7. The genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism according to claim 1, wherein After obtaining the to-be-tested missing sequence and inputting the to-be-tested missing sequence into the target genotype completion model, and the target genotype completion model outputs the genotypes of the missing sites in the to-be-tested missing sequence, the following steps are further included: According to the genotypes of the missing sites in the to-be-tested missing sequence, complete the missing sites in the to-be-tested missing sequence.

8. A genotype completion system based on a multi-reference gene template row-column collaborative attention mechanism, characterized in that, The genotype completion system based on the multi-reference gene template row-column collaborative attention mechanism includes: Sequence acquisition module: used to acquire a VCF file, process the VCF file to obtain a gene sequence, and acquire the position sequence corresponding to the gene sequence; Target MSA matrix construction module: used to perform masking processing on the gene sequence to obtain a query sequence, based on the query sequence, perform multi-sequence alignment-based sorting on the reference sequences in the multi-reference gene template according to multiple biological attributes to obtain an MSA matrix, and perform mapping processing on the MSA matrix to obtain a target MSA matrix; Genotype completion model training module: used to construct a genotype completion model based on the multi-reference gene template row-column collaborative attention mechanism, and train the genotype completion model according to the target MSA matrix and the position sequence to obtain a target genotype completion model; Genotype completion module: used to acquire a to-be-tested missing sequence and input the to-be-tested missing sequence into the target genotype completion model, and the target genotype completion model outputs the genotypes of the missing sites in the to-be-tested missing sequence.

9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a genotype completion program based on the multi-reference gene template row-column collaborative attention mechanism stored on the memory and executable on the processor. When the genotype completion program based on the multi-reference gene template row-column collaborative attention mechanism is executed by the processor, it implements the steps of the genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a genotype completion program based on a multi-reference gene template row-column collaborative attention mechanism. When the genotype completion program based on the multi-reference gene template row-column collaborative attention mechanism is executed by a processor, it implements the steps of the genotype completion method based on the multi-reference gene template row-column collaborative attention mechanism according to any one of claims 1-7.