Typing method of cow individual BAM file
Through the combination of deep learning and large language model, the problem of low typing accuracy of individual BAM files in dairy cows is solved, and high-precision genotyping is achieved, which is suitable for large-scale research and improves computing efficiency.
Patent Information
- Application Number
- CN202510555162.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-25
AI Technical Summary
The typing accuracy of existing individual BAM files of dairy cows is low, and it is impossible to effectively identify complex variant types and insufficient genotyping accuracy.
Using a method of combining deep learning models and large language models, the clustering results of all variant types signals in the BAM files of cow individuals are obtained, integrated into structural variant signals, and haplotype classification is performed. The model parameters are optimized using the comprehensive loss function until convergence is achieved, and the trained model is obtained to output the typing haplotype file.
It significantly improves the typing accuracy of individual BAM files of dairy cows, reduces data requirements, is suitable for large-scale research, and improves the computing efficiency and applicability of methods.
Smart Images

Figure CN120372323A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a genotyping method for individual dairy cow BAM files. Background Art
[0002] Dairy cows, also known as milk cows, are animals of the genus Bos in the family Bovidae. Dairy cows have a clear head contour, slightly long, with a thin neck and wrinkles; thin skin, short and fine hair, less subcutaneous fat, a well-proportioned and symmetrical body structure, delicate and compact, with distinct edges and corners; the hindquarters are more developed than the forequarters, and the udder is huge. They are cold-resistant and less heat-resistant. A healthy and prime dairy cow produces milk for 10 months out of 12 months in a year. The average lifespan is about 17 years, and it can normally produce milk for 7 - 13 years. Milk is the milk squeezed from female dairy cows and is one of the oldest natural beverages, known as white blood. Milk is a homogeneous fluid with the inherent flavor of cow's milk, a mellow taste, and has the effects of supplementing protein, calcium, calming the nerves, etc.
[0003] Currently, the breeding of existing breeding dairy cows is mostly carried out according to pedigrees, and the gene sequences of dairy cows are not overly concerned during the breeding process. The means of analyzing the gene sequences of dairy cows is to obtain a DNA sequence fragment with relatively high credibility through sequencing and then quality control, and then analyze the obtained fragments according to corresponding tasks through various tools and methods. The most important analysis is to obtain the differences from the known reference sequence, that is, the biological variations; the variations that occur in longer and more complex forms are called structural variations, including insertions, deletions, duplications, translocations, inversions, and their fusion forms. Existing evidence shows that structural variations have a huge impact on the traits of organisms, so detecting structural variations is very important. However, there are currently problems with incomplete detection, inaccurate detection, inability to identify complex variation types, and low accuracy of gene typing for structural variations. Summary of the Invention
[0004] The purpose of the present invention is to solve the problem of low genotyping accuracy of existing individual dairy cow BAM files, and to propose a genotyping method for individual dairy cow BAM files.
[0005] The specific process of a genotyping method for individual dairy cow BAM files is as follows:
[0006] Step 1: Obtain the BAM file of an individual dairy cow; each BAM file contains M SNP sites;
[0007] Step 2: Obtain all clustering cluster results of all variant type signals in the BAM file of the individual dairy cow;
[0008] The signals contained in each clustering cluster are integrated into a structural variant signal for output;
[0009] The structural variant signals include insertions, deletions, duplications, translocations, and inversions;
[0010] Step 3: Perform haplotype typing on each output structural variation signal to obtain the haplotype 1 file and haplotype 2 file after typing;
[0011] Step 4: Use the BAM file of the dairy cow individual obtained in Step 1 as the input of the deep learning model, and the haplotype 1 file and haplotype 2 file after typing as the output of the deep learning model;
[0012] Use the BAM file of the dairy cow individual obtained in Step 1 as the input of the large language model, and the haplotype 1 file and haplotype 2 file after typing as the output of the large language model;
[0013] Optimize the parameters of the deep learning model and the large language model using a comprehensive loss function, and perform gradient update in combination with the Adam optimizer until the comprehensive loss function converges to obtain the trained deep learning model and large language model;
[0014] Step 5: Input the BAM file of the dairy cow individual to be tested into the trained deep learning model, and the trained deep learning model outputs the haplotype 1 file and haplotype 2 file after typing the BAM file of the dairy cow individual to be tested.
[0015] Preferably, in Step 2, obtain all clustering cluster results of all variation type signals in the BAM file of the dairy cow individual;
[0016] Integrate the signals contained in each clustering cluster into a structural variation signal for output;
[0017] The structural variation signals include insertions, deletions, duplications, translocations and inversions;
[0018] The specific process is as follows:
[0019] Step 21: Obtain the clustering cluster results of all signals in the variation signal set of the "insertion" variation type; the specific process is as follows:
[0020] Step 211: Initialize an empty clustering cluster, and add the first signal in the variation signal set of the "insertion" variation type as the starting signal;
[0021] Step 212: Calculate the similarity between the current signal and the last signal in each existing clustering cluster;
[0022] Step 213:
[0023] If the similarity is less than the predetermined threshold, construct a new clustering cluster and add the current signal to the new clustering cluster;
[0024] If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current clustering cluster;
[0025] Step 214. Repeat steps 212 - 213 until all signals in the mutation signal set of the "insertion" mutation type are judged; obtain the clustering cluster results of all signals in the mutation signal set of the "insertion" mutation type;
[0026] Step 22. Obtain the clustering cluster results of all signals in the mutation signal set of the "deletion" mutation type; the specific process is as follows:
[0027] Step 221. Initialize an empty cluster and add the first signal in the mutation signal set of the "deletion" mutation type as the starting signal;
[0028] Step 222. Calculate the similarity between the current signal and the last signal in each existing cluster;
[0029] Step 223.
[0030] If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster;
[0031] If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster;
[0032] Step 224. Repeat steps 222 - 223 until all signals in the mutation signal set of the "deletion" mutation type are judged;
[0033] Step 23. Obtain the clustering cluster results of all signals in the mutation signal set of the "duplication" mutation type; the specific process is as follows:
[0034] Step 231. Initialize an empty cluster and add the first signal in the mutation signal set of the "duplication" mutation type as the starting signal;
[0035] Step 232. Calculate the similarity between the current signal and the last signal in each existing cluster;
[0036] Step 233.
[0037] If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster;
[0038] If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster;
[0039] Step 234. Move to the next signal and repeat steps 232 - 233 until all signals in the mutation signal set of the "duplication" mutation type are judged;
[0040] Step 24. Obtain the clustering cluster results of all signals in the mutation signal set of the "translocation" mutation type; the specific process is as follows:
[0041] Step 241: Initialize an empty cluster and add the first signal in the set of translocation mutation type mutation signals as the starting signal;
[0042] Step 242: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0043] Step 243:
[0044] If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster;
[0045] If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster;
[0046] Step 244: Repeat steps 242 - 243 until all signals in the set of translocation mutation type mutation signals are judged; obtain the clustering results of all signals in the set of translocation mutation type mutation signals;
[0047] Step 25: Obtain the clustering results of all signals in the set of inversion mutation type mutation signals; the specific process is as follows:
[0048] Step 251: Initialize an empty cluster and add the first signal in the set of inversion mutation type mutation signals as the starting signal;
[0049] Step 252: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0050] Step 253:
[0051] If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster;
[0052] If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster;
[0053] Step 254: Repeat steps 252 - 253 until all signals in the set of inversion mutation type mutation signals are judged;
[0054] Obtain the clustering results of all signals in the set of inversion mutation type mutation signals.
[0055] Preferably, the process of calculating the similarity between the current signal and the last signal in each existing cluster is as follows:
[0056] For insertion, deletion, inversion, and duplication mutation signals, obtain a comprehensive similarity score S; expressed as:
[0057] S = w1×S1 + w2×S2 + w3×S3
[0058] Among them, w1, w2, and w3 are the weight coefficients of the position similarity score, the mutation size similarity score / chromosome similarity score respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, and w1 + w2 + w3 = 1;
[0059] For translocation mutation signals, a comprehensive similarity score S is obtained; it is expressed as:
[0060] S = w4 × S4 + w5 × S5
[0061] Among them,
[0062] w4 is the weight coefficient of the starting chromosome similarity score S4, w4 = 0.5,
[0063] w5 is the weight coefficient of the ending chromosome similarity score S5, w5 = 0.5,
[0064] w4 + w5 = 1.
[0065] Preferably, for insertion, deletion, inversion, and duplication mutation signals, a comprehensive similarity score S is obtained; it is expressed as:
[0066] S = w1 × S1 + w2 × S2 + w3 × S3
[0067] Among them, w1, w2, and w3 are the weight coefficients of the position similarity score, the mutation size similarity score / chromosome similarity score respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, and w1 + w2 + w3 = 1;
[0068] The specific process is as follows:
[0069] 1), Calculate the position similarity S1 score; the specific process is as follows:
[0070] The calculation formula for the position similarity S1 score is as follows:
[0071] S1 = |Start1 - Start2|
[0072] Among them, Start1 and Start2 are the signal starting positions in the two mutation signals respectively;
[0073] 2), Calculate the mutation signal size similarity score S2; the specific process is as follows:
[0074] Record the mutation lengths of the two mutation signals as SVlen1 and SVlen2 respectively, and the mutation signal size similarity score S2 is calculated by the following formula:
[0075]
[0076] 3), Calculate the chromosome similarity score S3; the specific process is as follows:
[0077] Let the chromosome numbers where the two mutation signals are located be Chrom1 and Chrom2 respectively. The chromosome similarity score S3 is calculated by the following discrete function:
[0078]
[0079] When the chromosome names where the two mutation signals are located are the same, the chromosome similarity score is 1;
[0080] When the chromosome names where the two mutation signals are located are different, the chromosome similarity score is 0;
[0081] 4), Perform a weighted sum of the position similarity S1 score, the mutation size similarity score S2 score, and the chromosome similarity score S3 to obtain the comprehensive similarity score S; expressed as:
[0082] S = w1×S1 + w2×S2 + w3×S3
[0083] Among them, w1, w2, and w3 are the weight coefficients of the position similarity score, the mutation size similarity score / chromosome similarity score respectively. w1 = 0.2, w2 = 0.2, w3 = 0.6, and w1 + w2 + w3 = 1.
[0084] Preferably, for the translocation mutation signal, obtain the comprehensive similarity score S; expressed as:
[0085] S = w4×S4 + w5×S5
[0086] Among them,
[0087] w4 is the weight coefficient of the starting chromosome similarity score S4, w4 = 0.5,
[0088] w5 is the weight coefficient of the ending chromosome similarity score S5, w5 = 0.5,
[0089] w4 + w5 = 1;
[0090] The specific process is as follows:
[0091] Calculate the starting chromosome similarity score:
[0092]
[0093] Among them, S4 represents the starting chromosome similarity score;
[0094] represents the starting chromosome name of the first signal;
[0095] Indicates the starting chromosome name of the second signal;
[0096] Calculate the termination chromosome similarity score:
[0097]
[0098] Among them, S5 represents the termination chromosome similarity score;
[0099] Indicates the termination chromosome name of the first signal;
[0100] Indicates the termination chromosome name of the second signal;
[0101] Perform weighted summation on the starting chromosome similarity score S4 and the termination chromosome similarity score S5 to obtain the comprehensive similarity score S; expressed as:
[0102] S = w4 × S4 + w5 × S5
[0103] Among them,
[0104] w4 is the weight coefficient of the starting chromosome similarity score S4, w4 = 0.5,
[0105] w5 is the weight coefficient of the termination chromosome similarity score S5, w5 = 0.5,
[0106] w4 + w5 = 1.
[0107] Preferably, in the S3, haplotype typing is performed on each output structural variation signal to obtain the typed hap1 file and hap2 file; the specific process is as follows:
[0108] Step 31: Input the VCF file of each structural variation signal into the variant detection tool;
[0109] The variant detection tool outputs a typed VCF file, and the typed VCF file contains M SNP sites;
[0110] The typed VCF file contains 8 columns of information, which are respectively:
[0111] Chromosome name, SNP position, ID of the SNP, reference gene, alternative allele, quality value, filter flag, annotation information column;
[0112] Step 32: Use the Bcftools software to filter the typed VCF file to obtain the filtered typed VCF file;
[0113] Step 33: Input the filtered genotyped VCF file obtained in Step 22 into the WhatsHap genotyping tool, which outputs two genotyping files, namely Haplotype 1 file and Haplotype 2 file.
[0114] Preferably, in Step 32, the Bcftools software is used to filter the genotyped VCF file to obtain a filtered genotyped VCF file; the specific process is as follows:
[0115] Retain the rows in the genotyped VCF file where the value of the "filter flag" is equal to PASS;
[0116] Delete the rows in the genotyped VCF file where the value of the "filter flag" is not equal to PASS;
[0117] Finally, generate a genotyped VCF file;
[0118] Each SNP locus in the genotyped VCF file is divided into Haplotype 1 and Haplotype 2.
[0119] Preferably, in Step 33, the filtered genotyped VCF file obtained in Step 22 is input into the WhatsHap genotyping tool, which outputs two genotyping files, namely Haplotype 1 file and Haplotype 2 file; the specific process is as follows:
[0120] Step 331: Set the Modki software parameters: pileup, traditional;
[0121] Use the Modki software to convert the filtered genotyped VCF file obtained in Step 22 into a tsv format file;
[0122] Step 332: Use the WhatsHap genotyping tool to generate a BED.gz format file from the tsv format file;
[0123] Step 333: Input the genotyped VCF file obtained in Step 22 and the BED.gz format file obtained in Step 332 into the WhatsHap genotyping tool, which outputs a Haplotype 1 file and a Haplotype 2 file.
[0124] Preferably, in Step 4, the BAM file of the dairy cow individual obtained in Step 1 is used as the input of the deep learning model, and the genotyped Haplotype 1 file and Haplotype 2 file are used as the output of the deep learning model;
[0125] Use the BAM file of the dairy cow individual obtained in Step 1 as the input of the large language model, and the genotyped Haplotype 1 file and Haplotype 2 file as the output of the large language model;
[0126] Optimize the parameters of the deep learning model and the large language model using a comprehensive loss function, and combine with the Adam optimizer for gradient update until the comprehensive loss function converges to obtain a trained deep learning model and large language model;
[0127] The specific process is as follows:
[0128] Step 41: Construct a deep learning model, which successively includes:
[0129] The first 1×1 convolutional layer, BN layer, the first 7×7 depth convolutional layer, the first LN layer, the second 1×1 convolutional layer, the first GELU, the first GRN, the third 1×1 convolutional layer, the second 7×7 depth convolutional layer, the second LN layer, the fourth 1×1 convolutional layer, the second GELU, the second GRN, the fifth 1×1 convolutional layer, the third 7×7 depth convolutional layer, the third LN layer, the sixth 1×1 convolutional layer, the third GELU, the third GRN, the seventh 1×1 convolutional layer, fully connected layer, softmax layer;
[0130] The working process of the deep learning model is as follows:
[0131] Input the BAM file of the dairy cow individual obtained in step 1 into the first 1×1 convolutional layer and BN layer in sequence, and the BN layer outputs feature A;
[0132] The feature A output by the BN layer is input into the first 7×7 depth convolutional layer, the first LN layer, the second 1×1 convolutional layer, the first GELU, the first GRN, and the third 1×1 convolutional layer in sequence, and the third 1×1 convolutional layer outputs feature A';
[0133] The feature A' output by the third 1×1 convolutional layer and the feature A output by the BN layer are element-wise added to obtain feature A";
[0134] The feature A" is input into the second 7×7 depth convolutional layer, the second LN layer, the fourth 1×1 convolutional layer, the second GELU, the second GRN, and the fifth 1×1 convolutional layer in sequence, and the fifth 1×1 convolutional layer outputs feature A''';
[0135] The feature A''' output by the fifth 1×1 convolutional layer and the feature A" are element-wise added to obtain feature
[0136] Feature Is input into the third 7×7 depth convolutional layer, the third LN layer, the sixth 1×1 convolutional layer, the third GELU, the third GRN, and the seventh 1×1 convolutional layer in sequence, and the seventh 1×1 convolutional layer outputs feature
[0137] The feature output by the seventh 1×1 convolutional layer And feature Perform element-wise summation to obtain Feature B;
[0138] Feature B is sequentially input into a fully connected layer and a softmax layer, and the softmax layer outputs the classification result;
[0139] Step 42: Use a large language model to generate a haplotype 1 file and a haplotype 2 file corresponding to the BAM file of the dairy cow individual obtained in Step 1;
[0140] Step 43: Optimize the parameters of the deep learning model and the large language model using a comprehensive loss function, and perform gradient update in combination with the Adam optimizer until the comprehensive loss function converges to obtain a trained deep learning model and a large language model.
[0141] Preferably, the comprehensive loss function is
[0142] where N represents the total number of data in the BAM file of the dairy cow individual, i represents the i-th one; k represents the k-th one;
[0143] F i 1 represents the feature output by the deep learning model when the i-th data in the BAM file of the dairy cow individual is input into the deep learning model;
[0144] F i 2 represents the semantic feature of the text description information output by the large language model when the i-th data in the BAM file of the dairy cow individual is input into the large language model;
[0145] represents the semantic feature of the text description information output by the large language model when the k-th data in the BAM file of the dairy cow individual is input into the large language model;
[0146] s(F i 1 ,F i 2 ) represents the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the i-th data;
[0147] represents the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the k-th data;
[0148] τ represents a temperature hyperparameter.
[0149] The beneficial effects of the present invention are:
[0150] The present invention realizes the typing of the BAM file of individual dairy cows through the feature information obtained by the deep learning model and the semantic information of the large language model, significantly improving the typing accuracy; compared with the traditional classification method, the classification accuracy of the present invention has significant advantages. The present invention reduces the data requirements, does not rely on pedigree data, is applicable to large-scale research, and improves the applicability of the method. The method of the present invention avoids complex variant typing steps, directly analyzes the data, improves the calculation efficiency, and reduces the consumption of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0151] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0152] DETAILED DESCRIPTION OF THE EMBODIMENT 1: The specific process of a method for typing the BAM file of individual dairy cows in this embodiment is as follows:
[0153] Step 1: Obtain the BAM file of the individual dairy cow; each BAM file contains M SNP sites
[0154] Step 2: Obtain all clustering cluster results of all variant type signals in the BAM file of the individual dairy cow
[0155] The signals contained in each clustering cluster are integrated into a structural variant signal for output;
[0156] The structural variant signals include insertion, deletion, duplication, translocation, and inversion;
[0157] Step 3: Perform haplotype typing on each output structural variant signal to obtain the typed haplotype 1 file and haplotype 2 file;
[0158] Step 4: Take the BAM file of the individual dairy cow obtained in Step 1 as the input of the deep learning model, and the typed haplotype 1 file and haplotype 2 file as the output of the deep learning model;
[0159] Take the BAM file of the individual dairy cow obtained in Step 1 as the input of the large language model, and the typed haplotype 1 file and haplotype 2 file as the output of the large language model;
[0160] Use the comprehensive loss function to optimize the parameters of the deep learning model and the large language model, and combine the Adam optimizer to perform gradient update until the comprehensive loss function converges to obtain the trained deep learning model and large language model;
[0161] Step 5: Input the BAM file of the dairy cow to be tested into the trained deep learning model, and the trained deep learning model outputs the typed haplotype 1 file and haplotype 2 file of the BAM file of the dairy cow to be tested.
[0162] Embodiment 2: The difference between this embodiment and Embodiment 1 is that in step 2, all clustering cluster results of all mutation type signals in the BAM file of the individual cow are obtained;
[0163] The signals contained in each clustering cluster are integrated into a structural variation signal for output;
[0164] The structural variation signals include insertion, deletion, duplication, translocation and inversion;
[0165] The specific process is as follows:
[0166] Step 21: Obtain the clustering cluster results of all signals in the mutation signal set of the "insertion" mutation type; the specific process is as follows:
[0167] Step 211: Initialize an empty clustering cluster, and add the first signal in the mutation signal set of the "insertion" mutation type as the starting signal;
[0168] Step 212: Calculate the similarity between the current signal and the last signal in each existing clustering cluster;
[0169] Step 213:
[0170] If the similarity is less than the predetermined threshold, construct a new clustering cluster and add the current signal to the new clustering cluster;
[0171] If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current clustering cluster;
[0172] Step 214: Repeat steps 212 - 213 until all signals in the mutation signal set of the "insertion" mutation type are judged; obtain the clustering cluster results of all signals in the mutation signal set of the "insertion" mutation type;
[0173] Step 22: Obtain the clustering cluster results of all signals in the mutation signal set of the "deletion" mutation type; the specific process is as follows:
[0174] Step 221: Initialize an empty clustering cluster, and add the first signal in the mutation signal set of the "deletion" mutation type as the starting signal;
[0175] Step 222: Calculate the similarity between the current signal and the last signal in each existing clustering cluster;
[0176] Step 223:
[0177] If the similarity is less than the predetermined threshold, construct a new clustering cluster and add the current signal to the new clustering cluster;
[0178] If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current clustering cluster;
[0179] Step 224: Repeat Step 222 - Step 223 until all signals in the mutation signal set of the "deletion" mutation type are judged;
[0180] Step 23: Obtain the clustering results of all signals in the mutation signal set of the "duplication" mutation type; The specific process is as follows:
[0181] Step 231: Initialize an empty cluster and add the first signal in the mutation signal set of the "duplication" mutation type as the starting signal;
[0182] Step 232: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0183] Step 233:
[0184] If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster;
[0185] If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster;
[0186] Step 234: Move to the next signal and repeat Step 232 - Step 233 until all signals in the mutation signal set of the "duplication" mutation type are judged;
[0187] Step 24: Obtain the clustering results of all signals in the mutation signal set of the "translocation" mutation type; The specific process is as follows:
[0188] Step 241: Initialize an empty cluster and add the first signal in the mutation signal set of the "translocation" mutation type as the starting signal;
[0189] Step 242: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0190] Step 243:
[0191] If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster;
[0192] If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster;
[0193] Step 244: Repeat Step 242 - Step 243 until all signals in the mutation signal set of the "translocation" mutation type are judged; Obtain the clustering results of all signals in the mutation signal set of the "translocation" mutation type;
[0194] Step 25: Obtain the clustering results of all signals in the mutation signal set of the "inversion" mutation type; The specific process is as follows:
[0195] Step 251: Initialize an empty cluster, and add the first signal in the set of mutation signals of the "inversion" mutation type as the starting signal;
[0196] Step 252: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0197] Step 253:
[0198] If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster;
[0199] If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster;
[0200] Step 254: Repeat steps 252 - 253 until all signals in the set of mutation signals of the "inversion" mutation type are judged;
[0201] Obtain the clustering results of all signals in the set of mutation signals of the "inversion" mutation type.
[0202] Other steps and parameters are the same as those in the first specific implementation manner.
[0203] Specific implementation manner three: The difference between this implementation manner and the first or second specific implementation manner is that: calculating the similarity between the current signal and the last signal in each existing cluster; the specific process is as follows:
[0204] For insertion, deletion, inversion, and duplication mutation signals, obtain a comprehensive similarity score S; expressed as:
[0205] S = w1×S1 + w2×S2 + w3×S3
[0206] Among them, w1, w2, and w3 are the weight coefficients of the position similarity score, the mutation size similarity score / chromosome similarity score respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, w1 + w2 + w3 = 1;
[0207] For translocation mutation signals, obtain a comprehensive similarity score S; expressed as:
[0208] S = w4×S4 + w5×S5
[0209] Among them,
[0210] w4 is the weight coefficient of the starting chromosome similarity score S4, w4 = 0.5,
[0211] w5 is the weight coefficient of the ending chromosome similarity score S5, w5 = 0.5,
[0212] w4 + w5 = 1.
[0213] The other steps and parameters are the same as those in the first or second specific implementation manners.
[0214] Specific implementation manner four: The difference between this implementation manner and any one of the first to third specific implementation manners is that for the insertion, deletion, inversion, and duplication mutation signals, a comprehensive similarity score S is obtained, which is expressed as:
[0215] S = w1×S1 + w2×S2 + w3×S3
[0216] Among them, w1, w2, and w3 are the weight coefficients of the position similarity score, the mutation size similarity score / chromosome similarity score respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, and w1 + w2 + w3 = 1;
[0217] The specific process is as follows:
[0218] 1) Calculate the position similarity score S1; the specific process is as follows:
[0219] The calculation formula for the position similarity score S1 is as follows:
[0220] S1 = |Start1 - Start2|
[0221] Among them, Start1 and Start2 are the signal start positions in the two mutation signals respectively;
[0222] 2) Calculate the mutation signal size similarity score S2; the specific process is as follows:
[0223] Denote the mutation lengths of the two mutation signals as SVlen1 and SVlen2 respectively. The mutation signal size similarity score S2 is calculated by the following formula:
[0224]
[0225] 3) Calculate the chromosome similarity score S3; the specific process is as follows:
[0226] Suppose the chromosome numbers where the two mutation signals are located are Chrom1 and Chrom2 respectively. The chromosome similarity score S3 is calculated by the following discrete function:
[0227]
[0228] When the chromosome names where the two mutation signals are located are the same, the chromosome similarity score is 1;
[0229] When the chromosome names where the two mutation signals are located are different, the chromosome similarity score is 0;
[0230] 4), weighted sum the position similarity S1 score, the mutation size similarity score S2 score, and the chromosome similarity score S3 to obtain a comprehensive similarity score S, which is expressed as:
[0231] S = w1 × S1 + w2 × S2 + w3 × S3
[0232] Among them, w1, w2, and w3 are the weight coefficients of the position similarity score, the mutation size similarity score / chromosome similarity score respectively. w1 = 0.2, w2 = 0.2, w3 = 0.6, and w1 + w2 + w3 = 1.
[0233] Other steps and parameters are the same as those in any one of the specific embodiments one to three.
[0234] Specific embodiment five: The difference between this embodiment and any one of the specific embodiments one to four is that for the translocation mutation signal, a comprehensive similarity score S is obtained, which is expressed as:
[0235] S = w4 × S4 + w5 × S5
[0236] Among them,
[0237] w4 is the weight coefficient of the starting chromosome similarity score S4, w4 = 0.5,
[0238] w5 is the weight coefficient of the ending chromosome similarity score S5, w5 = 0.5,
[0239] w4 + w5 = 1;
[0240] The specific process is as follows:
[0241] Calculate the starting chromosome similarity score:
[0242]
[0243] Among them, S4 represents the starting chromosome similarity score;
[0244] represents the name of the starting chromosome of the first signal;
[0245] represents the name of the starting chromosome of the second signal;
[0246] Calculate the ending chromosome similarity score:
[0247]
[0248] Among them, S5 represents the ending chromosome similarity score;
[0249] represents the name of the ending chromosome of the first signal;
[0250] The name of the terminal chromosome representing the second signal;
[0251] The starting chromosome similarity score S4 and the terminal chromosome similarity score S5 are weighted and summed to obtain a comprehensive similarity score S; expressed as:
[0252] S = w4 × S4 + w5 × S5
[0253] Where,
[0254] w4 is the weight coefficient of the starting chromosome similarity score S4, w4 = 0.5,
[0255] w5 is the weight coefficient of the terminal chromosome similarity score S5, w5 = 0.5,
[0256] w4 + w5 = 1.
[0257] Other steps and parameters are the same as those in any one of the specific implementation manners one to four.
[0258] Specific implementation manner six: The difference between this implementation manner and any one of the specific implementation manners one to five is that: in the above S3, each output structural variation signal is haplotyped to obtain a haplotype 1 file and a haplotype 2 file after typing;
[0259] The specific process is as follows:
[0260] Step 31: Input the VCF file of each structural variation signal into a variant detection tool;
[0261] The variant detection tool outputs a typed VCF file, and the typed VCF file contains M SNP sites;
[0262] The typed VCF file contains 8 columns of information, which are respectively:
[0263] Chromosome name, SNP position, ID of the SNP, reference gene, alternative allele, quality value, filter flag, annotation information column;
[0264] Step 32: Use the Bcftools software to filter the typed VCF file to obtain a filtered typed VCF file;
[0265] Step 33: Input the filtered typed VCF file obtained in step 22 into the WhatsHap typing tool, and the WhatsHap typing tool outputs two typed files, namely a haplotype 1 file and a haplotype 2 file.
[0266] Other steps and parameters are the same as those in any one of the specific implementation manners one to five.
[0267] Embodiment 7: The difference between this embodiment and any one of Embodiments 1 to 6 is that: in step 32, the Bcftools software is used to filter the genotyped VCF file to obtain a filtered genotyped VCF file; the specific process is as follows:
[0268] Retain the rows in the genotyped VCF file where the "filter flag" value is equal to PASS;
[0269] Delete the rows in the genotyped VCF file where the "filter flag" value is not equal to PASS;
[0270] Finally, generate a genotyped VCF file;
[0271] Each SNP locus in the genotyped VCF file is divided into haplotype 1 and haplotype 2.
[0272] Other steps and parameters are the same as those in Embodiments 1 to 6.
[0273] Embodiment 8: The difference between this embodiment and any one of Embodiments 1 to 7 is that: in step 33, the filtered genotyped VCF file obtained in step 22 is input into the WhatsHap genotyping tool, and the WhatsHap genotyping tool outputs two genotyping files, namely a haplotype 1 file and a haplotype 2 file.
[0274] The specific process is as follows:
[0275] Step 331: Set the Modki software parameters: pileup, traditional;
[0276] Use the Modki software to convert the filtered genotyped VCF file obtained in step 22 into a tsv format file;
[0277] Step 332: Use the WhatsHap genotyping tool to generate a BED.gz format file from the tsv format file;
[0278] Step 333: Input the genotyped VCF file obtained in step 22 and the BED.gz format file obtained in step 332 into the WhatsHap genotyping tool, and the WhatsHap genotyping tool outputs a haplotype 1 file and a haplotype 2 file.
[0279] Other steps and parameters are the same as those in Embodiments 1 to 7.
[0280] Embodiment 9: The difference between this embodiment and any one of Embodiments 1 to 8 is that: in step 4, the BAM file obtained by step 1 for the dairy cow individual is used as the input of the deep learning model, and the genotyped haplotype 1 file and haplotype 2 file are used as the output of the deep learning model;
[0281] Use the BAM file obtained in step 1 of the dairy cow individual as the input of the large language model, and the phased haplotype 1 file and haplotype 2 file as the output of the large language model;
[0282] Optimize the parameters of the deep learning model and the large language model using a comprehensive loss function, and perform gradient updates in combination with the Adam optimizer until the comprehensive loss function converges to obtain a trained deep learning model and large language model;
[0283] The specific process is as follows:
[0284] Step 41: Construct a deep learning model, which successively includes:
[0285] The first 1×1 convolutional layer, BN layer, first 7×7 depth convolutional layer, first LN layer, second 1×1 convolutional layer, first GELU, first GRN, third 1×1 convolutional layer, second 7×7 depth convolutional layer, second LN layer, fourth 1×1 convolutional layer, second GELU, second GRN, fifth 1×1 convolutional layer, third 7×7 depth convolutional layer, third LN layer, sixth 1×1 convolutional layer, third GELU, third GRN, seventh 1×1 convolutional layer, fully connected layer, softmax layer;
[0286] The working process of the deep learning model is as follows:
[0287] Input the BAM file of the dairy cow individual obtained in step 1 into the first 1×1 convolutional layer and BN layer in sequence, and the BN layer outputs feature A;
[0288] The feature A output by the BN layer is successively input into the first 7×7 depth convolutional layer, first LN layer, second 1×1 convolutional layer, first GELU, first GRN, third 1×1 convolutional layer, and the third 1×1 convolutional layer outputs feature A';
[0289] The feature A' output by the third 1×1 convolutional layer and the feature A output by the BN layer are added element-wise to obtain feature A'';
[0290] Feature A'' is successively input into the second 7×7 depth convolutional layer, second LN layer, fourth 1×1 convolutional layer, second GELU, second GRN, fifth 1×1 convolutional layer, and the fifth 1×1 convolutional layer outputs feature A''';
[0291] The feature A''' output by the fifth 1×1 convolutional layer and the feature A'' are added element-wise to obtain feature
[0292] Feature Is successively input into the third 7×7 depth convolutional layer, third LN layer, sixth 1×1 convolutional layer, third GELU, third GRN, seventh 1×1 convolutional layer, and the seventh 1×1 convolutional layer outputs feature
[0293] Output features of the seventh 1×1 convolutional layer and the features are added element-wise to obtain Feature B;
[0294] Feature B is sequentially input into the fully connected layer and the softmax layer, and the softmax layer outputs the classification result;
[0295] Step 42: Use the large language model to generate the haplotype 1 file and the haplotype 2 file corresponding to the BAM file of the dairy cow individual obtained in Step 1;
[0296] Step 43: Optimize the parameters of the deep learning model and the large language model using the comprehensive loss function, and perform gradient update in combination with the Adam optimizer until the comprehensive loss function converges to obtain the trained deep learning model and large language model.
[0297] Other steps and parameters are the same as those in the first to eighth specific embodiments.
[0298] Specific Embodiment Ten: The difference between this embodiment and any one of the first to ninth specific embodiments is that: the comprehensive loss function is
[0299] where N represents the total number of data in the BAM file of the dairy cow individual, i represents the i-th one; k represents the k-th one;
[0300] F i 1 represents the feature output by the deep learning model when the i-th data in the BAM file of the dairy cow individual is input into the deep learning model;
[0301] F i 2 represents the semantic feature of the text description information output by the large language model when the i-th data in the BAM file of the dairy cow individual is input into the large language model;
[0302] represents the semantic feature of the text description information output by the large language model when the k-th data in the BAM file of the dairy cow individual is input into the large language model;
[0303] s(F i 1 ,F i 2 ) represents the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the i-th data;
[0304] represents the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the k-th data;
[0305] τ represents a temperature hyperparameter.
[0306] Other steps and parameters are the same as those in the first to ninth specific embodiments.
[0307] The present invention may also have various other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention. However, these corresponding changes and modifications should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A typing method for individual BAM files of dairy cows, characterized in that: The specific process of the method is as follows: Step 1: Obtain the BAM file of each dairy cow individual; each BAM file contains M SNP loci. Step 2: Obtain all the clustering cluster results of all variant type signals in the BAM file of each dairy cow individual. The signals contained in each clustering cluster are integrated into a structural variant signal for output. The structural variant signals include insertions, deletions, duplications, translocations, and inversions. Step 3: Perform haplotype typing on each output structural variant signal to obtain the typed haplotype 1 file and haplotype 2 file. Step 4: Use the BAM file of the dairy cow individual obtained in Step 1 as the input of the deep learning model, and the typed haplotype 1 file and haplotype 2 file as the output of the deep learning model. Use the BAM file of the dairy cow individual obtained in Step 1 as the input of the large language model, and the typed haplotype 1 file and haplotype 2 file as the output of the large language model. Adopt a comprehensive loss function to optimize the parameters of the deep learning model and the large language model, and combine the Adam optimizer for gradient update until the comprehensive loss function converges to obtain the trained deep learning model and large language model. Step 5: Input the BAM file of the dairy cow individual to be tested into the trained deep learning model, and the trained deep learning model outputs the typed haplotype 1 file and haplotype 2 file of the BAM file of the dairy cow individual to be tested.
2. The genotyping method of an individual dairy cow BAM file according to claim 1, wherein: In Step 2, obtain all the clustering cluster results of all variant type signals in the BAM file of each dairy cow individual. The signals contained in each clustering cluster are integrated into a structural variant signal for output. The structural variant signals include insertions, deletions, duplications, translocations, and inversions. The specific process is as follows: Step 21: Obtain the clustering cluster results of all signals in the variant signal set of the "insertion" variant type; the specific process is as follows: Step 211: Initialize an empty cluster and add the first signal in the variant signal set of the "insertion" variant type as the starting signal. Step 212: Calculate the similarity between the current signal and the last signal in each existing cluster. Step 213: If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster. If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster. Step 214: Repeat Step 212 - Step 213 until all signals in the variant signal set of the "insertion" variant type are judged; obtain the clustering cluster results of all signals in the variant signal set of the "insertion" variant type. Step 22: Obtain the clustering cluster results of all signals in the variant signal set of the "deletion" variant type; the specific process is as follows: Step 221: Initialize an empty cluster and add the first signal in the variant signal set of the "deletion" variant type as the starting signal. Step 222: Calculate the similarity between the current signal and the last signal in each existing cluster. Step 223: If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster. If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster. Step 224. Repeat Step 222 - Step 223 until all signals in the mutation signal set of the "deletion" mutation type are judged; Step 23. Obtain the clustering results of all signals in the mutation signal set of the "duplication" mutation type; The specific process is as follows: Step 231. Initialize an empty cluster and add the first signal in the mutation signal set of the "duplication" mutation type as the starting signal; Step 232. Calculate the similarity between the current signal and the last signal in each existing cluster; Step 233. If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster; If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster; Step 234. Move to the next signal and repeat Step 232 - Step 233 until all signals in the mutation signal set of the "duplication" mutation type are judged; Step 24. Obtain the clustering results of all signals in the mutation signal set of the "translocation" mutation type; The specific process is as follows: Step 241. Initialize an empty cluster and add the first signal in the mutation signal set of the "translocation" mutation type as the starting signal; Step 242. Calculate the similarity between the current signal and the last signal in each existing cluster; Step 243. If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster; If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster; Step 244. Repeat Step 242 - Step 243 until all signals in the mutation signal set of the "translocation" mutation type are judged; Obtain the clustering results of all signals in the mutation signal set of the "translocation" mutation type; Step 25. Obtain the clustering results of all signals in the mutation signal set of the "inversion" mutation type; The specific process is as follows: Step 251. Initialize an empty cluster and add the first signal in the mutation signal set of the "inversion" mutation type as the starting signal; Step 252. Calculate the similarity between the current signal and the last signal in each existing cluster; Step 253. If the similarity is less than the predetermined threshold, construct a new cluster and add the current signal to the new cluster; If the similarity is greater than or equal to the predetermined threshold, add the current signal to the current cluster; Step 254. Repeat Step 252 - Step 253 until all signals in the mutation signal set of the "inversion" mutation type are judged; Obtain the clustering results of all signals in the mutation signal set of the "inversion" mutation type.
3. The genotyping method of an individual dairy cow BAM file according to claim 2, characterized in that: The process of calculating the similarity between the current signal and the last signal in each existing cluster is as follows: For insertion, deletion, inversion, and duplication mutation signals, obtain a comprehensive similarity score S; expressed as: S = w1 × S1 + w2 × S2 + w3 × S3 Wherein, w1, w2, and w3 are the weight coefficients of the position similarity score, the mutation size similarity score / chromosome similarity score respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, and w1 + w2 + w3 = 1; For translocation mutation signals, obtain a comprehensive similarity score S; expressed as: S = w4 × S4 + w5 × S5 Wherein, w4 is the weight coefficient of the starting chromosome similarity score S4, w4 = 0.5, w5 is the weight coefficient of the ending chromosome similarity score S5, w5 = 0.5, w4 + w5 = 1.
4. The genotyping method of an individual dairy cow BAM file according to claim 3, characterized in that: For the insertion, deletion, inversion, and duplication mutation signals, the comprehensive similarity score S is obtained; it is expressed as: S = w1 × S1 + w2 × S2 + w3 × S3 Wherein, w1, w2, and w3 are the weight coefficients of the position similarity score, the mutation size similarity score / chromosome similarity score respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, w1 + w2 + w3 = 1; The specific process is as follows: 1), Calculate the position similarity S1 score; the specific process is as follows: The calculation formula for the position similarity S1 score is as follows: S1 = |Start1 - Start2| Wherein, Start1 and Start2 are the signal starting positions in the two mutation signals respectively; 2), Calculate the mutation signal size similarity score S2; the specific process is as follows: Denote the mutation lengths of the two mutation signals as SVlen1 and SVlen2 respectively, and the mutation signal size similarity score S2 is calculated by the following formula: 3), Calculate the chromosome similarity score S3; the specific process is as follows: Suppose the chromosome numbers where the two mutation signals are located are Chrom1 and Chrom2 respectively, and the chromosome similarity score S3 is calculated by the following discrete function: When the chromosome names where the two mutation signals are located are the same, the chromosome similarity score is 1; When the chromosome names where the two mutation signals are located are different, the chromosome similarity score is 0; 4), Perform weighted summation on the position similarity S1 score, the mutation size similarity score S2 score, and the chromosome similarity score S3 to obtain the comprehensive similarity score S; it is expressed as: S = w1 × S1 + w2 × S2 + w3 × S3 Wherein, w1, w2, and w3 are the weight coefficients of the position similarity score, the mutation size similarity score / chromosome similarity score respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, w1 + w2 + w3 = 1.
5. A genotyping method for individual dairy cow BAM files according to claim 4, characterized in that: For the translocation mutation signal, the comprehensive similarity score S is obtained; it is expressed as: S = w4 × S4 + w5 × S5 Wherein, w4 is the weight coefficient of the starting chromosome similarity score S4, w4 = 0.5, w5 is the weight coefficient of the ending chromosome similarity score S5, w5 = 0.5, w4 + w5 = 1; The specific process is as follows: Calculate the starting chromosome similarity score: Wherein, S4 represents the starting chromosome similarity score; Indicates the starting chromosome name of the first signal; Indicates the starting chromosome name of the second signal; Calculate the ending chromosome similarity score: Wherein, S5 represents the ending chromosome similarity score; The name of the terminating chromosome representing the first signal; The name of the terminal chromosome representing the second signal; Perform weighted summation on the starting chromosome similarity score S4 and the ending chromosome similarity score S5 to obtain the comprehensive similarity score S; it is expressed as: S = w4 × S4 + w5 × S5 Wherein, w4 is the weight coefficient of the starting chromosome similarity score S4, w4 = 0.5, w5 is the weight coefficient of the ending chromosome similarity score S5, w5 = 0.5, w4 + w5 = 1.
6. The genotyping method of an individual dairy cow BAM file according to claim 5, characterized in that: In S3, haplotype typing is performed on each output structural variation signal to obtain the typed hap1 file and hap2 file; the specific process is as follows: Step 31: Input the VCF file of each structural variation signal into a variant detection tool; The variant detection tool outputs a typed VCF file, and the typed VCF file contains M SNP sites; The typed VCF file contains 8 columns of information, namely: Chromosome name, SNP position, SNP ID, reference gene, alternative allele, quality value, filter flag, annotation information column; Step 32: Use the Bcftools software to filter the typed VCF file to obtain a filtered typed VCF file; Step 33: Input the filtered typed VCF file obtained in step 22 into the WhatsHap typing tool, and the WhatsHap typing tool outputs two typed files, namely the haplotype 1 file and the haplotype 2 file.
7. A genotyping method for individual dairy cow BAM files according to claim 6, characterized in that: In step 32, the Bcftools software is used to filter the typed VCF file to obtain a filtered typed VCF file; the specific process is as follows: Retain the rows in the typed VCF file where the "filter flag" value is equal to PASS; Delete the rows in the typed VCF file where the "filter flag" value is not equal to PASS; Finally, a typed VCF file is generated; Each SNP site in the typed VCF file is divided into haplotype 1 and haplotype 2.
8. A genotyping method for individual dairy cow BAM files according to claim 7, characterized in that: In step 33, the filtered typed VCF file obtained in step 22 is input into the WhatsHap typing tool, and the WhatsHap typing tool outputs two typed files, namely the haplotype 1 file and the haplotype 2 file; the specific process is as follows: Step 331: Set the Modki software parameters: pileup, traditional; Use the Modki software to convert the filtered typed VCF file obtained in step 22 into a tsv format file; Step 332: Use the WhatsHap typing tool to generate a BED.gz format file from the tsv format file; Step 333: Input the typed VCF file obtained in step 22 and the BED.gz format file obtained in step 332 into the WhatsHap typing tool, and the WhatsHap typing tool outputs the haplotype 1 file and the haplotype 2 file.
9. A genotyping method for individual dairy cow BAM files according to claim 8, characterized in that: In step 4, the BAM file of the dairy cow individual obtained in step 1 is used as the input of the deep learning model, and the typed haplotype 1 file and haplotype 2 file are used as the output of the deep learning model; Use the BAM file of the dairy cow individual obtained in step 1 as the input of the large language model, and the typed haplotype 1 file and haplotype 2 file are used as the output of the large language model; Adopt a comprehensive loss function to optimize the parameters of the deep learning model and the large language model, and combine the Adam optimizer for gradient update until the comprehensive loss function converges to obtain a trained deep learning model and large language model; The specific process is as follows: Step 41: Construct a deep learning model, and the deep learning model successively includes: The first 1×1 convolutional layer, BN layer, first 7×7 depth convolutional layer, first LN layer, second 1×1 convolutional layer, first GELU, first GRN, third 1×1 convolutional layer, second 7×7 depth convolutional layer, second LN layer, fourth 1×1 convolutional layer, second GELU, second GRN, fifth 1×1 convolutional layer, third 7×7 depth convolutional layer, third LN layer, sixth 1×1 convolutional layer, third GELU, third GRN, seventh 1×1 convolutional layer, fully connected layer, softmax layer; The working process of the deep learning model is as follows: The BAM files of the dairy cow individuals obtained in step 1 are sequentially input into the first 1×1 convolutional layer and the BN layer, and the BN layer outputs feature A; The feature A output by the BN layer is sequentially input into the first 7×7 depth convolutional layer, the first LN layer, the second 1×1 convolutional layer, the first GELU, the first GRN, and the third 1×1 convolutional layer, and the third 1×1 convolutional layer outputs feature A'; The feature A' output by the third 1×1 convolutional layer and the feature A output by the BN layer are element-wise added to obtain feature A''; The feature A'' is sequentially input into the second 7×7 depth convolutional layer, the second LN layer, the fourth 1×1 convolutional layer, the second GELU, the second GRN, and the fifth 1×1 convolutional layer, and the fifth 1×1 convolutional layer outputs feature A''; The output feature A″ of the fifth 1×1 convolutional layer is element-wise added to the feature A″ to obtain the feature Feature Successively input the third 7×7 depth convolution layer, the third LN layer, the sixth 1×1 convolution layer, the third GELU, the third GRN, and the seventh 1×1 convolution layer. The output of the seventh 1×1 convolution layer is the feature Output features of the seventh 1×1 convolutional layer and the features are element-wise added to obtain Feature B; The feature B is sequentially input into the fully connected layer and the softmax layer, and the softmax layer outputs the classification result; Step 42: Use the large language model to generate the haplotype 1 file and haplotype 2 file corresponding to the BAM file of the dairy cow individual obtained in step 1; Step 43: Optimize the parameters of the deep learning model and the large language model using the comprehensive loss function, and perform gradient update in combination with the Adam optimizer until the comprehensive loss function converges to obtain the trained deep learning model and large language model.
10. A typing method for individual dairy cow BAM files according to claim 9, characterized in that: The comprehensive loss function is Where N represents the total number of data in the BAM file of the dairy cow individual, i represents the i-th one; k represents the k-th one; The i-th data in the BAM file representing an individual dairy cow is input into a deep learning model, and the features output by the deep learning model; The semantic features of the text description information output by the large language model for the i-th data input from the data in the BAM file representing an individual dairy cow are input into the large language model. The k-th data in the BAM file representing an individual dairy cow is input into the large language model, and the semantic features of the text description information output by the large language model; Indicates the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the i-th data; Indicates the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the k-th data; τ represents a temperature hyperparameter.
Citation Information
Patent Citations
Deep learning-based tetraploid oyster whole genome SNP (Single Nucleotide Polymorphism) typing method
CN117637020A
Mutation detection method based on end-to-end assembly genome
CN119785877A
Method of determining kinship using gene sequence variation
KR102391084B1
Sparse coding and extraction of ultrasound knowledge for explainable point-of-care ultrasound artificial intelligence
WO2024097623A1