A genotype detection algorithm system and electronic device for a genetic disease in dogs

By constructing a genotype detection algorithm system for canine genetic diseases, the entire process from raw data to detection results has been optimized. This solves the problems of insufficient correlation between variant site screening and genotyping and the single dimension of pathogenic gene association analysis, improving the accuracy and efficiency of detection and providing reliable technical support for the diagnosis and breeding of canine genetic diseases.

CN122157771AInactive Publication Date: 2026-06-05AGSINO GENSOURCES CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AGSINO GENSOURCES CO LTD
Filing Date
2026-03-02
Publication Date
2026-06-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies for genotyping of canine genetic diseases lack sufficient correlation between variant site screening and genotyping, and the pathogenic gene association analysis is limited in scope, resulting in insufficient reliability and accuracy of test results, which cannot meet the needs of large-scale screening and diagnosis of complex diseases.

Method used

A genotype detection algorithm system for canine genetic diseases was constructed, including a gene sequence feature extraction module, a site polymorphism identification module, a variant site screening module, a genotyping calculation module, and a pathogenic gene association module. Through multi-dimensional feature extraction and probability calculation, combined with pathogenic gene database access, sequence homology analysis, and functional pathway matching, the entire process from raw data to detection results was optimized.

Benefits of technology

It significantly improves the accuracy and efficiency of detection, reduces false positive screening and typing errors, comprehensively explores the relationship between genes and diseases, and provides comprehensive and reliable technical support for the accurate diagnosis, breeding optimization and disease prevention and control of canine genetic diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122157771A_ABST
    Figure CN122157771A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of canine genotype detection, and discloses a canine genetic disease genotype detection algorithm system and electronic equipment, which comprises a gene sequence feature extraction module, a site polymorphism identification module, a variation site screening module, a genotype grouping operation module, a pathogenic gene correlation module and a detection result output module. Each module is sequentially connected to realize efficient data transmission and processing. The variation site screening module, the genotype grouping operation module and the pathogenic gene correlation module each comprise four functional units, which are used for completing site screening, genotype grouping and pathogenic gene correlation analysis through multi-step fine processing. The system is provided with a built-in special analysis model, can integrate multidimensional data features to carry out detection operation, realizes whole-process processing from raw data to a standardized detection report, significantly improves the accuracy and comprehensiveness of canine genetic disease genotype detection, and provides reliable technical support for canine genetic disease diagnosis, breeding optimization and prevention and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of canine genotype detection technology, and in particular to a genotype detection algorithm system and electronic device for canine genetic diseases. Background Technology

[0002] With the increasing popularity of dogs as companion animals, the impact of genetic diseases on canine health and the quality of life of their owners is becoming increasingly prominent. Accurate genotyping of canine genetic diseases has become a key requirement for veterinary diagnosis, breeding optimization, and disease prevention and control. Currently, there are numerous types of canine genetic diseases, most of which are caused by mutations at specific gene loci. Traditional detection methods rely on targeted sequencing, which struggles to achieve simultaneous detection of multiple genes and loci, failing to meet the needs of large-scale screening and complex disease diagnosis. The development of high-throughput sequencing technology has provided a data foundation for genome-wide genotyping. However, how to efficiently extract features from massive sequencing data, accurately identify mutation sites, accurately classify genotypes, and associate them with pathogenic genes has become the core issue restricting detection efficiency and accuracy. There is an urgent need to build a systematic and professional genotyping algorithm system to provide end-to-end technical support from raw data to detection results.

[0003] Existing technologies for genotyping in canine genetic diseases have two significant drawbacks: First, the correlation between variant screening and genotyping is insufficient. Current technologies often conduct variant screening and genotyping independently, failing to fully utilize the biological association information between the two. This results in low matching between screened variants and genotyping results, leading to false positives or genotyping errors, which affect the reliability of subsequent pathogenic gene association analysis. Second, pathogenic gene association analysis is limited in scope. Existing methods often rely solely on sequence homology for association judgment, failing to integrate multi-dimensional information such as gene functional pathways and population frequency characteristics. This makes it difficult to comprehensively assess the association between genes and diseases, resulting in the omission of some low-homology but functionally critical pathogenic genes, thus failing to provide comprehensive and accurate genotyping evidence for disease diagnosis. Summary of the Invention

[0004] In order to overcome the shortcomings and deficiencies of the existing technology, the present invention provides a genotype detection algorithm system and electronic device for canine genetic diseases.

[0005] The technical solution adopted in this invention is a genotype detection algorithm system for canine genetic diseases, comprising a gene sequence feature extraction module, a site polymorphism identification module, a variant site screening module, a genotype typing calculation module, a pathogenic gene association module, and a detection result output module.

[0006] The gene sequence feature extraction module captures features from the raw canine genome sequencing data and transmits them to the locus polymorphism identification module. The locus polymorphism identification module analyzes the gene locus polymorphism features and transmits them to the variant site screening module. The variant site screening module identifies potential pathogenic variant sites and sends them to the genotyping calculation module. The genotyping calculation module performs genotyping calculations on the canine genotypes and sends the results to the pathogenic gene association module. The pathogenic gene association module performs association analysis between the genotyping results and known pathogenic genes and transmits the results to the detection result output module. The detection result output module formats and outputs the detection data after association analysis.

[0007] Furthermore, the variant site screening module includes: a sequence alignment unit, a site frequency statistics unit, a variant harmfulness prediction unit, and a screening threshold determination unit. The sequence alignment unit performs base-by-base alignment between the canine gene sequence and the reference genome sequence to obtain site difference information. The site frequency statistics unit performs statistical analysis on the frequency of occurrence of the differential sites in the population. The variant harmfulness prediction unit predicts and evaluates the biological functional impact of the differential sites. The screening threshold determination unit sets a screening threshold based on the frequency statistics results and the harmfulness prediction evaluation results and screens out variant sites that meet the threshold requirements.

[0008] Furthermore, the genotyping calculation module includes: a sequencing data denoising unit, an allele identification unit, a genotype probability calculation unit, and a genotyping result determination unit. The sequencing data denoising unit removes noise from the input raw sequencing data. The allele identification unit identifies the sequence characteristics of different alleles from the denoised data. The genotype probability calculation unit calculates the probability of different genotype combinations based on the allele sequence characteristics. The genotyping result determination unit selects the genotype with the highest probability as the final genotyping result based on the probability calculation result.

[0009] Furthermore, the pathogenic gene association module includes: a pathogenic gene database retrieval unit, a sequence homology analysis unit, a functional pathway matching unit, and an association degree quantification unit. The pathogenic gene database retrieval unit retrieves a preset canine genetic disease pathogenic gene database. The sequence homology analysis unit performs homology comparison between the gene sequence corresponding to the typing result and the pathogenic gene sequence in the database. The functional pathway matching unit matches gene sequences with high homology with known pathogenic functional pathways. The association degree quantification unit quantifies the matching results to obtain the association degree data between the gene and the disease.

[0010] Furthermore, the gene sequence feature extraction module employs a feature extraction model: ,in, The extracted gene sequence comprehensive feature value, For the first Feature weights of each gene locus For the first Sequencing signal intensity at each site, For the first The reference genomic base values ​​for each locus, where ⊕ represents the base feature XOR operation. For the first Sequencing coverage depth at each site, The regularization coefficient is . This is the weight matrix. For the weight matrix Norm, This represents the total number of sites in the gene sequence.

[0011] Furthermore, the site polymorphism identification module employs a polymorphism identification model: ,in, This is a site polymorphism identification index. For the first Locus heterozygosity markers for each sequencing fragment For the first The quality value of each sequencing fragment, This is the distance attenuation coefficient. For the first The base mismatch distance between each sequencing fragment and the reference sequence The coefficient of variation is 1. For the first Allele frequencies at each locus The average allele frequency across all loci. This represents the total number of sequencing fragments covering this site.

[0012] Furthermore, the variant site screening module employs a screening model: ,in, The score is used to screen for variant sites. This is the weighting coefficient for harmfulness. Here, represents the predicted harmfulness of the variant site, and Freq represents the frequency of the variant site in the population. It is the minimum value. is the conservation weighting coefficient, Cons is the sequence conservation score of the variant site, and Len is the length of the coding region of the gene containing the variant.

[0013] Furthermore, the genotyping module employs a genotyping calculation model: ,in, For the final genotype, For the set of all possible genotypes, The number of fractal feature dimensions. For the first The variance of each feature, Genotype In the Observations on each feature Genotype In the The mean of each feature Genotype The prior probability;

[0014] Furthermore, the pathogenic gene association module employs an association analysis model: ,in, The correlation coefficient between genes and diseases. The number of functional feature dimensions. For the gene to be detected in the first The numerical values ​​of each functional characteristic For known disease-causing genes, in the first... The numerical values ​​of each functional characteristic This is the correlation strength adjustment coefficient. The gene to be detected and the pathogenic gene were compared at the first... The number of matches on each functional feature For the first The total number of features for each functional feature.

[0015] A genotyping algorithm system for canine genetic diseases comprises the following steps: First, peripheral blood samples from dogs are obtained and genomic DNA is extracted. The genomic DNA is then sequenced using high-throughput sequencing technology to obtain raw sequencing data. Second, the raw sequencing data is input into a gene sequence feature extraction module, where a feature capture algorithm extracts features such as base arrangement patterns, site mutation characteristics, and sequence length distribution from the gene sequence. Third, the extracted feature data is transmitted to a site polymorphism identification module, where a polymorphism analysis algorithm identifies allelic variation types and variation site distribution patterns at gene sites. Fourth, the identified polymorphic site data is sent to a variation site screening module, where a site screening algorithm screens variation sites with potential pathogenicity. Fifth, the screened variation site data is transmitted to a genotyping calculation module and a pathogenic gene association module, sequentially completing genotyping calculation and pathogenic gene association analysis. Sixth, the association analysis results are transmitted to a detection result output module, where a data format normalization algorithm standardizes the results, generating and outputting a detection report including genotype information and pathogenic gene association data.

[0016] An electronic device includes a memory and a processor, the memory storing a computer program and the processor being configured to execute the computer program to implement a genotype detection algorithm system for canine genetic diseases.

[0017] Beneficial Effects: This invention proposes a genotype detection algorithm system and electronic device for canine genetic diseases. By constructing a complete technical system with six modules working collaboratively, it achieves full-process optimization of canine genetic disease genotype detection, significantly improving the accuracy and efficiency of detection. Through the coherent connection of gene sequence feature extraction, site polymorphism identification, variant site screening, genotyping calculation, pathogenic gene association, and detection result output, variant site screening and genotyping are closely integrated. The biological correlation information between the two is fully utilized. Through multi-dimensional feature extraction and probability calculation, the matching degree between variant sites and genotyping results is greatly improved, effectively reducing false positive screening and genotyping errors, and successfully overcoming the reliability problem caused by the independent operation of the two in traditional technologies. Meanwhile, the pathogenic gene association module integrates multiple functions such as pathogenic gene database access, sequence homology analysis, functional pathway matching, and association degree quantification. It breaks through the limitations of single sequence homology judgment, incorporates multi-dimensional information such as gene functional pathways and population frequency characteristics, comprehensively explores the association between genes and diseases, avoids the omission of pathogenic genes with low homology but key functions, and completely solves the shortcomings of the single dimension of the existing association analysis. It provides comprehensive and reliable technical support for the accurate diagnosis, breeding optimization, and disease prevention and control of canine genetic diseases. Attached Figure Description

[0018] Figure 1 This is a diagram showing the system module composition of the present invention;

[0019] Figure 2 This is a flowchart of the system operation steps of the present invention;

[0020] Figure 3 This is a diagram illustrating the electronic device components of the present invention. Detailed Implementation

[0021] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] like Figure 1 As shown, a genotype detection algorithm system for canine genetic diseases includes: a gene sequence feature extraction module, a site polymorphism identification module, a variant site screening module, a genotype typing calculation module, a pathogenic gene association module, and a detection result output module.

[0023] The gene sequence feature extraction module captures features from the raw canine genome sequencing data and transmits them to the locus polymorphism identification module. The locus polymorphism identification module analyzes the gene locus polymorphism features and transmits them to the variant site screening module. The variant site screening module identifies potential pathogenic variant sites and sends them to the genotyping calculation module. The genotyping calculation module performs genotyping calculations on the canine genotypes and sends the results to the pathogenic gene association module. The pathogenic gene association module performs association analysis between the genotyping results and known pathogenic genes and transmits the results to the detection result output module. The detection result output module formats and outputs the detection data after association analysis.

[0024] The gene sequence feature extraction module, serving as the initial stage of the system's data processing, receives raw canine genome data generated by high-throughput sequencing. This data covers approximately 2.4 billion base pairs across the entire canine genome, with a sequencing depth set at 30-50× to ensure data reliability. This module employs a feature capture algorithm to perform multi-dimensional analysis of the raw data, focusing on extracting core features for each gene locus, including sequencing signal intensity, reference genome base matching, and sequencing coverage depth. A signal intensity threshold greater than 30dB is set to filter low-quality signals, and loci with a coverage depth less than 10× are marked as low-confidence loci. During implementation, the module assigns a feature weight of 0.1-0.9 to each locus. This weight allocation is determined based on the locus's conservation among genes related to canine genetic diseases; higher conservation results in a larger weight. Regularization is also used to control feature dimensionality and avoid data redundancy. The core significance of this module lies in transforming massive amounts of raw sequencing data into structured feature data that can be used for subsequent analysis. By accurately extracting key information from gene sequences, it provides high-quality data input for subsequent steps such as site polymorphism identification and variant screening. Its processing efficiency reaches 1 million base pairs per second, ensuring the efficient advancement of the overall detection process of the system.

[0025] The site polymorphism identification module receives structured feature data output from the gene sequence feature extraction module and analyzes approximately 150 million potential polymorphic sites in the canine genome. During implementation, the module first performs quality screening on the sequencing fragments at each site, retaining fragments with a quality score greater than 30. Each site must meet the condition of being covered by at least 15 valid sequencing fragments. The module analyzes the base variations in the sequencing fragments using a polymorphism analysis algorithm, identifying allele variation types, including single nucleotide polymorphisms and insertion / deletion variations. It also calculates the allele frequency at each variation site and determines the site's heterozygosity. The module sets a distance attenuation coefficient to correct for the impact of base mismatch distances; sequencing fragments with mismatch distances exceeding 3 bases are assigned lower weights. Furthermore, the module standardizes the degree of variation at individual sites by calculating the average allele frequency across all sites. The significance of this module lies in accurately identifying polymorphic sites in the genome, eliminating false positive variants caused by sequencing errors, and controlling the identification accuracy rate to over 99%. It provides an accurate set of polymorphic sites for subsequent variant site screening, ensuring that subsequent analysis is carried out only on truly existing variant sites and avoiding unnecessary computational waste.

[0026] The variant site screening module receives polymorphic site data identified by the polymorphism identification module and, through the collaborative efforts of four functional units, identifies potential pathogenic variant sites. During implementation, the sequence alignment unit performs base-by-base alignment of the canine gene sequence with the canine reference genome sequence, achieving alignment accuracy at the single-base level to obtain site difference information. The site frequency statistics unit counts the frequency of differentially expressed sites in a canine population database containing more than 10,000 individuals, setting a population frequency threshold of less than 1%. Common variant sites with frequencies exceeding this threshold are initially excluded from pathogenicity. The variant harmfulness prediction unit analyzes the impact of differentially expressed sites on the structure and function of gene-encoded proteins, providing a harmfulness prediction value between 0 and 1. Sites with prediction values ​​greater than 0.7 are marked as highly harmful. The screening threshold determination unit combines the frequency statistics results with the harmfulness prediction value to set a comprehensive screening threshold, retaining only variant sites with a population frequency below 1% and a harmfulness prediction value greater than 0.7. Simultaneously, the sequence conservation score of the variant site is considered, with sites having a conservation score higher than 80 being prioritized for retention. The significance of this module lies in its ability to accurately screen out variant sites with potential pathogenic risks from a massive number of polymorphic sites, achieving a screening efficiency of 500,000 sites per hour, effectively reducing the amount of data for subsequent analysis and improving the system's targeting and accuracy.

[0027] The genotyping module receives pathogenic variant candidate site data output by the variant site screening module and performs refined genotyping calculations for each candidate site. During implementation, the sequencing data denoising unit uses a signal filtering algorithm to remove background noise from the sequencing data. The noise filtering threshold is set to 10% of the signal intensity; signals below this threshold are identified as noise and removed. The allele identification unit extracts sequence features of different alleles from the denoised data, including base composition and length differences. Each allele must meet the condition supported by at least five independent sequencing fragments. The genotype probability calculation unit calculates the probability of homozygous dominant, homozygous recessive, and heterozygous genotype combinations based on allele sequence features and the known genotype frequency distribution in the population. The genotyping result determination unit selects the genotype with the highest probability as the final genotyping result. When the difference between the highest and second-highest probabilities is less than 10%, it is marked as genotyping uncertain and requires further verification based on subsequent association analysis results. The significance of this module lies in enabling accurate genotyping of candidate sites for pathogenic variants, with a typing accuracy rate of over 98.5%, providing core genotyping data support for pathogenic gene association analysis.

[0028] The pathogenic gene association module receives the genotype results output by the genotyping module and performs association analysis in conjunction with a pre-set database of pathogenic genes for canine genetic diseases. This database includes over 8,000 pathogenic gene sequences and functional information corresponding to more than 1,200 common canine genetic diseases, and is updated quarterly. During implementation, the pathogenic gene database retrieval unit retrieves relevant information of homologous genes from the database based on the gene name and location corresponding to the typing results. The sequence homology analysis unit uses a global alignment algorithm to compare the typing gene sequences with the pathogenic gene sequences in the database, and sequences with homology higher than 85% are included in subsequent analysis. The functional pathway matching unit matches gene sequences with high homology with known pathogenic functional pathways, including more than 200 core pathways such as metabolic pathways and signal transduction pathways. The association degree quantification unit quantifies the matching results, assigning weights of 0.4, 0.3, and 0.3 to three dimensions: sequence homology, functional relevance, and pathway matching degree, respectively, and calculates the association coefficient between the gene and the disease. An association coefficient higher than 0.6 is considered a strong association, 0.4-0.6 is a moderate association, and lower than 0.4 is a weak association. The significance of this module lies in establishing the association between genotypes and canine genetic diseases, improving the reliability of association judgment through multi-dimensional quantitative analysis, providing clear gene-level evidence for disease diagnosis, and achieving a specificity of over 97% in the association analysis.

[0029] The detection result output module receives the association analysis data output by the pathogenic gene association module and, as the system's terminal link, is responsible for data formatting and final output. During implementation, the module first standardizes the association analysis data, integrating multi-dimensional data such as genotype information, pathogenic gene association degree, population frequency, and harmfulness prediction value into a unified format, retaining data precision to four decimal places. The module sets a fixed report template, including core sections such as basic sample information, detection site details, genotype results, pathogenic gene association analysis, and detection conclusions. The detection site details list parameters such as the chromosomal location, reference bases, variant bases, and sequencing coverage depth for each site. The association analysis section clearly indicates the association coefficient and association strength level, along with the pathogenic gene database version information and the algorithm parameters used in the analysis. The output format supports two common formats: PDF and Excel, which users can choose. The output file size is controlled within 5MB to facilitate transmission and storage. The significance of this module lies in transforming complex test and analysis data into clear, standardized, and easy-to-interpret test reports, ensuring that users such as breeders and veterinarians can quickly obtain key information. Its data format is standardized and the accuracy rate reaches 100%, providing direct technical support for the diagnosis, treatment, and breeding decisions of canine genetic diseases.

[0030] Preferably, the variant site screening module includes: a sequence alignment unit, a site frequency statistics unit, a variant harmfulness prediction unit, and a screening threshold determination unit. The sequence alignment unit performs base-by-base alignment between the canine gene sequence and the reference genome sequence to obtain site difference information. The site frequency statistics unit performs statistical analysis on the frequency of occurrence of the differential sites in the population. The variant harmfulness prediction unit predicts and evaluates the biological functional impact of the differential sites. The screening threshold determination unit sets a screening threshold based on the frequency statistics results and the harmfulness prediction evaluation results and screens out variant sites that meet the threshold requirements.

[0031] Specifically, the variant site screening module achieves accurate identification of pathogenic variant sites through the coordinated operation of four functional units, and its implementation process strictly follows multi-dimensional verification logic. The sequence alignment unit uses a global alignment algorithm to perform base-by-base alignment of the canine gene sequence with the latest version of the canine reference genome sequence. During the alignment process, a base matching threshold of 99% is set, and mismatched bases are marked and the difference type is recorded to ensure the completeness and accuracy of the site difference information. The site frequency statistics unit retrieves a population gene database including more than 10,000 dogs of different breeds and regions, and counts the frequency of differentially expressed sites in the population. The sample size is allocated according to a standard of at least 500 individuals per breed to ensure the representativeness of the frequency data. The variant harmfulness prediction unit, based on a gene function annotation database, evaluates differentially expressed sites from three dimensions: amino acid sequence changes, protein structural stability, and impact on gene expression regulation. Each dimension has a scoring range of 0-100, and sites with a comprehensive score higher than 70 are judged as highly harmful potential sites. The screening threshold determination unit combines frequency statistics and harmfulness prediction scores, setting a population frequency threshold of 1% and a harmfulness score threshold of 70. Only variant sites with a frequency below 1% and a score above 70 are retained. Sequence conservation score is introduced as an auxiliary screening condition, with sites with a conservation score above 85 being given priority for inclusion in the candidate set. This module effectively eliminates harmless variants and false positive sites through multi-unit layer-by-layer screening, providing highly reliable candidate site data for subsequent genotyping and pathogenic gene association analysis.

[0032] Preferably, the genotyping calculation module includes: a sequencing data denoising unit, an allele identification unit, a genotype probability calculation unit, and a genotyping result determination unit. The sequencing data denoising unit removes noise from the input raw sequencing data. The allele identification unit identifies the sequence characteristics of different alleles from the denoised data. The genotype probability calculation unit calculates the probability of occurrence of different genotype combinations based on the allele sequence characteristics. The genotyping result determination unit selects the genotype with the highest probability as the final genotyping result based on the probability calculation result.

[0033] Specifically, the genotyping module achieves accurate genotyping of candidate variant sites through the orderly connection of four functional units. The sequencing data denoising unit uses a sliding window filtering algorithm to process the input raw sequencing data. The window size is set to 5 bases. Bases with a signal intensity below 30dB within the window are identified as noise and corrected. Simultaneously, sequencing fragments with a quality value below 20 are removed to ensure the reliability of the input data. The allele identification unit extracts allele sequence features from the denoised data, including key information such as base composition, variant site location, and sequence length. Each allele must meet the condition of being supported by at least 8 independent sequencing fragments, and the fragment coverage must be evenly distributed within 50 bases upstream and downstream of the variant site. The genotype probability calculation unit, based on allele sequence features and combined with a pre-set population genotype frequency database, uses a Bayesian probability model to calculate the probability of homozygous dominant, homozygous recessive, and heterozygous genotype combinations. A sequencing error rate correction coefficient (range 0.001-0.01) is introduced during the calculation to correct the probability results. The genotyping result determination unit sorts the probability values ​​of the three genotypes and selects the genotype with the highest probability as the final genotyping result. When the difference between the highest probability and the second highest probability is less than 15%, it is marked as genotyping uncertainty and a probability confidence interval is attached. This module ensures the accuracy of genotyping results through refined data processing and probability calculation. The genotyping accuracy rate is stable at over 98%, providing core data support for pathogenic gene association analysis.

[0034] Preferably, the pathogenic gene association module includes: a pathogenic gene database retrieval unit, a sequence homology analysis unit, a functional pathway matching unit, and an association degree quantification unit. The pathogenic gene database retrieval unit retrieves a preset canine genetic disease pathogenic gene database. The sequence homology analysis unit performs homology comparison between the gene sequence corresponding to the typing result and the pathogenic gene sequence in the database. The functional pathway matching unit matches gene sequences with high homology with known pathogenic functional pathways. The association degree quantification unit quantifies the matching results to obtain the association degree data between the gene and the disease.

[0035] Specifically, the pathogenic gene association module establishes the association between genotypes and canine genetic diseases through the synergistic effect of four functional units. The pathogenic gene database retrieval unit accesses a pre-defined database of pathogenic genes for canine genetic diseases. This database includes over 8,500 pathogenic gene sequences, functional annotations, and clinical phenotype information corresponding to over 1,200 common canine genetic diseases. The database is updated quarterly to ensure data timeliness and completeness. The sequence homology analysis unit uses the Smith-Waterman local alignment algorithm to perform homology comparisons between the gene sequences corresponding to the genotyping results and the pathogenic gene sequences in the database. The alignment window size is set to 10 bases. Sequences with homology scores higher than 85 are included in subsequent analysis, and key parameters such as the location, length, and matching base ratio of homologous regions are recorded. The functional pathway matching unit matches gene sequences with high homology with over 200 pathogenic functional pathways in the Kyoto Encyclopedia of Genomes and Genomes (KEGG), including metabolic pathways, signal transduction pathways, and immune regulation pathways. During the matching process, the key roles and regulatory relationships of genes within these pathways are emphasized. The association degree quantification unit quantifies the matching results from three dimensions: sequence homology, functional relevance, and pathway matching degree. The weights of the three dimensions are set to 0.4, 0.3, and 0.3, respectively. The association coefficient between the gene and the disease is calculated (the value ranges from 0 to 1). An association coefficient higher than 0.6 is considered a strong association, 0.4-0.6 is considered a moderate association, and a value lower than 0.4 is considered a weak association. This module comprehensively explores the intrinsic relationship between genotype and disease through multi-dimensional association analysis and quantitative evaluation, providing a scientific basis for the diagnosis of canine genetic diseases.

[0036] Preferably, the gene sequence feature extraction module employs a feature extraction model: ,in, The extracted gene sequence comprehensive feature value, For the first Feature weights of each gene locus For the first Sequencing signal intensity at each site, For the first The reference genomic base values ​​for each locus, where ⊕ represents the base feature XOR operation. For the first Sequencing coverage depth at each site, The regularization coefficient is . This is the weight matrix. For the weight matrix Norm, This represents the total number of sites in the gene sequence.

[0037] Specifically, the gene sequence feature extraction module employs a feature extraction model. During implementation, the total number of loci in the gene sequence is determined based on the actual situation of the canine whole genome, typically ranging from 1000 to 10000 loci for a single gene. The feature weight of each gene locus is allocated according to its functional importance in canine genetic disease-related genes, with a value range of 0.1-0.9. Specifically, the weight of functionally critical loci, such as core loci in coding regions, is set at 0.7-0.9, while the weight of non-critical loci in non-coding regions is set at 0.1-0.3. The sequencing signal intensity of each locus is directly acquired using sequencing equipment, with a value range of 0-100 dB. The base values ​​of the reference genome are digitally encoded according to the base types of the standard canine reference genome. Base feature XOR operations quantify features by comparing the differences between the sequenced bases and the reference bases. The sequencing coverage depth is set to 30-50× to ensure sufficient data support for each locus. The regularization coefficient, ranging from 0.001 to 0.01, controls the complexity of the weight matrix to prevent overfitting. The dimension of the weight matrix is ​​consistent with the total number of loci in the gene sequence, and the weights are constrained by calculating the L2 norm of the weight matrix. This model transforms the multidimensional raw data of the gene sequence into unified comprehensive feature values ​​through multi-parameter collaborative computation. During implementation, processing 1 million loci takes no more than 30 seconds, ensuring both comprehensive feature extraction and processing efficiency. This provides high-quality feature input for subsequent modules such as locus polymorphism identification and variant screening, ensuring the accuracy and efficiency of the entire detection system.

[0038] Preferably, the site polymorphism identification module employs a polymorphism identification model: ,in, This is a site polymorphism identification index. For the first Locus heterozygosity markers for each sequencing fragment For the first The quality value of each sequencing fragment, This is the distance attenuation coefficient. For the first The base mismatch distance between each sequencing fragment and the reference sequence The coefficient of variation is 1. For the first Allele frequencies at each locus The average allele frequency across all loci. This represents the total number of sequencing fragments covering this site.

[0039] Specifically, the site polymorphism identification module employs a polymorphism identification model. During implementation, the total number of valid sequencing fragments covering each site is first determined, with the total number of valid sequencing fragments for a single site controlled between 15 and 50 to ensure data reliability and representativeness. The site heterozygosity flag for each sequencing fragment is assigned a value based on actual conditions: 1 for heterozygous sites and 0 for homozygous sites. The quality value of the sequencing fragment is output by the sequencing equipment, ranging from 0 to 60. Fragments with quality values ​​below 20 are marked as low quality and given lower weight in the calculation. The distance attenuation coefficient is set at 0.1-0.5 to correct for the impact of base mismatch distance. The base mismatch distance between each sequencing fragment and the reference sequence is calculated based on the actual number of mismatched bases. Fragments with mismatch distances exceeding 3 bases will have their influence in the calculation reduced using the distance attenuation coefficient. The coefficient of variation is set at 0.2-0.8. The allele frequency at each site is obtained by statistically analyzing the proportion of different alleles at that site in the valid sequencing fragments. The average allele frequency across all sites is calculated by averaging the allele frequencies of all polymorphic sites throughout the genome. This model calculates a site polymorphism identification index by integrating multi-dimensional data such as sequencing fragment heterozygosity, quality value, mismatch distance, and allele frequency. The identification index ranges from 0 to 10, and sites with an index higher than 6 are identified as high-confidence polymorphic sites. During implementation, the time taken to process 10,000 sites is no more than 20 seconds, which not only ensures the accuracy of polymorphic site identification but also improves the identification efficiency, providing a precise set of candidate sites for subsequent variant site screening.

[0040] Preferably, the variant site screening module employs a screening model: ,in, The score is used to screen for variant sites. This is the weighting coefficient for harmfulness. Here, represents the predicted harmfulness of the variant site, and Freq represents the frequency of the variant site in the population. It is the minimum value. is the conservation weighting coefficient, Cons is the sequence conservation score of the variant site, and Len is the length of the coding region of the gene containing the variant.

[0041] Specifically, the variant site screening module employs a screening model. During implementation, the harmfulness weight coefficient is set at 1.0-3.0. Based on the pathogenic mechanisms of canine genetic diseases, higher weights are assigned to variant types that affect protein function. The variant site harmfulness prediction value is obtained through professional bioinformatics tools, ranging from 0 to 1. Higher values ​​indicate a stronger harmfulness of the variant site to gene function. The frequency of the variant site in the population is obtained by statistically analyzing a gene database of over 10,000 dogs from different breeds and regions, with a minimum value set at 0.0001 to avoid calculation anomalies where the denominator is 0 when the population frequency is 0. The conservation weight coefficient is set at 0.5-2.0. The sequence conservation score of the variant site is obtained by comparing homologous gene sequences from different species, ranging from 0 to 100. Higher scores indicate that the site is more conserved during evolution, and its variation may have a greater impact on gene function. The length of the coding region of the gene containing the mutation is statistically determined based on the actual number of bases in the coding region, ranging from 100 to 10,000 bases. This model calculates a screening score for mutation sites by integrating multiple dimensions of parameters, including harmfulness prediction value, population frequency, conservation score, and coding region length. Sites with a screening score higher than 5.0 are identified as high-risk pathogenic mutation sites. During implementation, processing 50,000 sites takes no more than 30 seconds, effectively eliminating harmless mutations and false positive sites, significantly improving the accuracy and specificity of mutation site screening, and laying a solid foundation for subsequent genotyping and pathogenic gene association analysis.

[0042] Preferably, the genotyping module employs a genotyping calculation model: ,in, For the final genotype, For the set of all possible genotypes, The number of fractal feature dimensions. For the first The variance of each feature, Genotype In the Observations on each feature Genotype In the The mean of each feature Genotype The prior probability;

[0043] Specifically, the genotyping module employs a genotyping calculation model. During implementation, the number of genotyping feature dimensions is set to 5-10, including key features such as sequencing coverage depth, allele signal intensity, and base quality value. The variance of each feature is obtained by calculating the dispersion of that feature across all sequencing fragments. The value range varies depending on the feature type; the variance of sequencing coverage depth is typically between 10-100, and the variance of signal intensity is typically between 100-1000. The observed values ​​for each genotype on each feature are obtained by extracting the feature data from the corresponding sequencing fragment. The mean is obtained by calculating the average of all observed values ​​for each feature for that genotype. The prior probability of a genotype is determined based on the known genotype frequency distribution in the canine population, ranging from 0 to 1. The sum of the prior probabilities of the three genotypes is 1. This model calculates the probability density product of different genotypes on each feature and combines it with prior probabilities to obtain the comprehensive probability of each genotype. The genotype with the highest comprehensive probability is selected as the final genotyping result. When the difference between the highest and second-highest probabilities is less than 15%, it is marked as genotyping uncertainty and a probability confidence interval is attached. During implementation, the genotyping calculation for each variant site takes no more than 0.5 seconds, and the genotyping accuracy remains stable at over 98%. This ensures both the accuracy of the genotyping results and meets the efficiency requirements of large-scale testing, providing reliable genotyping data for pathogenic gene association analysis.

[0044] Preferably, the pathogenic gene association module employs an association analysis model: ,in, The correlation coefficient between genes and diseases. The number of functional feature dimensions. For the gene to be detected in the first The numerical values ​​of each functional characteristic For known disease-causing genes, in the first... The numerical values ​​of each functional characteristic This is the correlation strength adjustment coefficient. The gene to be detected and the pathogenic gene were compared at the first... The number of matches on each functional feature For the first The total number of features for each functional feature.

[0045] Specifically, the pathogenic gene association module employs an association analysis model. During implementation, the number of functional feature dimensions is set to 8-15, including key features such as sequence homology, gene expression level, and functional domain matching degree. The numerical values ​​of the gene to be detected and known pathogenic genes on each functional feature are calculated using professional analysis tools. The value range is adjusted according to the feature type: sequence homology values ​​range from 0-100, and gene expression level values ​​range from 0-1000. The association strength adjustment coefficient is set to 0.8-1.2 to adjust the sensitivity of the association analysis based on the pathogenic characteristics of different genetic diseases. The number of matches between the gene to be detected and the pathogenic gene on each functional feature is obtained by counting the number of times their values ​​on that feature meet a preset matching threshold. The matching threshold is set according to the importance of the feature: a matching threshold of 80% or higher for key functional features and 60% or higher for minor features. The total number of features for each functional feature is determined according to the specific type of the feature: 10-20 for sequence homology features and 5-10 for functional domain matching degree features. This model calculates the correlation between the gene to be detected and known disease-causing genes across multiple functional characteristics. It then performs a logarithmic operation based on the ratio of the number of matches to the total number of features to obtain the association coefficient between the gene and the disease. The association coefficient ranges from 0 to 1, with values ​​above 0.6 indicating a strong association, 0.4-0.6 indicating a moderate association, and below 0.4 indicating a weak association. During implementation, the association analysis for each gene takes less than one second, and the specificity of the association analysis exceeds 97%. This comprehensive approach uncovers the intrinsic link between genotype and disease, providing a scientific basis for the accurate diagnosis of canine genetic diseases.

[0046] like Figure 2 As shown, a genotype detection algorithm system for canine genetic diseases includes the following steps: First, peripheral blood samples from dogs are obtained and genomic DNA is extracted. The genomic DNA is then sequenced using high-throughput sequencing technology to obtain raw sequencing data. Second, the raw sequencing data is input into a gene sequence feature extraction module, where a feature capture algorithm extracts features such as base arrangement patterns, site mutation characteristics, and sequence length distribution from the gene sequence. Third, the extracted feature data is transmitted to a site polymorphism identification module, where a polymorphism analysis algorithm identifies allelic variation types and variation site distribution patterns at gene sites. Fourth, the identified polymorphic site data is sent to a variation site screening module, where a site screening algorithm screens variation sites with potential pathogenicity. Fifth, the screened variation site data is transmitted to a genotype typing calculation module and a pathogenic gene association module, sequentially completing genotype typing calculation and pathogenic gene association analysis. Sixth, the association analysis results are transmitted to a detection result output module, where a data format normalization algorithm standardizes the results, generating and outputting a detection report including genotype information and pathogenic gene association data.

[0047] like Figure 3 As shown, an electronic device includes a memory and a processor, the memory storing a computer program and the processor being configured to execute the computer program to implement a genotype detection algorithm system for canine genetic diseases.

[0048] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0049] The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0050] A genotyping algorithm system and electronic device for canine genetic diseases have been developed, constructing a six-module collaborative end-to-end technical architecture that achieves standardized and precise processing from raw sequencing data to test results. Through the orderly connection of gene sequence feature extraction, site polymorphism identification, variant site screening, genotyping calculation, pathogenic gene association, and test result output, each module is functionally complementary and data transmission is highly efficient. This ensures the integrity of the testing process while improving the processing accuracy of each step through professional division of labor. Simultaneously, key modules incorporate multi-unit refined design and dedicated analysis models, integrating multi-dimensional data considerations from feature extraction to association analysis, significantly improving the comprehensiveness and reliability of genotyping and meeting the dual needs of large-scale screening and precise diagnosis.

[0051] This invention addresses the insufficient correlation between variant site screening and genotyping. The system integrates both into a coherent process, with pathogenic variant sites output by the screening module directly serving as the core input for genotyping calculations. The screening logic is then optimized by combining probability calculations from the genotyping process, enabling information exchange and mutual verification between the two, significantly improving matching accuracy and reducing errors. Furthermore, addressing the limitation of single-dimensional pathogenic gene association analysis, the pathogenic gene association module integrates multiple functions such as database access, homology analysis, functional pathway matching, and association quantification. It comprehensively considers sequence characteristics, functional pathways, population frequencies, and other multi-dimensional information for association judgment, fully exploring the intrinsic connection between genes and diseases, avoiding the omission of key pathogenic genes, and providing stronger technical support for detection results.

[0052] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0053] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A genotype detection algorithm system for canine genetic diseases, characterized in that, It includes a gene sequence feature extraction module, a site polymorphism identification module, a variant site screening module, a genotyping calculation module, a pathogenic gene association module, and a test result output module; The gene sequence feature extraction module captures features from the raw canine genome sequencing data and transmits them to the locus polymorphism identification module. The locus polymorphism identification module analyzes the gene locus polymorphism features and transmits them to the variant site screening module. The variant site screening module identifies potential pathogenic variant sites and sends them to the genotyping calculation module. The genotyping calculation module performs genotyping calculations on the canine genotypes and sends the results to the pathogenic gene association module. The pathogenic gene association module performs association analysis between the genotyping results and known pathogenic genes and transmits the results to the detection result output module. The detection result output module formats and outputs the detection data after association analysis.

2. The genotype detection algorithm system for canine genetic diseases according to claim 1, characterized in that, The variant site screening module includes: a sequence alignment unit, a site frequency statistics unit, a variant harmfulness prediction unit, and a screening threshold determination unit; The sequence alignment unit performs base-by-base alignment of the canine gene sequence with the reference genome sequence to obtain site difference information. The site frequency statistics unit performs statistical analysis on the frequency of occurrence of the differential sites in the population. The variant harmfulness prediction unit predicts and evaluates the biological functional impact of the differential sites. The screening threshold determination unit sets the screening threshold based on the frequency statistics results and the harmfulness prediction evaluation results and screens out variant sites that meet the threshold requirements.

3. The genotype detection algorithm system for canine genetic diseases according to claim 1, characterized in that, The genotyping calculation module includes: a sequencing data noise reduction unit, an allele identification unit, a genotype probability calculation unit, and a genotyping result determination unit; The sequencing data noise reduction unit removes noise from the input raw sequencing data. The allele identification unit identifies the sequence characteristics of different alleles from the noise-reduced data. The genotype probability calculation unit calculates the probability of different genotype combinations based on the allele sequence characteristics. The typing result determination unit selects the genotype with the highest probability as the final typing result based on the probability calculation results.

4. The genotype detection algorithm system for canine genetic diseases according to claim 1, characterized in that, The pathogenic gene association module includes: a pathogenic gene database retrieval unit, a sequence homology analysis unit, a functional pathway matching unit, and an association degree quantification unit; The pathogenic gene database retrieval unit calls up the preset canine genetic disease pathogenic gene database. The sequence homology analysis unit compares the gene sequence corresponding to the typing result with the pathogenic gene sequence in the database for homology. The functional pathway matching unit matches the gene sequence with high homology with known pathogenic functional pathways. The association degree quantification unit quantifies the matching results to obtain the association degree data between the gene and the disease.

5. The genotype detection algorithm system for canine genetic diseases according to claim 1, characterized in that, The gene sequence feature extraction module employs a feature extraction model: , in, The extracted gene sequence comprehensive feature value, For the first Feature weights of each gene locus For the first Sequencing signal intensity at each site, For the first The reference genomic base values ​​for each locus, where ⊕ represents the base feature XOR operation. For the first Sequencing coverage depth at each site, The regularization coefficient is . This is the weight matrix. For the weight matrix Norm, This represents the total number of sites in the gene sequence.

6. The genotype detection algorithm system for canine genetic diseases according to claim 1, characterized in that, The site polymorphism identification module employs a polymorphism identification model: , in, This is a site polymorphism identification index. For the first Locus heterozygosity markers for each sequencing fragment For the first The quality value of each sequencing fragment, This is the distance attenuation coefficient. For the first The base mismatch distance between each sequencing fragment and the reference sequence The coefficient of variation is 1. For the first Allele frequencies at each locus The average allele frequency across all loci. This represents the total number of sequencing fragments covering this site.

7. The genotype detection algorithm system for canine genetic diseases according to claim 1, characterized in that, The variant site screening module uses a screening model: , in, The score is used to screen for variant sites. This is the weighting coefficient for harmfulness. Here, represents the predicted harmfulness of the variant site, and Freq represents the frequency of the variant site in the population. It is the minimum value. is the conservation weighting coefficient, Cons is the sequence conservation score of the variant site, and Len is the length of the coding region of the gene containing the variant.

8. The genotype detection algorithm system for canine genetic diseases according to claim 1, characterized in that, The genotyping module employs a genotyping calculation model: , in, For the final genotype, For the set of all possible genotypes, The number of fractal feature dimensions. For the first The variance of each feature Genotype In the Observations on each feature Genotype In the The mean of each feature Genotype The prior probability; The pathogenic gene association module employs an association analysis model: , in, The correlation coefficient between genes and diseases. The number of functional feature dimensions. For the gene to be detected in the first The numerical values ​​of each functional characteristic For known disease-causing genes, in the first... The numerical values ​​of each functional characteristic This is the correlation strength adjustment coefficient. The gene to be detected and the pathogenic gene were compared at the first... The number of matches on each functional feature For the first The total number of features for each functional feature.

9. A genotype detection algorithm system for canine genetic diseases according to any one of claims 1-8, characterized in that, The system is running Includes the following steps: The first step is to obtain peripheral blood samples from dogs and extract genomic DNA. The genomic DNA is then sequenced using high-throughput sequencing technology to obtain raw sequencing data. The second step involves inputting the raw sequencing data into the gene sequence feature extraction module, which uses a feature capture algorithm to extract features such as base arrangement patterns, site mutation characteristics, and sequence length distribution from the gene sequence. The third step is to transmit the extracted feature data to the site polymorphism identification module, and use the polymorphism analysis algorithm to identify the allele variation type and the distribution pattern of the variation site of the gene site. The fourth step is to send the identified polymorphic site data to the variant site screening module, and use the site screening algorithm to screen variant sites with potential pathogenicity risk. The fifth step involves sending the selected variant site data to the genotyping calculation module and the pathogenic gene association module to complete the genotyping calculation and pathogenic gene association analysis in sequence. The sixth step involves transmitting the association analysis results to the test result output module, standardizing the results using a data format normalization algorithm, and generating and outputting a test report that includes genotype information and pathogenic gene association data.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program and the processor being configured to execute the computer program to implement a genotype detection algorithm system for canine genetic diseases as described in any one of claims 1 to 8.