Animal individual recognition method based on multi-source feature fusion

CN122531472APending Publication Date: 2026-08-07MIANYANG TEACHERS COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MIANYANG TEACHERS COLLEGE
Filing Date
2026-07-07
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0008]针对现有技术的不足,本发明提供了基于多源特征融合的动物个体识别方法,解决了现有方法依赖固定经验进行同源性判断,难以适配不同物种、测序批次及样本质量条件下单核苷酸多态性一致率分布的系统性偏移的问题

Benefits of technology

[0022] (1) This invention performs unified structured processing on sample registration information, basic identification information, sequencing data and reference genome information, and converts multi-source data into feature expressions that can participate in subsequent calculations, so that the animal individual identification process is transformed from a single detection result comparison to a multi-source data collaborative analysis, thereby realizing the integrated effect of automated processing of animal identification data, feature fusion and relationship determination, effectively solving the problems of animal individual identification relying on a single data source, scattered data processing links and difficulty in forming a unified calculation basis in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531472A_ABST
    Figure CN122531472A_ABST
Patent Text Reader

Abstract

The application discloses an animal individual identification method based on multi-source feature fusion and relates to the technical field of data processing.The animal individual identification method based on multi-source feature fusion comprises the following steps: S1, acquiring sample registration records, basic identification records, original sequencing files and reference genome files and performing pretreatment; S2, constructing a conditional domain label and establishing a multi-source site evidence matrix; S3, calculating a quality-weighted consistency rate and a quality-weighted mismatch rate, and performing missing correction and conditional domain layer calibration; S4, dividing relationship distribution clusters; performing homology determination; S5, constructing a sample similarity graph with symbols, performing clustering merging, and generating a same individual candidate sample group; and S6, performing evidence discount fusion and generating a candidate sample group confidence degree.The method solves the problem that existing methods rely on fixed experience for homology determination and are difficult to adapt to systematic deviation of single nucleotide polymorphism consistency rate distribution under different species, sequencing batches and sample quality conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to an animal individual identification method based on multi-source feature fusion. Background Technology

[0002] With the increasing demand for ecological environmental protection, wildlife resource supervision, and identification of biological samples involved in cases, animal species identification and individual origin determination have significant application value in fields such as wildlife protection, forensic identification, population monitoring, and biodiversity research. In real-world scenarios, submitted samples may originate from animal tissues, remains, hair, blood, bones, or other biological materials. The morphological integrity, preservation conditions, and identifiable characteristics of these samples vary, making accurate identification often difficult to achieve solely based on external morphological features. With the development of molecular biology and high-throughput sequencing technologies, animal identification technology based on genetic markers has gradually become an important tool. Related methods can analyze the species origin, phylogenetic relationships, and individual affiliation of samples using genetic information such as mitochondrial sequences, nuclear genome variation sites, microsatellite markers, or single nucleotide polymorphism sites. Existing bioinformatics analysis workflows typically combine sequencing data quality control, sequence alignment, variation site detection, genetic difference statistics, and sample relationship determination, providing a data foundation for animal individual identification and quantity assessment under multi-sample conditions. As sample sources diversify and the scale of testing data increases, animal individual identification technology is gradually developing towards multi-source information integration, automated analysis, and data-driven determination.

[0003] For example, the invention patent with publication number CN116246705B discloses a method and apparatus for analyzing whole-genome sequencing data, including: acquiring whole-genome sequencing sequence data; determining a fusion feature dataset and a feature dictionary based on the sequencing sequence data; inputting the fusion feature dataset into a pre-established prediction model to obtain phenotypic prediction data corresponding to the sequencing sequence data; determining weight data corresponding to the prediction model based on the fusion feature dataset and the prediction model; filtering and sorting the fusion data in the fusion feature dataset based on the weight data, pre-set goodness of fit and accuracy thresholds to obtain a fusion feature evaluation ranking list; and determining the target gene sequence or target non-coding functional region sequence in the whole-genome sequencing sequence data based on the fusion feature evaluation ranking list and the feature dictionary, thereby improving the accuracy of detection results, the efficiency of computing resource utilization, the scalability of applications, and the completeness of the analysis process.

[0004] For example, the invention patent with publication number CN114974414B discloses a method, apparatus, device, and storage medium for constructing population phylogenetic trees. The method includes: obtaining a chromosome map file of a population sample, wherein the population sample includes multiple individual samples, and the chromosome map file includes SNP locus information of the individual samples; using the linkage disequilibrium principle, calculating the fragment ratios of various target chromosome segments among the individual samples based on the chromosome map file, wherein the number and type of common ancestral segments contained in each target chromosome segment are different; using the similarity between SNP loci, calculating the genomic kinship correlation coefficients between individual samples based on the chromosome map file; determining the kinship between individual samples based on the fragment ratios and genomic kinship correlation coefficients; and constructing a population phylogenetic tree of the population sample based on the genomic kinship correlation coefficients and kinship, thereby realizing the reconstruction of population phylogenetic trees across multiple generations and varieties, filling the gap in methods for real-world large-scale population phylogenetic tree reconstruction.

[0005] However, existing analytical methods based on whole-genome sequencing data or single nucleotide polymorphism (SNP) sites primarily focus on feature extraction, phenotypic prediction, phylogenetic calculation, or population phylogenetic construction. When used for animal individual identification, they typically rely on statistical results such as inter-sample locus consistency rate, mismatch rate, or phylogenetic correlation coefficient. Due to differences in genetic background, reference genome fit, sequencing batches, sample preservation status, and the number of valid loci across species, the distribution of SNP consistency rates obtained under the same judgment scale is prone to shift. Directly using fixed empirical thresholds or a single similarity index for homology judgment fails to adequately reflect the impact of sample quality, locus deletions, reference alignment bias, and batch differences on identification results. This can easily lead to the incorrect splitting of low-quality samples or the incorrect merging of individuals with similar genetic backgrounds, thus affecting the accuracy of animal individual count identification and sample attribution results.

[0006] Therefore, in order to address the above problems, there is an urgent need for an animal individual identification method based on multi-source feature fusion. Summary of the Invention

[0007] Technical problems to be solved

[0008] To address the shortcomings of existing technologies, this invention provides an animal individual identification method based on multi-source feature fusion, which solves the problem that existing methods rely on fixed experience for homology judgment and are difficult to adapt to the systematic shift in the single nucleotide polymorphism consistency rate distribution under different species, sequencing batches and sample quality conditions.

[0009] Technical solution

[0010] To achieve the above objectives, this invention provides the following technical solution: an animal individual identification method based on multi-source feature fusion, comprising: S1, acquiring and preprocessing sample registration records, basic identification records, raw sequencing files, and reference genome files within the same identification batch; S2, constructing conditional domain labels based on the basic identification records, sample registration records, and reference genome files, establishing a multi-source locus evidence matrix, and generating conditional domain offsets and distribution migrations; S3, calculating the quality-weighted consistency rate and quality-weighted mismatch rate based on the conditional domain offsets and distribution migrations, and... S4. Perform missing correction and conditional domain hierarchical calibration to obtain conditional scale calibration consistency rate, conditional scale calibration mismatch rate, and scale drift compensation amount; S5. Based on the conditional scale calibration consistency rate, conditional scale calibration mismatch rate, scale drift compensation amount, and distribution migration amount, divide the relationship distribution clusters through dimensionality reduction clustering; generate sample pair homology support results and perform homology determination; S6. Construct a signed sample similarity map, perform clustering and merging to generate candidate sample groups for the same individual; S7. Perform evidence discount fusion to generate the confidence of candidate sample groups and output the number of animal individuals and individual attribution labels.

[0011] Further, the specific process of obtaining and preprocessing the sample registration records, basic identification records, raw sequencing files, and reference genome files within the same identification batch is as follows: Obtain the sample registration records, basic identification records, raw sequencing files, and reference genome files for multiple animal samples to be identified within the same identification batch; the sample registration records include sample number, evidence number, tissue type, sampling site, storage time record, sequencing batch number, and reference genome file identifier; the basic identification records include morphological observation records and mitochondrial species confirmation records; the raw sequencing files include read numbers, read base sequences, and read quality characters; the reference genome file includes chromosome numbers, reference coordinates, and reference base sequences, where the reference coordinates are the positional identifiers in the reference genome coordinate system; the raw sequencing files undergo quality character parsing, adapter removal, and quality... After quantitative shearing, the base sequence of the reads is aligned to the reference coordinate system of the reference genome file to obtain the alignment result file. Then, duplicate read removal and variant site detection are performed to obtain a candidate site evidence library and determine the site reference coordinates. Genotype interpretation results, genotype likelihood values, site sequencing depth values, site quality values, and genotype quality values ​​are obtained from the effective coverage read statistics and genotype likelihood estimation of candidate variant sites. The alignment ratio, average sequencing depth, coverage balance, and soft shearing ratio are statistically analyzed using alignment read markers, reference coordinate coverage counts, and soft shearing operation markers in the alignment result file to form a set of sample alignment statistical features. The average sequencing depth and storage duration are processed using the Z-score normalization algorithm. The site quality value, genotype quality value, alignment ratio, coverage balance, and soft shearing ratio are processed using the min-max normalization algorithm.

[0012] Furthermore, the specific process of constructing a multi-source locus evidence matrix based on conditional domain labels using basic identification records, sample registration records, and reference genome files is as follows: Hierarchical combination coding is performed based on mitochondrial species confirmation records, reference genome file identifiers, sequencing batch numbers, and tissue types. Morphological observation records are used as species anchoring auxiliary information, and sampling sites are used as tissue type supplementary verification information. The hierarchical combination coding results are hashed to generate conditional domain labels. Under the same mitochondrial species confirmation results, animal samples to be identified are paired to obtain candidate sample pairs. A multi-source locus evidence matrix is ​​constructed based on the candidate locus evidence library. The multi-source locus evidence matrix is ​​indexed and matched according to sample number, chromosome number, and locus reference coordinates. The corresponding genotype interpretation results, genotype likelihood values, locus sequencing depth values, locus quality values, and genotype quality values ​​are read from the candidate locus evidence library and filled into the matrix cells.

[0013] Further, the specific process for generating conditional domain offsets and distribution migrations is as follows: The sample alignment statistical feature set and standardized storage duration records are spliced ​​and whitened by principal component analysis to obtain a sample quality conditional vector; the reference genome file identifier, alignment ratio, coverage balance, and soft shearing ratio are hashed and whitened by principal component analysis to obtain a reference adaptation feature vector; the evidence matrix of multi-source sites under the same conditional domain label is processed using the conditional quantile calibration method to obtain the conditional domain site interpretation boundary; sites with genotype interpretation results are screened based on the conditional domain site interpretation boundary to obtain a set of interpretable sites for the samples; candidate sample pairs are matched according to chromosome number and site reference coordinates, and both animal samples to be identified belong to the same category. Paired sites in the readable site set are categorized into the common single nucleotide polymorphism (SNP) site set; paired sites with the same genotype interpretation result are accumulated as the number of identical sites, paired sites with different genotype interpretation results are accumulated as the number of mismatch sites, and the total number of sites in the common SNP site set is accumulated as the number of common SNP sites; a minimum hash sketch is constructed for the readable site set of the sample to obtain the site overlap prior value; the conditional domain offset is obtained by statistically processing the sample quality conditional vector and the reference adaptation feature vector through the conditional vector mean difference and quantile difference; the site sequencing depth value distribution, site quality value distribution, and genotype likelihood value distribution are processed through the optimal transfer distance to obtain the depth distribution migration, quality distribution migration, and likelihood distribution migration.

[0014] Furthermore, the specific process for calculating the quality-weighted consistency rate and quality-weighted mismatch rate based on conditional domain offset and distribution migration is as follows: Empirical distribution function quantile mapping is performed on the sequencing depth, quality, and genotype likelihood values ​​of the common single nucleotide polymorphism (SNP) site set to obtain depth quantiles, quality quantiles, and likelihood quantiles; the depth quantiles, quality quantiles, likelihood quantiles, conditional domain offset, depth distribution migration, quality distribution migration, and likelihood distribution migration are fused using a non-additive evidence integration method to obtain the site information contribution value; the summation of the site information contribution values ​​for sites with the same genotype interpretation results is divided by the summation of all site information contribution values ​​to obtain the quality-weighted consistency rate; the summation of the site information contribution values ​​for sites with different genotype interpretation results is divided by the summation of all site information contribution values ​​to obtain the quality-weighted mismatch rate; the distribution of the number of common SNP sites, the prior value of site overlap, and the number of common SNP sites for all candidate sample pairs under the same conditional domain label are input into the beta binomial empirical contraction rule to obtain the site cardinality confidence coefficient.

[0015] Further, the specific process of performing deletion correction and conditional domain hierarchical calibration to obtain the conditional scale calibration consistency rate, conditional scale calibration mismatch rate, and scale drift compensation is as follows: The union of the reference coordinates of all animal samples to be identified under the same conditional domain label in the candidate site evidence library is taken as the complete set of candidate sites for the conditional domain; the number of sites in the complete set of candidate sites for the conditional domain where only one animal sample has a genotype interpretation result is accumulated as the number of unilateral deletion sites, and the number of sites where neither animal sample has a genotype interpretation result is accumulated as the number of bilateral deletion sites; the sum of the number of unilateral deletion sites and the number of bilateral deletion sites is divided by the number of sites in the complete set of candidate sites for the conditional domain to obtain the deletion ratio; the number of unilateral deletion sites... The quantity and the number of bilaterally missing sites are proportionalized according to the number of sites in the full set of candidate sites for the conditional domain. The proportionalized result is synthesized with the missing ratio, and the synthesized result is corrected according to the conditional domain offset, depth distribution migration, and quality distribution migration to obtain the missing perturbation coefficient. The credibility is recalibrated according to the quality-weighted consistency rate, quality-weighted mismatch rate, site cardinality confidence coefficient, and missing perturbation coefficient to obtain the missing correction consistency rate and missing correction mismatch rate. The conditional domain is hierarchically calibrated according to the missing correction consistency rate, missing correction mismatch rate, site cardinality confidence coefficient, conditional domain offset, and conditional domain label to obtain the conditional scale calibration consistency rate, conditional scale calibration mismatch rate, and scale drift compensation amount.

[0016] Furthermore, based on the conditional scale calibration consistency rate, conditional scale calibration mismatch rate, scale drift compensation, and distribution migration, the specific process of dividing the relationship distribution clusters through dimensionality reduction clustering is as follows: The result after subtracting the conditional scale calibration consistency rate, the conditional scale calibration mismatch rate, the scale drift compensation, the missing proportion, the depth distribution migration, the quality distribution migration, and the likelihood distribution migration are written into the sample pair difference matrix; the sample pair difference matrix is ​​processed using principal component analysis to obtain the sample pair difference features; and the relationship distribution is divided according to the difference features of all sample pairs under the same conditional domain label using the GMM clustering method to obtain homologous relationship distribution clusters, separated relationship distribution clusters, and transitional relationship distribution clusters.

[0017] Furthermore, the specific process of generating sample pair homology support results and determining homology is as follows: using the Bayesian evidence fusion method, with the posterior probability of the homology relationship distribution cluster, the posterior probability of the separation relationship distribution cluster, the conditional scale calibration consistency rate, the conditional scale calibration mismatch rate, the site cardinality confidence coefficient, and the scale drift compensation amount as inputs, the sample pair homology support value is obtained. Based on the distribution characteristics of the posterior probabilities of the homology relationship distribution cluster and the separation relationship distribution cluster, the homology determination threshold is determined. When the sample pair homology support value is greater than or equal to the homology determination threshold, a homology relationship edge is generated; when the sample pair homology support value is less than the homology determination threshold and belongs to the separation relationship distribution cluster, a separation relationship edge is generated; when the sample pair belongs to the transition relationship distribution cluster, a relationship label to be verified is generated.

[0018] Furthermore, the specific process of constructing a signed sample similarity graph, performing clustering and merging to generate candidate sample groups for the same individual is as follows: using sample numbers as graph nodes, using homology edges and separation edges as signed graph edges, and using the homology support value of sample pairs as graph edge weights, a signed sample similarity graph is constructed; the graph node representations in the signed sample similarity graph are updated through signed graph attention to obtain sample node embedding vectors; the initial sample merging results are clustered and merged through intra-group consistency statistics and inter-group separation statistics to obtain candidate sample groups for the same individual.

[0019] Further, the specific process of performing evidence discount fusion to generate candidate sample group confidence scores and output the number of animal individuals and their attribution labels is as follows: Traverse the sample pairs within the same candidate sample group, counting the number of separating edges within the group and the sum of their corresponding graph edge weights; traverse the sample pairs between different candidate sample groups of the same individual, counting the number of homologous edges across groups and the sum of their corresponding graph edge weights; based on the number of separating edges within the group, the sum of their weights, the number of homologous edges across groups, and the sum of their weights, obtain the merge conflict value; and employ evidence discount fusion based on conflict factors. Then, the homology support value, locus cardinality confidence coefficient, scale drift compensation, and merging conflict value of the sample pairs are fused to obtain the confidence of the candidate sample group; the number of candidate sample groups of the same individual is taken as the number of animal individuals, the group identifier of the candidate sample group of the same individual is taken as the individual attribution label corresponding to the sample number within the group, the sample number within the group is associated with the physical evidence number, and the sample number in the sample pair carrying the relationship label to be verified is taken as the sample number to be verified; the output is the number of animal individuals, the individual attribution label corresponding to the sample number and the associated physical evidence number, the confidence of the candidate sample group, and the sample number to be verified.

[0020] Beneficial effects

[0021] The present invention has the following beneficial effects:

[0022] (1) This invention performs unified structured processing on sample registration information, basic identification information, sequencing data and reference genome information, and converts multi-source data into feature expressions that can participate in subsequent calculations, so that the animal individual identification process is transformed from a single detection result comparison to a multi-source data collaborative analysis, thereby realizing the integrated effect of automated processing of animal identification data, feature fusion and relationship determination, effectively solving the problems of animal individual identification relying on a single data source, scattered data processing links and difficulty in forming a unified calculation basis in the prior art.

[0023] (2) This invention constructs conditional domain labels to incorporate different species origins, sequencing batches, tissue types and reference genome adaptation into a unified grouping basis, so that samples in the same identification batch can be compared under similar detection conditions, thereby achieving the effect of adaptive adjustment of the sample relationship judgment scale as the detection conditions change, effectively solving the problem that the fixed empirical scale in the prior art is difficult to adapt to different species, different batches and different sample quality conditions.

[0024] (3) In this invention, by performing quality weighting on single nucleotide polymorphism site evidence and correcting it in combination with site deletion, site evidence of different depths, different qualities and different deletion states can participate in homology judgment in a differentiated manner, thereby achieving the effect of recalibrating the credibility of site evidence, and effectively solving the problem that low-quality sites, deletion sites and abnormal sites in the prior art can easily interfere with the consistency rate judgment.

[0025] (4) This invention, by scaling and compensating for the consistency rate and mismatch rate under different condition domains, enables the sample difference results formed under cross-species, cross-batch and cross-reference genome conditions to enter a unified comparison space, thereby achieving the effect of cross-condition domain sample relationship comparability, effectively solving the problem of unstable judgment results caused by the distribution shift of single nucleotide polymorphism consistency rate under different detection conditions in the prior art.

[0026] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0027] Figure 1 This is a flowchart of the animal individual identification method based on multi-source feature fusion of the present invention;

[0028] Figure 2 This is a comparison chart of the changes in the reference genome alignment ratio and consistency rate in this invention;

[0029] Figure 3 This is a stacked decomposition diagram of the quality-weighted consistency rate and missing perturbation coefficient of the candidate samples in this invention;

[0030] Figure 4 This is a flowchart of the clustering and merging process of the symbolic sample similarity graph of the present invention. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] Please see Figures 1-4This invention provides a technical solution: an animal individual identification method based on multi-source feature fusion, comprising the following steps: S1, acquiring sample registration records, basic identification records, original sequencing files, and reference genome files within the same identification batch and performing preprocessing; S2, constructing conditional domain labels based on the basic identification records, sample registration records, and reference genome files, establishing a multi-source locus evidence matrix, and generating conditional domain offsets and distribution migrations; S3, calculating quality-weighted consistency rate and quality-weighted mismatch rate based on conditional domain offsets and distribution migrations, and performing missing correction and conditional domain hierarchical calibration to obtain conditional scale calibration consistency rate, conditional scale calibration mismatch rate, and scale drift compensation; S4, dividing relationship distribution clusters through dimensionality reduction clustering based on conditional scale calibration consistency rate, conditional scale calibration mismatch rate, scale drift compensation, and distribution migration; generating sample pair homology support results and performing homology determination; S5, constructing a signed sample similarity graph, performing clustering and merging to generate candidate sample groups for the same individual; S6, performing evidence discount fusion to generate the confidence score of the candidate sample group and outputting the number of animal individuals and individual attribution markers.

[0033] Specifically, the process of obtaining and preprocessing the sample registration records, basic identification records, raw sequencing files, and reference genome files within the same identification batch is as follows: The sample registration records, basic identification records, raw sequencing files, and reference genome files of multiple animal samples to be identified within the same identification batch are obtained. When a fixed species determination standard is lacking for the animal samples to be identified, an initial association is established between the physical evidence number in the sample registration record and the morphological observation record. The sample registration record includes the sample number, physical evidence number, tissue type, sampling site, storage duration record, sequencing batch number, and reference genome file identifier. The storage duration record is the time interval between the completion of sampling and the completion of sequencing detection for the animal sample to be identified, and a uniform time unit is used within the same identification batch. The basic identification record includes the morphological observation record and the mitochondrial species confirmation record. The raw sequencing file includes the read number, read base sequence, and read quality characters. The reference genome file includes the chromosome number, reference coordinates, and reference base sequence, where the reference coordinates are the positional identifiers in the reference genome coordinate system.

[0034] After quality character parsing, adapter removal, and quality splicing of the original sequencing file, the quality splicing process, based on the base error probability obtained from read quality character conversion, cuts off consecutive low-quality bases at the ends of reads, retaining the read sequence that still meets the read length requirements after splicing. The read sequence is then aligned to the reference coordinate system of the reference genome file to obtain the alignment result file. Quality character parsing uses ASCII code offset conversion to obtain the base error probability. Adapter removal uses the FASTP tool for automatic detection and trimming, followed by duplicate read removal and variant site detection to obtain a candidate site evidence library and determine the site reference coordinates. Variant site detection is based on effective coverage read statistics, site quality values, genotype quality values, and genotype likelihood values. Variant sites that meet the candidate writing criteria are added to the candidate site evidence library. Duplicate read removal uses the MarkDuplicates strategy to remove redundant alignment records introduced by PCR amplification. The MarkDuplicates strategy is a method that identifies and marks sets of duplicate reads originating from the same DNA template based on alignment coordinates and insert fragment length. The effective coverage read statistics and genotype likelihood values ​​of the candidate variant sites are used to determine the candidate site evidence library. Genotype likelihood estimation yields genotype interpretation results, genotype likelihood values, sequencing depth values, site quality values, and genotype quality values. Genotype likelihood estimation utilizes the haplotype call algorithm from the Genome Analysis Toolkit, combined with base quality recalibration. Alignment read markers, reference coordinate coverage counts, and soft shearing operation markers in the alignment results file are used to statistically analyze alignment proportions, average sequencing depth, coverage uniformity, and soft shearing proportions, forming a set of statistical features for sample alignment. Coverage uniformity is quantified by calculating the coefficient of variation of coverage depth, and the soft shearing proportion reflects the sequence differences between the reference genome and the sample species. The average sequencing depth and storage time records are processed using the Z-score normalization algorithm. Z-score normalization transforms each feature into a distribution with a mean of 0 and a standard deviation of 1, eliminating dimensional differences. The locus quality value, genotype quality value, alignment ratio, coverage balance, and soft shearing ratio are processed using the min-max normalization algorithm. The minimum and maximum values ​​are derived from the statistical boundaries of the corresponding features within the same identification batch. Feature values ​​exceeding the statistical boundaries are truncated according to the boundary values ​​and then normalized. Min-max normalization linearly maps all feature values ​​to the [0, 1] interval, ensuring a balance of contributions from each feature during subsequent fusion.

[0035] This implementation plan establishes a unified data foundation for animal samples to be identified within the same identification batch, and transforms the original sequencing files into structured data with quality control, reference coordinate constraints, variant site evidence, and sample alignment quality descriptions. This reduces the impact of sample source confusion, species prior bias, low-quality reads, duplicate reads, reference fit differences, and anomalous eigenvalues ​​on subsequent analyses. It ensures that the candidate site evidence library, site reference coordinates, genotype interpretation results, sample alignment statistical feature set, and normalized quality features are traceable and comparable, providing stable data support for conditional domain label construction, multi-source site evidence matrix establishment, common single nucleotide polymorphism site set screening, and animal individual identification result output.

[0036] Specifically, the process of constructing a multi-source locus evidence matrix based on conditional domain labels using basic identification records, sample registration records, and reference genome files is as follows: Hierarchical combination coding is performed based on mitochondrial species confirmation records, reference genome file identifiers, sequencing batch numbers, and tissue types. This hierarchical combination coding is performed in the order of mitochondrial species confirmation records, reference genome file identifiers, sequencing batch numbers, and tissue types. By sequentially concatenating multi-layered metadata into a single combination key according to biological meaning, fine-grained conditional grouping of samples across batches is achieved. Morphological observation records serve as auxiliary information for species anchoring, and sampling sites serve as supplementary verification information for tissue types. Morphological observation records are used to verify the consistency of species attribution corresponding to mitochondrial species confirmation records, while sampling sites are used to verify the consistency of tissue types. Animal samples with inconsistent verification are not directly incorporated into the same conditional domain label. Morphological observation records provide species prior constraints independent of molecular markers, while sampling sites perform cross-validation of tissue type labels, reducing the introduction of tissue mislabeling. The conditional domain segmentation error is addressed by providing morphological evidence when there are geographical subspecies differences between the reference genome and the target species. The hierarchical combination coding results are hashed to generate conditional domain labels, and the hierarchical combination coding results are retained as a basis for reverse lookup of conditional domain labels. When the hash mapping results are the same but the hierarchical combination coding results are different, conditional domain labels are regenerated. The hash mapping converts the variable-length hierarchical combination coding into fixed-length identifiers, ensuring the uniqueness and comparability of conditional domain labels between different batches. Under the same mitochondrial species confirmation result, the animal samples to be identified are paired to obtain candidate sample pairs. A multi-source locus evidence matrix is ​​constructed based on the candidate locus evidence library. The multi-source locus evidence matrix is ​​indexed and matched according to sample number, chromosome number, and locus reference coordinates. The three-key index structure enables the matrix to quickly locate and extract corresponding evidence units from three dimensions: sample, chromosome, and locus. The corresponding genotype interpretation results, genotype likelihood values, locus sequencing depth values, locus quality values, and genotype quality values ​​are read from the candidate locus evidence library and filled into the matrix units.

[0037] In this implementation plan, the original sequencing files are converted into a structured data foundation with quality control, reference coordinate constraints, site evidence support, and sample quality description. Low-quality reads, adapter residues, duplicate reads, and abnormal feature values ​​are eliminated to prevent interference with subsequent analysis. This ensures that the candidate site evidence library, site reference coordinates, genotype interpretation results, and sample comparison statistical feature set have high credibility and comparability, providing stable data support for screening common single nucleotide polymorphism site sets, calculating quality-weighted consistency rates, correcting deletion perturbation coefficients, and outputting animal individual identification results.

[0038] Specifically, the process for generating conditional domain offsets and distribution migrations is as follows: The sample alignment statistical feature set and standardized storage duration records are spliced ​​and subjected to principal component whitening. Principal component whitening calculates the covariance matrix based on the spliced ​​features corresponding to the animal samples to be identified within the same identification batch, and retains principal components according to their cumulative contribution. The splicing operation merges alignment statistical features with sample storage degradation factors, enabling the quality vector to simultaneously capture sequencing process bias and exogenous noise caused by sample storage, thus obtaining the sample quality condition vector. The reference genome file identifier, alignment ratio, coverage balance, and soft shearing ratio are feature-hash encoded and subjected to principal component whitening. The principal component whitening corresponding to the reference adaptation feature vector is calculated based on the reference genome file identifier and alignment statistics within the same identification batch. Feature hashing mapping maps the categorical variable of the reference genome file identifier and the continuous alignment statistics to a high-dimensional sparse space, preserving the interaction information between features and obtaining a reference-fit feature vector. The reference-fit feature vector quantifies the species closeness between the reference genome and the sample population to be tested. The conditional quantile calibration method processes the evidence matrix of multi-source sites under the same conditional domain label to obtain the conditional domain site interpretation boundary. The conditional quantile calibration method statistically analyzes the quantile distribution of sequencing depth, site quality, and genotype quality values ​​of sites under the same conditional domain label, and generates the conditional domain site interpretation boundary accordingly. Conditional quantile calibration replaces the fixed threshold with a quantile threshold, so that the interpretation boundary is adaptively adjusted according to the data quality distribution within each conditional domain. Based on the conditional domain site interpretation boundary, sites with genotype interpretation results are screened to obtain the set of interpretable sites of the sample.

[0039] Candidate sample pairs are paired based on chromosome number and locus reference coordinates. Coordinate pairing ensures that the two samples are compared at the same physical location in the genome, eliminating false mismatches caused by coordinate offset. Paired loci belonging to the identifiable locus set of both animal samples are included in the common single nucleotide polymorphism (SNP) locus set, i.e., the common valid SNP locus set. Paired loci with the same genotype interpretation result are accumulated as the number of identical loci, and paired loci with different genotype interpretation results are accumulated as the number of mismatch loci. The total number of loci in the common SNP locus set is accumulated as the common SNP locus number. A minimum hash draft is constructed for the identifiable locus set to obtain the locus overlap prior value. The minimum hash draft uses the same set of hash functions to process the identifiable locus sets on both sides of the candidate sample pair, and the locus overlap prior value is determined based on the consistency ratio of hash values ​​at the same position in the two drafts. The minimum hash draft is then processed through random permutation and minimum hash function. Small hash value extraction enables fast and unbiased estimation of site set overlap without storing the complete site set. By statistically processing the sample quality condition vector and reference adaptation feature vector using condition vector mean and quantile differences, condition domain offset is obtained. This offset is determined by the mean and quantile differences of the current candidate sample's corresponding condition vector relative to the overall condition vector under the same condition domain label, characterizing the degree of data deviation across condition domains. The optimal transfer distance is used to process the site sequencing depth distribution, site quality distribution, and genotype likelihood distribution, yielding depth migration, quality migration, and likelihood migration. The input distribution for the optimal transfer distance is discretized from the corresponding values ​​under the same condition domain label using a unified binning method. The transfer cost is constructed based on binning position differences. The optimal transfer distance measures the minimum cost required to transfer quality from one probability distribution to another, geometrically capturing the overall deformation of the distribution.

[0040] In this implementation plan, sample quality status, reference adaptation status, and site interpretation status are uniformly incorporated into the conditional domain analysis framework, forming a sample quality conditional vector, a reference adaptation feature vector, a conditional domain site interpretation boundary, and a set of interpretable sites in the sample. This improves the consistency of site screening criteria across different sequencing batches, different levels of adaptation to the reference genome, and different sample preservation states. Simultaneously, by generating coordinate pairing, site overlap prior values, conditional domain offsets, and distribution migrations, the comparability and reliability of common single nucleotide polymorphism (SNP) site sets are enhanced, providing a stable conditional domain evidence basis for subsequent quality-weighted consistency rate calculations, deletion perturbation coefficient corrections, conditional scale calibration, and homology determination.

[0041] Specifically, the process of calculating the quality-weighted consistency rate and quality-weighted mismatch rate based on conditional domain offset and distribution migration is as follows: Empirical distribution function quantile mapping is performed on the sequencing depth, quality, and genotype likelihood values ​​of sites in the common single nucleotide polymorphism (SNP) site set to obtain depth quantiles, quality quantiles, and likelihood quantiles. This empirical distribution function quantile mapping uniformly transforms each original metric to a quantile scale, eliminating dimensional differences and making values ​​from different distribution sources directly comparable. The depth quantiles, quality quantiles, likelihood quantiles, conditional domain offset, depth distribution migration, quality distribution migration, and likelihood distribution migration are fused using a non-additive evidence integration method to obtain the site information contribution value. This site information contribution value measures the evidence weight of a single SNP site in homology determination. The non-additive integration considers the nonlinear interaction between evidence sources during the fusion process, avoiding the duplication of information in the additive model when evidence is redundant.

[0042] The quality-weighted consistency rate is obtained by summing the contribution values ​​of loci with the same genotype interpretation and dividing by the sum of all loci's contribution values. The quality-weighted mismatch rate is obtained by summing the contribution values ​​of loci with different genotype interpretations and dividing by the sum of all loci's contribution values. Compared with the ordinary consistency rate, the quality-weighted method reduces the interference of low-depth and low-quality loci on the judgment results. The number of common single nucleotide polymorphism (SNP) sites, the prior value of site overlap, and the distribution of the number of common SNP sites of all candidate sample pairs under the same condition domain label are input into the beta binomial empirical contraction rule to obtain the site cardinality confidence coefficient. The beta binomial empirical contraction uses the distribution of the number of common SNP sites of sample pairs under the same condition domain as an empirical prior to perform contraction estimation towards the population mean for sample pairs with insufficient site cardinality, thus suppressing the variance expansion caused by small cardinality.

[0043] In this implementation scheme, the differences in site quality, sequencing depth, genotype interpretation confidence, and conditional domain distribution offset in the common single nucleotide polymorphism site set are uniformly converted into site information contribution values. This allows sites with higher quality, stable interpretation, and smaller conditional domain offsets to receive higher evidence weight in homology determination, reducing the interference of low-quality sites, offset sites, and insufficient site base on consistency judgment. At the same time, by generating quality-weighted consistency rate, quality-weighted mismatch rate, and site base confidence coefficient, the homology evidence of candidate sample pairs no longer depends on ordinary consistency rate or fixed empirical scale, providing a more reliable site-level confidence basis for subsequent calculation of missing correction consistency rate, conditional scale calibration consistency rate, and homology support value of sample pairs.

[0044] Specifically, the process of performing missing data correction and conditional domain hierarchical calibration to obtain the conditional scale calibration consistency rate, conditional scale calibration mismatch rate, and scale drift compensation is as follows: The union of the reference coordinates of all animal samples to be identified under the same conditional domain label in the candidate site evidence library is taken as the complete set of candidate sites for the conditional domain; for the current candidate sample pair, the number of sites in the complete set of candidate sites where only one animal sample to be identified has a genotype interpretation result is accumulated as the number of unilaterally deleted sites, and the number of sites in the current candidate sample pair where neither animal sample to be identified has a genotype interpretation result is accumulated as the number of bilaterally deleted sites; the sum of the number of unilaterally deleted sites and the number of bilaterally deleted sites is divided by the number of sites in the complete set of candidate sites for the conditional domain to obtain the missing ratio; the number of unilaterally deleted sites and the number of bilaterally deleted sites are proportionalized according to the number of sites in the complete set of candidate sites for the conditional domain, respectively, and the proportionalized result is synthesized with the missing ratio, and the synthesized result is corrected according to the conditional domain offset, depth distribution migration, and quality distribution migration to obtain the missing perturbation coefficient.

[0045] Reliability recalibration is performed based on quality-weighted consistency rate, quality-weighted mismatch rate, locus cardinality confidence coefficient, and missing perturbation coefficient to obtain missing correction consistency rate and missing correction mismatch rate. Conditional domain hierarchical calibration is then performed based on missing correction consistency rate, missing correction mismatch rate, locus cardinality confidence coefficient, conditional domain offset, and conditional domain label to obtain conditional scale calibration consistency rate, conditional scale calibration mismatch rate, and scale drift compensation. Specifically, the missing correction consistency rate is obtained by multiplying the quality-weighted consistency rate, the locus cardinality confidence coefficient, and the result after subtracting the missing perturbation coefficient. The missing match rate is obtained by summing the match rate and the missing perturbation coefficient, and then multiplying the result by subtracting the site base confidence coefficient. Conditional domain hierarchical calibration is performed by scaling the missing match rate and missing mismatch rate using the consistency rate distribution and mismatch rate distribution under the same conditional domain label. The offset compensation is recorded according to the conditional domain offset. The scale drift compensation is determined by the change in missing match rate and missing mismatch rate before and after conditional domain hierarchical calibration. The scale drift compensation records the offset amplitude during the conditional domain hierarchical calibration process for retrospective comparison in subsequent cross-domain comparisons.

[0046] In this embodiment, individual identification was performed using multiple suspected Asiatic black bear (Ursus thibetanus) tissue samples. After sequencing, samples BZ-2, BZ-3, and BZ-4 were assigned to the same reference genome adaptation batch via conditional domain tags, with initial quality-weighted consistency rates of 84.76%, 85.08%, and 84.64%, respectively. Since the reference genome alignment ratio of sample BZ-7 was only 18.23%, the reference adaptation feature vector deviated significantly, and the conditional domain offset was dynamically increased. After deletion perturbation coefficient correction, the deletion correction consistency rates of BZ-2, BZ-3, and BZ-4 all exceeded 0.95, while the correction consistency rate of BZ-7 with other samples was below 0.55. Finally, BZ-2, BZ-3, and BZ-4 were merged into the same individual, consistent with subsequent mitochondrial genome and morphological verification results, demonstrating that this method can still stably output individual assignment results even in the absence of highly closely related reference genomes.

[0047] like Figure 2 The figure shows a comparison of the alignment ratio and consistency rate of the reference genome. The horizontal axis represents the alignment ratio of the reference genome, expressed as a percentage, indicating the proportion of sequencing reads from each sample that have aligned to the reference genome. The vertical axis represents the consistency rate, showing the consistency changes for different samples. The solid dotted lines in the figure represent the deletion correction consistency rate, and the dashed square lines represent the quality-weighted consistency rate. The BZ-2, BZ-3, BZ-4, and BZ-7 labels in the figure are the corresponding sample numbers. As shown in the figure, the reference genome alignment rates for BZ-2, BZ-3, and BZ-4 were 89.5%, 91.2%, and 88.7%, respectively, with corresponding quality-weighted consistency rates of 0.848, 0.850, and 0.846. After deletion correction, the consistency rates increased to 0.957, 0.958, and 0.955, respectively, indicating that these samples achieved relatively stable consistency results after deletion correction, provided that the reference genome fit was good. BZ-7 had a reference genome alignment rate of 18.23%, with a corresponding quality-weighted consistency rate of 0.412 and a deletion correction consistency rate of 0.515, both lower than BZ-2, BZ-3, and BZ-4. This suggests that when the reference genome alignment rate was low, the consistency results of BZ-7 did not enter the same high consistency range after deletion correction. Overall, the figure can intuitively show the impact of changes in the reference genome alignment ratio on the quality-weighted consistency rate and the deletion correction consistency rate, and corresponds to the judgment process in which samples BZ-2, BZ-3 and BZ-4 are merged into one individual, while BZ-7 is not involved in the merging result.

[0048] In this implementation scheme, the site missing information in candidate sample pairs is transformed from simple statistical analysis into a missing perturbation coefficient that can be used for homology determination. This distinguishes the different impacts of unilateral missing, bilateral missing, and overall missing proportions on the consistency results, reducing the interference of low coverage, insufficient reference fit, and conditional domain offset on the quality-weighted consistency rate. At the same time, by generating missing correction consistency rate, missing correction mismatch rate, conditional scale calibration consistency rate, conditional scale calibration mismatch rate, and scale drift compensation, candidate sample pairs under different conditional domain labels can be compared at a unified scale. This provides stable inputs that have undergone missing correction and conditional domain calibration for subsequent sample pair difference matrix construction, relation distribution division, and sample pair homology support value calculation.

[0049] Specifically, the process of dividing relational distribution clusters by dimensionality reduction clustering based on conditional scale calibration consistency rate, conditional scale calibration mismatch rate, scale drift compensation, and distribution migration is as follows: The result after subtracting the conditional scale calibration consistency rate, the conditional scale calibration mismatch rate, the scale drift compensation, the missing proportion, the depth distribution migration, the quality distribution migration, and the likelihood distribution migration are written into the sample pair difference matrix. Each row of the difference matrix corresponds to a candidate sample pair, and each column corresponds to a calibrated difference metric dimension, which constitutes the structured feature input for subsequent dimensionality reduction clustering. The sample pair difference matrix is ​​processed by principal component analysis to obtain the sample pair difference features. The principal component analysis method is used to compress the repetitive information in the sample pair difference matrix and retain the main difference features that can distinguish homologous sample pairs, separated sample pairs, and transitional sample pairs.

[0050] The Gaussian Mixture Model (GMM) clustering method is used to divide the differential features of all samples under the same conditional domain label into relationship distribution clusters, resulting in homologous relationship clusters, separating relationship clusters, and transitional relationship clusters. The number of clusters in the GMM clustering method is consistent with the number of homologous relationship clusters, separating relationship clusters, and transitional relationship clusters. Clustering results with high conditional scale calibration consistency rate and low conditional scale calibration mismatch rate are labeled as homologous relationship clusters, clustering results with low conditional scale calibration consistency rate and high conditional scale calibration mismatch rate are labeled as separating relationship clusters, and the remaining clustering results are labeled as transitional relationship clusters. Sample pairs within the transitional clusters need to be verified with additional evidence to avoid misjudgments caused by forced merging.

[0051] In this implementation scheme, the multidimensional difference information of candidate sample pairs is transformed into clusterable and comparable sample pair difference features. This allows conditional scale calibration consistency rate, conditional scale calibration mismatch rate, scale drift compensation, missing proportion, and distribution migration to jointly participate in relation distribution partitioning, enhancing the distinguishability between homologous sample pairs, separated sample pairs, and transitional sample pairs. At the same time, by partitioning relation distribution clusters under the same conditional domain label, homology determination no longer relies on a single fixed threshold. This allows for adaptation to sample difference distributions under different sequencing qualities, site missing degrees, and reference genome adaptation states, providing a stable relation determination basis for subsequent generation of homology support values, homology relation edges, separated relation edges, and output of relation labels to be verified.

[0052] Specifically, the process of generating sample pair homology support results and determining homology is as follows: Using a Bayesian evidence fusion method, the posterior probability of the homology relationship distribution cluster, the posterior probability of the separation relationship distribution cluster, the conditional scale calibration consistency rate, the conditional scale calibration mismatch rate, the site cardinality confidence coefficient, and the scale drift compensation amount are used as inputs to obtain the sample pair homology support value. The Bayesian evidence fusion method uses the proportion of sample pairs in each relationship distribution cluster under the same conditional domain label as a prior, and the statistical distributions corresponding to the conditional scale calibration consistency rate, conditional scale calibration mismatch rate, site cardinality confidence coefficient, and scale drift compensation amount as the likelihood basis. Through the Bayesian evidence fusion method, the posterior probability assigned by clustering is used as a prior constraint, and evidence fusion calculation is performed with the site-level calibration statistics. The homology support value of sample pairs is obtained, and the homology determination threshold is determined based on the distribution characteristics of the posterior probabilities of the homology relationship distribution cluster and the separation relationship distribution cluster. The homology determination threshold is determined by the boundary position of the homology support value of the corresponding sample pairs in the homology relationship distribution cluster and the separation relationship distribution cluster. The homology determination threshold adaptively fluctuates with the overall sequencing quality and site retention rate of the current identification batch. When the homology support value of a sample pair is greater than or equal to the homology determination threshold, a homology relationship edge is generated; when the homology support value of a sample pair is less than the homology determination threshold and belongs to the separation relationship distribution cluster, a separation relationship edge is generated; when the sample pair belongs to the transition relationship distribution cluster, a relationship label to be verified is generated; when the homology support value of a sample pair is less than the homology determination threshold and does not belong to the separation relationship distribution cluster, a relationship label to be verified is generated.

[0053] In this embodiment, Table 1 is a data table for determining the homology support of candidate sample pairs. It records in detail the quality-weighted consistency rate, missing perturbation coefficient, conditional scale calibration consistency rate, relation distribution cluster, homology support value of sample pairs, and generated relation markers of different candidate sample pairs in the homology determination process. It is used to show the changes of different candidate sample pairs from site evidence calibration to relation determination output. Among them, the generated relation marker is used to indicate the relation output type formed by the candidate sample pair after homology determination, including homology relation edge, separation relation edge, and relation marker to be verified. The weighted consistency rates between BZ-2 and BZ-3, BZ-2 and BZ-4, and BZ-3 and BZ-4 were 0.8476, 0.8508, and 0.8464, respectively, with missing perturbation coefficients of 0.061, 0.054, and 0.063, respectively. After conditional scaling, the consistency rates improved to 0.956, 0.962, and 0.954, respectively, with corresponding homology support values ​​of 0.942, 0.951, and 0.936, respectively. All samples were assigned to the homology distribution cluster and homology edges were generated. The weighted consistency rates between BZ-2 and BZ-7, and between BZ-4 and BZ-7 were 0. The missing perturbation coefficients of BZ-5 and BZ-6 increased to 0.386 and 0.421, respectively, while the conditional scale calibration consistency rates were 0.528 and 0.501. The corresponding sample pair homology support values ​​decreased to 0.183 and 0.154, respectively, all falling into the separation relation distribution cluster and generating separation relation edges. The quality-weighted consistency rate of BZ-5 and BZ-6 was 0.7020, the missing perturbation coefficient was 0.184, the conditional scale calibration consistency rate was 0.748, and the sample pair homology support value was 0.622, placing them in a transitional state between homology and separation relations. Therefore, they were classified into the transitional relation distribution cluster and generated a relation label for review. Thus, Table 1 reflects the corresponding changes in missing perturbation, scale calibration, and homology support values.

[0054] Table 1. Data on candidate samples' ability to determine homology.

[0055] Candidate sample pairs Quality-weighted consistency rate Missing perturbation coefficient Conditional scale calibration consistency rate Relationship distribution cluster Samples of homologous support values Generate relation tags BZ-2 and BZ-3 0.8476 0.061 0.956 Homology distribution clusters 0.942 Homologous edge BZ-2 and BZ-4 0.8508 0.054 0.962 Homology distribution clusters 0.951 Homologous edge BZ-3 and BZ-4 0.8464 0.063 0.954 Homology distribution clusters 0.936 Homologous edge BZ-2 and BZ-7 0.4120 0.386 0.528 Separation Relationship Distribution Cluster 0.183 Separate relation edges BZ-4 and BZ-7 0.3860 0.421 0.501 Separation Relationship Distribution Cluster 0.154 Separate relation edges BZ-5 and BZ-6 0.7020 0.184 0.748 Transitional relation distribution clusters 0.622 Pending review relationship marker

[0056] like Figure 3The figure shows a stacked decomposition diagram of the quality-weighted consistency rate and the missing perturbation coefficient of the candidate sample pair. The horizontal axis represents the candidate sample pair, and the vertical axis represents the numerical value. The "Net Contribution of Quality-Weighted Consistency Rate" in the legend is a graphical display item, which is obtained by subtracting the missing perturbation coefficient from the quality-weighted consistency rate. It is used to visually display the part of the quality-weighted consistency rate that is not weakened by the missing perturbation, and is not used as a new judgment parameter in subsequent calculations. As shown in Table 1, the net contribution of quality-weighted consistency rate in the bars corresponding to BZ-2 and BZ-3, BZ-2 and BZ-4, and BZ-3 and BZ-4 is relatively large, while the proportion of missing perturbation coefficient is relatively small. This indicates that the effective consistency of the above candidate sample pairs is relatively strong, which is consistent with the results of classifying them into the homology distribution cluster and generating homology relationship edges. The missing perturbation coefficient in the bars corresponding to BZ-2 and BZ-7, and BZ-4 and BZ-7 is relatively large, and the net contribution of quality-weighted consistency rate is significantly reduced. This indicates that the corresponding candidate sample pairs are significantly affected by site missing and insufficient evidence quality, which is consistent with the results of classifying them into the separation relationship distribution cluster and generating separation relationship edges. The bars corresponding to BZ-5 and BZ-6 are in an intermediate state. Neither the net contribution of quality-weighted consistency rate nor the missing perturbation coefficient shows obvious homology or obvious separation characteristics, which is consistent with the results of classifying them into the transitional relationship distribution cluster and generating relationship markers to be verified. Overall, the figure can intuitively show how the quality-weighted consistency rate of different candidate sample pairs changes after being weakened by missing perturbations, and it corresponds to the homology support value and generation relationship label of sample pairs in Table 1.

[0057] In this implementation plan, the results of relation distribution division, conditional scale calibration, and site base confidence are uniformly converted into the homology support value of sample pairs. The homology determination threshold is adaptively determined based on the support distribution of homology relation distribution clusters and separation relation distribution clusters within the current identification batch, so that the generation of homology relation edges, separation relation edges, and relation markers to be verified has clear data basis.

[0058] Specifically, the process of constructing a signed sample similarity graph, performing clustering and merging to generate candidate sample groups for the same individual is as follows: Figure 4The diagram shows the flowchart for clustering and merging signed sample similarity graphs. Sample numbers are used as graph nodes, and edges representing common origins and separations are used as signed graph edges. Positive signs for common origin edges indicate attraction, while negative signs for separations indicate repulsion. This signification mechanism allows the graph structure to simultaneously encode both similarity and dissimilarity information. The sample's support value for its common origin is used as the edge weight to construct the signed sample similarity graph. The graph node representations in the signed sample similarity graph are updated using signed graph attention, resulting in sample node embedding vectors. Signed graph attention takes sample nodes, signed graph edges, and edge weights as input, performing forward aggregation on neighboring nodes connected by edges representing common origins and backward constraint on neighboring nodes connected by edges representing separations. The update round is determined according to the sample size within the same identification batch. The signed graph attention mechanism performs directional aggregation of neighbor information based on edge symbols during message passing, so that the node embedding vectors tend to move closer to attractive neighbors and move further away from repulsive neighbors. The initial sample merging results are clustered and merged through intra-group consistency statistics and inter-group separation statistics to obtain the same individual candidate sample group. Intra-group consistency statistics are used to retain sample groups with concentrated homology edges and fewer separation edges, while inter-group separation statistics are used to ensure that there is sufficient separation support between different sample groups. The same individual candidate sample group enters the conflict verification and result output module to calculate the confidence of the candidate sample group, and then serves as the basis for the output of the animal individual identification results.

[0059] In this implementation scheme, the homology edges, separation edges, and homology support values ​​of candidate sample pairs are converted into a signed sample similarity graph at the overall sample level. This allows homology support and separation constraints between different samples to participate in the merging judgment within the same graph structure. By combining sample node embedding vectors with intra-group consistency statistics and inter-group separation statistics, the impact of misjudgment of a single sample pair on the overall individual attribution can be reduced, improving the structural consistency and inter-group exclusivity of the same candidate sample group. This provides stable graph structure support for subsequent calculation of merging conflict values, generation of candidate sample group confidence, and output of animal individual identification results.

[0060] Specifically, the process of performing evidence discount fusion to generate candidate sample group confidence scores and output the number of animal individuals and their attribution labels is as follows: Traverse sample pairs within the same individual candidate sample group, counting the number of separating edges within the group and the sum of their corresponding graph edge weights; traverse sample pairs between different candidate sample groups of the same individual, counting the number of homology edges across groups and the sum of their corresponding graph edge weights; obtain the merge conflict value based on the number of separating edges within the group, the sum of their weights, the number of homology edges across groups, and the sum of their weights; employ evidence discount fusion rules based on conflict factors to fuse the homology support value of sample pairs, the locus cardinality confidence coefficient, the scale drift compensation, and the merge conflict value to obtain the candidate sample group confidence score, where a larger merge conflict value results in a larger discount; the homology support value of sample pairs and the locus cardinality confidence coefficient are used to improve the candidate sample group confidence score, and the scale drift compensation is used to adjust the scale drift compensation. Drift compensation and merging conflict values ​​are used to reduce the confidence of candidate sample groups. Evidence discount fusion maps the merging conflict value to a discount factor. The greater the conflict, the deeper the discount on positive evidence within the group, so that the final confidence simultaneously reflects intra-group consistency and inter-group exclusivity. The number of candidate sample groups of the same individual whose confidence meets the output conditions and has not triggered the pending verification relationship marker is taken as the number of animal individuals. The group identifier corresponding to the candidate sample group of the same individual is taken as the individual attribution marker corresponding to the sample number within the group. The sample number within the group is associated with the physical evidence number and output. The sample number in the sample pair carrying the pending verification relationship marker is taken as the sample number to be verified. The sample number to be verified is obtained by deduplicating all sample numbers in the sample pair carrying the pending verification relationship marker. The output includes the number of animal individuals, the individual attribution marker corresponding to the sample number and the associated physical evidence number, the confidence of the candidate sample group and the sample number to be verified.

[0061] In this implementation plan, intra-group and inter-group conflict verification is performed on the same candidate sample group of an individual. The weights of separation relationship edges, cross-group homology relationship edges, and corresponding graph edges are converted into merged conflict values. The confidence level of the candidate sample group simultaneously reflects intra-group consistency, inter-group exclusivity, the reliability of the locus cardinality, and the impact of scale drift, avoiding the direct output of low-confidence candidate sample groups as animal individual identification results. At the same time, the same candidate sample group of an individual is split by output conditions and the relationship marker to be verified, so that the number of animal individuals, individual attribution markers, physical evidence number association results, and sample numbers to be verified have clear conflict verification basis and traceability.

[0062] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0063] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. An animal individual identification method based on multi-source feature fusion, characterized in that, Includes the following steps: S1: Obtain sample registration records, basic identification records, raw sequencing files, and reference genome files within the same identification batch and perform preprocessing; S2, based on basic identification records, sample registration records and reference genome files, constructs conditional domain labels, establishes a multi-source site evidence matrix, and generates conditional domain offsets and distribution migrations; S3. Based on the conditional domain offset and distribution migration, calculate the quality-weighted consistency rate and the quality-weighted mismatch rate, and perform missing correction and conditional domain hierarchical calibration to obtain the conditional scale calibration consistency rate, conditional scale calibration mismatch rate and scale drift compensation amount. S4, based on conditional scale calibration consistency rate, conditional scale calibration mismatch rate, scale drift compensation amount and distribution migration amount, divides the relational distribution clusters through dimensionality reduction clustering; Generate samples to support the homology results and determine homology; S5. Construct a similarity graph of signed samples, perform clustering and merging, and generate candidate sample groups of the same individual; S6 performs evidence discounting fusion, generates candidate sample group confidence scores, and outputs the number of animal individuals and their attribution markers.

2. The animal individual identification method based on multi-source feature fusion according to claim 1, characterized in that: The specific process for obtaining and preprocessing sample registration records, basic identification records, raw sequencing files, and reference genome files within the same identification batch is as follows: Obtain sample registration records, basic identification records, raw sequencing files, and reference genome files for multiple animal samples to be identified within the same identification batch. Sample registration records include sample number, physical evidence number, tissue type, sampling site, storage duration, sequencing batch number, and reference genome file identifier. Basic identification records include morphological observation records and mitochondrial species confirmation records. Raw sequencing files include read numbers, read base sequences, and read quality characters. Reference genome files include chromosome numbers, reference coordinates, and reference base sequences; the reference coordinates are the positional identifiers within the reference genome coordinate system. After quality character parsing, adapter removal, and quality splicing of the original sequencing file, the base sequences of the reads are aligned to the reference coordinate system of the reference genome file to obtain the alignment result file. Then, duplicate read removal and variant site detection are performed to obtain a candidate site evidence library and determine the site reference coordinates. Genotype interpretation results, genotype likelihood values, site sequencing depth values, site quality values, and genotype quality values ​​are obtained from the effective coverage read statistics of candidate variant sites and genotype likelihood estimation. By statistically analyzing the alignment read markers, reference coordinate coverage counts, and soft splicing operation markers in the alignment result file, the alignment ratio, average sequencing depth, coverage uniformity, and soft splicing ratio are statistically analyzed to form a set of sample alignment statistical features. The average sequencing depth and storage duration records were processed using the Z-score normalization algorithm; the locus quality value, genotype quality value, alignment ratio, coverage balance, and soft shearing ratio were processed using the min-max normalization algorithm.

3. The animal individual identification method based on multi-source feature fusion according to claim 1, characterized in that: The specific process for constructing a multi-source locus evidence matrix based on conditional domain labels derived from basic identification records, sample registration records, and reference genome files is as follows: Hierarchical combination coding was performed based on mitochondrial species confirmation records, reference genome file identifiers, sequencing batch numbers, and tissue types. Morphological observation records were used as species anchoring auxiliary information, and sampling sites were used as tissue type supplementary verification information. The hierarchical combination coding results were hashed to generate conditional domain labels. Under the same mitochondrial species confirmation results, animal samples to be identified were paired to obtain candidate sample pairs. A multi-source locus evidence matrix was constructed based on the candidate locus evidence library. The multi-source locus evidence matrix was indexed and matched according to sample number, chromosome number, and locus reference coordinates. The corresponding genotype interpretation results, genotype likelihood values, locus sequencing depth values, locus quality values, and genotype quality values ​​were read from the candidate locus evidence library and filled into the matrix cells.

4. The animal individual identification method based on multi-source feature fusion according to claim 1, characterized in that: The specific process for generating the conditional domain offset and distribution migration is as follows: The sample quality condition vector is obtained by performing feature splicing and principal component whitening on the sample alignment statistical feature set and the standardized storage duration record; the reference genome file identifier, alignment ratio, coverage balance and soft splicing ratio are obtained by performing feature hash encoding and principal component whitening on the reference adaptation feature vector. The evidence matrix of multi-source sites under the same conditional domain label is processed by the conditional quantile calibration method to obtain the conditional domain site interpretation boundary; based on the conditional domain site interpretation boundary, sites with genotype interpretation results are screened to obtain the set of interpretable sites in the sample. Candidate sample pairs are matched based on chromosome number and locus reference coordinates. Paired loci belonging to the set of identifiable loci for both animal samples to be identified are included in the set of common single nucleotide polymorphism (SNP) loci. Paired loci with the same genotype interpretation result are counted as the number of identical loci, and paired loci with different genotype interpretation results are counted as the number of mismatch loci. The total number of loci in the set of common SNP loci is counted as the number of common SNP loci. A minimum hash sketch is constructed for the set of identifiable sites in the sample to obtain the prior value of site overlap; the conditional domain offset is obtained by statistically processing the sample quality conditional vector and the reference adaptation feature vector through the difference in the mean and quantile differences of the conditional vector; the distribution of site sequencing depth, site quality, and genotype likelihood is processed by the optimal transmission distance to obtain the depth distribution migration, quality distribution migration, and likelihood distribution migration.

5. The animal individual identification method based on multi-source feature fusion according to claim 1, characterized in that: The specific process for calculating the quality-weighted consistency rate and quality-weighted mismatch rate based on conditional domain offset and distribution migration is as follows: Empirical distribution function quantile mapping was performed on the sequencing depth, quality, and genotype likelihood values ​​of sites in the common single nucleotide polymorphism (SNP) site set to obtain depth quantiles, quality quantiles, and likelihood quantiles. The depth quantiles, quality quantiles, likelihood quantiles, conditional domain offsets, depth distribution migration, quality distribution migration, and likelihood distribution migration were fused using a non-additive evidence integration method to obtain the site information contribution value. The quality-weighted consistency rate is obtained by summing the contribution values ​​of loci with the same genotype interpretation and dividing by the sum of all loci's contribution values. The quality-weighted mismatch rate is obtained by summing the contribution values ​​of loci with different genotype interpretations and dividing by the sum of all loci's contribution values. The distribution of the number of common single nucleotide polymorphism (SNP) sites, the prior value of site overlap, and the number of common SNP sites of all candidate sample pairs under the same conditional domain label are input into the beta binomial empirical shrinkage rule to obtain the locus cardinality confidence coefficient.

6. The animal individual identification method based on multi-source feature fusion according to claim 1, characterized in that: The specific process of performing missing correction and conditional domain hierarchical calibration to obtain the conditional scale calibration consistency rate, conditional scale calibration mismatch rate, and scale drift compensation is as follows: The union of the reference coordinates of all animal samples to be identified under the same condition domain label in the candidate site evidence library is taken as the complete set of candidate sites for the condition domain; the total number of sites in the complete set of candidate sites for the condition domain where only one animal sample to be identified has a genotype interpretation result is the number of unilateral deletion sites, and the total number of sites where neither animal sample to be identified has a genotype interpretation result is the number of bilateral deletion sites. The missing proportion is obtained by summing the number of unilateral and bilateral missing sites and dividing by the number of sites in the entire set of candidate sites for the conditional domain. The number of unilateral and bilateral missing sites were scaled up according to the number of sites in the full set of candidate sites in the conditional domain. The scaled-up results were combined with the missing site ratio, and the combined results were corrected according to the conditional domain offset, depth distribution migration, and quality distribution migration to obtain the missing perturbation coefficient. The credibility is recalibrated based on the quality-weighted consistency rate, quality-weighted mismatch rate, site cardinality confidence coefficient, and missing perturbation coefficient to obtain the missing correction consistency rate and missing correction mismatch rate. The conditional domain is hierarchically calibrated based on the missing correction consistency rate, missing correction mismatch rate, site cardinality confidence coefficient, conditional domain offset, and conditional domain label to obtain the conditional scale calibration consistency rate, conditional scale calibration mismatch rate, and scale drift compensation amount.

7. The animal individual identification method based on multi-source feature fusion according to claim 1, characterized in that: The specific process of dividing relational distribution clusters through dimensionality reduction clustering based on conditional scale calibration consistency rate, conditional scale calibration mismatch rate, scale drift compensation, and distribution migration is as follows: The result after subtracting the conditional scale calibration consistency rate, the conditional scale calibration mismatch rate, the scale drift compensation, the missing proportion, the depth distribution migration, the quality distribution migration, and the likelihood distribution migration are written into the sample pair difference matrix; the sample pair difference matrix is ​​processed by principal component analysis to obtain the sample pair difference features. The GMM clustering method is used to divide the differential features into relational distributions based on all samples under the same conditional domain label, resulting in homologous relational distribution clusters, disjoint relational distribution clusters, and transitional relational distribution clusters.

8. The animal individual identification method based on multi-source feature fusion according to claim 1, characterized in that: The specific process of generating samples to support the homology results and determining homology is as follows: Using a Bayesian evidence fusion method, the posterior probabilities of the homology relationship distribution cluster, the posterior probabilities of the separation relationship distribution cluster, the conditional scale calibration consistency rate, the conditional scale calibration mismatch rate, the site cardinality confidence coefficient, and the scale drift compensation amount are used as inputs to obtain the homology support value of the sample pair. Based on the distribution characteristics of the posterior probabilities of the homology relationship distribution cluster and the separation relationship distribution cluster, the homology determination threshold is determined. When the homology support value of the sample pair is greater than or equal to the homology determination threshold, a homology relationship edge is generated; when the homology support value of the sample pair is less than the homology determination threshold and belongs to the separation relationship distribution cluster, a separation relationship edge is generated; when the sample pair belongs to the transition relationship distribution cluster, a relationship label to be verified is generated.

9. The animal individual identification method based on multi-source feature fusion according to claim 1, characterized in that: The specific process of constructing a signed sample similarity graph, performing clustering and merging to generate candidate sample groups of the same individual is as follows: Using sample IDs as graph nodes, homology edges and separation edges as signed graph edges, and sample pair homology support values ​​as graph edge weights, a signed sample similarity graph is constructed. The graph node representations in the signed sample similarity graph are updated through signed graph attention to obtain sample node embedding vectors. The initial sample merging results are clustered and merged through intra-group consistency statistics and inter-group separation statistics to obtain candidate sample groups for the same individual.

10. The animal individual identification method based on multi-source feature fusion according to claim 1, characterized in that: The specific process of performing evidence discounting fusion, generating candidate sample group confidence scores, and outputting the number of animal individuals and individual attribution labels is as follows: Traverse the sample pairs within the same candidate sample group of an individual, and count the number of separating edges within the group and the sum of the corresponding graph edge weights; traverse the sample pairs between different candidate sample groups of the same individual, and count the number of homologous edges across groups and the sum of the corresponding graph edge weights. The merging conflict value is obtained based on the number of separation edges within a group, the sum of the weights of separation edges within a group, the number of homologous edges across groups, and the sum of the weights of homologous edges across groups. An evidence discount fusion rule based on conflict factors is used to fuse the homology support value of sample pairs, the confidence coefficient of the locus cardinality, the scale drift compensation, and the merge conflict value to obtain the confidence score of candidate sample groups. The number of candidate sample groups of the same individual is taken as the number of animal individuals, and the group identifier of the candidate sample group of the same individual is taken as the individual attribution label corresponding to the sample number within the group. The sample number within the group is associated with the physical evidence number and output. The sample number in the sample pair carrying the relationship label to be verified is taken as the sample number to be verified. The output includes the number of animal individuals, the individual attribution label corresponding to the sample number and the associated physical evidence number, the confidence score of the candidate sample group, and the sample number to be verified.

Citation Information

Patent Citations

  • Method, device, equipment and storage medium for constructing population pedigree

    CN114974414B

  • Methods and apparatus for analyzing whole genome sequencing data

    CN116246705B