Bovine embryo pre-implantation genetic assessment method and system based on whole genome sequencing

By using whole-genome sequencing and Bayesian inference technology, the problems of insufficient sample size and unstable identification of interfering factors in pre-implantation genetic assessment of bovine embryos have been solved, enabling reliable genetic risk assessment and breeding decisions, and improving breeding efficiency and stability.

CN121838874APending Publication Date: 2026-04-10HENAN QINGNIU SIYUAN BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing bovine embryo preimplantation genetic assessment technologies face challenges such as insufficient signal-to-noise ratio due to low sample size, fragmented utilization of information from materials from different sources, unstable identification of interfering factors such as chimeras and contamination, and lack of a closed-loop judgment mechanism for iterative verification, leading to the risk of misjudgment and uncertainty in decision-making.

Method used

DNA data from the paternal, maternal, trophoblast, and culture medium were obtained through whole-genome sequencing. Haplotype data were constructed, and genome segmentation was performed by combining Bayesian inference and read depth observation. The posterior probabilities of aneuploidy, copy number variation, chimerism, and contamination events were calculated, and the final conclusion was output through iterative decision-making.

Benefits of technology

It enables reliable output of embryonic genetic risk conclusions under low sample size and interference conditions, supports review and iterative decision-making, and improves breeding efficiency and industry stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838874A_ABST
    Figure CN121838874A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biological information processing, in particular to a cattle embryo pre-implantation genetic evaluation method and system based on whole genome sequencing, and the method comprises the following steps: obtaining whole genome sequencing data of a male parent and a female parent, embryo trophoblast live detection sequencing data and culture solution free desoxyribonucleic acid sequencing data; performing variation detection and haplotype phasing on male and female parent data to form haplotype data, and calculating a haplotype transmission posterior at an anchor point based on trophoblast data; performing genome segmentation based on reading depth observation and allele frequency observation of a trophoblast and a culture solution, and performing Bayesian inference by combining haplotype transfer posteriori as priori to obtain posteriori probabilities of aneuploid, copy number variation, chimera, pathogenic homozygosis and pollution events; and outputting a passing / rechecking / elimination conclusion, adding sequencing and updating the posteriori during rechecking until ending, and outputting a final conclusion and a genome breeding value, so that the evaluation reliability and the decision interpretability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of biological information processing, and in particular to a bovine embryo pre-implantation genetic evaluation method and system based on whole genome sequencing. BACKGROUND

[0002] In the large-scale breeding of dairy cows and beef cattle, the genetic quality at the embryonic stage directly affects the subsequent survival rate, production performance consistency, and population improvement speed. Earlier and more accurate identification of chromosomal abnormalities, structural variations, and single-gene pathogenic risks helps to reduce ineffective implantation and empty pregnancies, reduce breeding costs, and shorten the expansion cycle of excellent genetic resources. At the same time, combining genetic information with breeding values can promote the shift from "post-selection" to "source screening" in population improvement, and improve breeding efficiency and industry stability. Existing evaluation methods generally face the following problems: low sample size leading to insufficient signal-to-noise ratio, fragmented information utilization of different source materials (trophoblast and culture medium), unstable identification of interference factors such as chimeras and contamination, and lack of a closed-loop decision mechanism for iterative review, resulting in misjudgment risk and decision uncertainty. SUMMARY

[0003] The present application provides a bovine embryo pre-implantation genetic evaluation method and system based on whole genome sequencing, which at least solves the problem of how to reliably output embryo genetic risk conclusions and support iterative review decisions under the conditions of low volume, heterogeneous sequencing information, and contamination / chimeric interference.

[0004] In a first aspect, the present application provides a bovine embryo pre-implantation genetic evaluation method based on whole genome sequencing, comprising the following steps: Obtaining paternal whole genome sequencing data, maternal whole genome sequencing data, bovine embryo trophoblast live detection sequencing data, and culture medium free deoxyribonucleic acid sequencing data; Constructing haplotype data based on the paternal whole genome sequencing data and the maternal whole genome sequencing data, and calculating haplotype transmission posteriori based on bovine embryo trophoblast live detection sequencing data at anchor sites; Segmenting the genome based on the read depth observation and allele frequency observation of the bovine embryo trophoblast live detection sequencing data and the culture medium free deoxyribonucleic acid sequencing data, and obtaining posterior probabilities of aneuploidy events, copy number variation events, chimeric events, pathogenic allele homozygous events, and sample contamination events using Bayesian inference; Outputting a pass conclusion, a review conclusion, or a discard conclusion according to the posterior probabilities, obtaining review sequencing data to update the posterior probabilities until the termination condition is met under the review conclusion, and outputting a final conclusion and a genomic breeding value.

[0005] In a second aspect, the present application provides a bovine embryo pre-implantation genetic evaluation system based on whole genome sequencing, which is used to implement the bovine embryo pre-implantation genetic evaluation method based on whole genome sequencing. The system comprises: a data acquisition module, configured to acquire paternal whole genome sequencing data, maternal whole genome sequencing data, bovine embryo trophoblast viability detection sequencing data, and culture fluid free deoxyribonucleic acid sequencing data; a haplotype module, configured to construct haplotype data based on the paternal whole genome sequencing data and the maternal whole genome sequencing data, and calculate haplotype transmission posteriori based on the bovine embryo trophoblast viability detection sequencing data at anchor point sites; a segmentation inference module, configured to perform genome segmentation based on read depth observation and allele frequency observation of the bovine embryo trophoblast viability detection sequencing data and the culture fluid free deoxyribonucleic acid sequencing data, and obtain posterior probability of aneuploidy event, copy number variation event, chimera event, pathogenic allele homozygous event and sample contamination event by Bayesian inference; a conclusion review module, configured to output a passing conclusion, a review conclusion or a discarded conclusion according to the posterior probability, acquire review sequencing data to update the posterior probability until a termination condition is met under the review conclusion, and output a final conclusion and a genomic breeding value.

[0006] Compared with the prior art, the application has the following advantages and beneficial effects: Through parental sequencing variation detection and haplotype phasing technical means, traceable modeling effect of the genetic source of the embryo is achieved; through the technical means of counting allele support at anchor point sites and calculating haplotype transmission posteriori, the effect of converting parental information into embryo inference priori is achieved; through the technical means of simultaneously utilizing read depth observation and allele frequency observation of the trophoblast and the culture fluid and performing genome segmentation, the stable positioning effect of the local abnormal segment is achieved; through the technical means of fusing priori and observation and performing Bayesian inference, the probabilistic determination effect of aneuploidy, copy number variation, chimera, pathogenic homozygous and contamination event is achieved; through the technical means of review region driven additional sequencing and posteriori iteration, the closed loop decision effect of “reviewable and terminable” is achieved, and the genomic breeding value for selection is simultaneously output. BRIEF DESCRIPTION OF DRAWINGS

[0007] Figure 1 FIG. 1 is a schematic diagram of the execution process of the method of the application; Figure 2 FIG. 2 is a schematic diagram of haplotype transmission posteriori at anchor point sites in a specific embodiment of the application; Figure 3 FIG. 3 is a schematic diagram of read depth observation segmentation of a chromosome window in a specific embodiment of the application; Figure 4 FIG. 4 is a schematic diagram of allele frequency observation segmentation of a chromosome window in a specific embodiment of the application; Figure 5 FIG. 5 is a schematic diagram of comparison of posterior probability of multiple events in a specific embodiment of the application; Figure 6 A structural diagram of the system of the present application. DETAILED DESCRIPTION

[0008] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure.

[0009] The pre-implantation genetic evaluation of bovine embryos refers to detecting and interpreting genetic variations and genetic burdens carried by embryos before embryo transfer, to determine whether they have the genetic basis to continue cultivation, transfer and breeding. Its focus includes not only numerical or fragment abnormalities at the chromosomal level, but also pathogenic variations affecting important economic traits or health traits, and complex situations such as genetic chimerism that may lead to subsequent developmental instability. Compared with genetic testing at the postnatal or adult stage, pre-implantation evaluation emphasizes completing reliable determination within a very limited sample size and time window, and maintaining the continuity of embryo activity and cultivation process as much as possible, so as to identify genetic risks in advance to the transfer decision-making link without changing the existing breeding rhythm. Based on this, the present application proposes a genetic evaluation scheme for the pre-implantation scenario of bovine embryos, which takes whole genome sequencing data as the core input, considers multiple sources of sample information, and outputs reviewable conclusions.

[0010] As shown in Figure 1 A bovine embryo pre-implantation genetic evaluation method based on whole genome sequencing, comprising the following steps: Obtaining paternal whole genome sequencing data, maternal whole genome sequencing data, bovine embryo trophoblast biopsy sequencing data, and culture fluid free deoxyribonucleic acid sequencing data; The same batch of paternal and maternal individuals are selected as genetic source individuals, and blood or ear tissue samples are collected and genomic deoxyribonucleic acid is extracted. Whole genome sequencing is used to obtain paternal whole genome sequencing data and maternal whole genome sequencing data, which are used to represent the whole genome genetic variation information of the parents. For the bovine embryo to be evaluated, a trophoblast biopsy is performed at the blastocyst stage to obtain a trophoblast cell sample, genomic deoxyribonucleic acid is extracted and whole genome sequencing is performed to obtain bovine embryo trophoblast biopsy sequencing data, which is used to represent the nuclear genomic signal of the bovine embryo. The culture fluid sample corresponding to the bovine embryo is collected synchronously, and the free deoxyribonucleic acid is extracted and sequenced to obtain the culture fluid free deoxyribonucleic acid sequencing data, which is used to provide an independent evidence source related to the bovine embryo. By establishing the correspondence between the bovine embryo identification and the identification of various samples at the collection stage, the input consistency of haplotype transmission inference and event posterior evaluation is ensured, thereby reducing the evaluation deviation caused by sample mismatch.

[0011] After acquiring paternal whole-genome sequencing data, maternal whole-genome sequencing data, bovine embryo trophoblast viability detection sequencing data, and culture medium free DNA sequencing data, read quality control and reference genome alignment were performed on the paternal whole-genome sequencing data, maternal whole-genome sequencing data, bovine embryo trophoblast viability detection sequencing data, and culture medium free DNA sequencing data, respectively. Based on the reference genome alignment results, read depth observation and allele frequency observation were generated.

[0012] Read quality control was performed on paternal whole-genome sequencing data, maternal whole-genome sequencing data, bovine embryo trophoblast viability detection sequencing data, and culture medium-free DNA sequencing data to ensure consistency with the data foundation for subsequent alignments and observations. Read quality control included removing sequencing adapter sequences, filtering low-quality reads, and removing reads of abnormal length. The data after read quality control were also deduplicated to reduce the impact of amplification bias on read depth observations. Subsequently, reference genome alignment was performed on each type of data after read quality control based on the bovine reference genome. Alignment result files in a uniform format were output, and these files were sorted and indexed to support rapid statistical analysis at genomic windows and anchor sites. To avoid systematic errors introduced by cross-sample alignment differences, consistent alignment parameters and filtering rules were used for all data, including retaining only unique aligned reads and removing low-quality reads.

[0013] When generating read depth observations, the reference genome is divided into a set of genome windows according to a preset window length. The read depth of each window is obtained by counting the number of covered reads based on the alignment results file. Genome base composition correction and standardization are then performed on the read depths to reduce coverage fluctuations caused by differences in base composition between different windows. Genome base composition correction can be performed using a regression correction method with the window base composition as the independent variable, outputting the corrected read depths, which are then used to generate read depth observations for genome segmentation and event inference.

[0014] When generating allele frequency observations, the first step is to determine the set of loci used for statistics. This set of loci is consistent with the anchor loci used in subsequent haplotype transmission posterior calculations. Based on the alignment result file, the number of supporting reads for reference alleles and alternative alleles is counted at the locus set. After base quality and alignment quality filtering, the locus allele frequencies are calculated. Subsequently, the locus allele frequencies are summarized according to the genomic window to obtain the windowed allele frequencies, thus forming the allele frequency observations. Through the above processing, read depth observations and allele frequency observations are established under the same reference genomic coordinate system and generated by the same read quality control and reference genome alignment process. This allows them to serve as direct inputs for subsequent genome segmentation and Bayesian inference, thereby reducing posterior probability bias caused by differences in data processing.

[0015] Haplotype data were constructed based on paternal and maternal whole-genome sequencing data, and haplotype transfer posterior was calculated at anchor sites based on bovine embryo trophoblast viability detection sequencing data. When constructing haplotype data based on paternal and maternal whole-genome sequencing data, variant detection is first performed in both datasets to obtain a set of variant sites for phasing. These variant sites are then divided into multiple haplotype segments based on linkage relationships, forming paternal and maternal haplotypes. Anchor sites are selected as heterozygous variant sites in either the paternal or maternal whole-genome sequencing data that can distinguish parental haplotypes. These sites are used to locate the transmission source in bovine embryo trophoblast viability sequencing data. At the anchor sites, the number of supporting reads for the reference and alternative alleles in the bovine embryo trophoblast viability sequencing data is counted, and low-quality supporting reads are filtered out under unified rules for read quality control and reference genome alignment. Subsequently, the relative support levels of paternal and maternal haplotype transmission were calculated for each anchor site, and the transmission continuity of adjacent anchor sites was combined to jointly solve the whole chromosome, obtaining the haplotype transmission posterior covering each haplotype segment. This haplotype transmission posterior provides parental source constraints for subsequent genome segmentation and event posterior probability calculations, thereby reducing the fluctuations in transmission determination caused by uneven coverage of bovine embryo trophoblast viability sequencing data and allele deletions.

[0016] The construction of haplotype data based on paternal and maternal whole-genome sequencing data includes: performing variant detection on paternal and maternal whole-genome sequencing data to obtain variant detection results, and performing haplotype phasing based on the variant detection results to obtain paternal and maternal haplotypes, thus forming haplotype data.

[0017] In this embodiment, paternal and maternal whole-genome sequencing data are first used to generate haplotype data that can be used for subsequent haplotype transmission inference. To ensure the reproducibility of the haplotype data, variant detection is performed after read quality control and reference genome alignment are completed for both paternal and maternal whole-genome sequencing data. Variance detection targets single nucleotide variants and small insertion / deletion variants on the reference genome. The process includes counting the number of supporting reads for reference and alternative alleles at candidate sites, removing low-confidence supporting reads based on sequencing and alignment quality filtering rules, and outputting genotype determination and genotype quality markers for each site, thus forming the variant detection results. To avoid inconsistencies between the paternal and maternal variant sets leading to a lack of alignment basis for subsequent transmission inference, the variant detection results are merged and aligned using a unified reference genome coordinate system to obtain a joint site set containing paternal genotype, maternal genotype, and site quality markers.

[0018] When performing haplotype phasing, a set of combined loci is used as input. Loci that meet the locus quality requirements and possess parental information are first selected as the phasing point set. Parental informational loci include loci where the father is heterozygous and the mother is homozygous, loci where the mother is heterozygous and the father is homozygous, and loci where both the father and mother are heterozygous and phasing can be determined. Subsequently, the phasing point set is divided into multiple phasing segments along the chromosome direction using linkage disequilibrium information between adjacent loci. Within each phasing segment, the allele arrangement of the paternal and maternal haplotypes is determined, thus obtaining the paternal and maternal haplotypes. To facilitate subsequent attribution at anchor sites, the paternal and maternal haplotypes are further structured and stored in a format that includes chromosomes, segment start and end positions, and locus allele lists, forming haplotype data.

[0019] Through the above processing, haplotype data serves two purposes: first, it provides candidate sites for the selection of subsequent anchor sites, ensuring a one-to-one correspondence between anchor sites and paternal and maternal haplotypes; second, it provides parental source constraints for the calculation of haplotype transmission posterior, enabling consistency in transmission status at the segmental scale even when bovine embryo trophoblast liveness detection data exhibits allele deletions or uneven coverage, thereby reducing the instability of transmission inference and providing a reliable basic input for subsequent genome segmentation and event posterior probability calculation.

[0020] The calculation of haplotype transmission posterior based on bovine embryo trophoblast liveness detection sequence data at anchor sites includes: selecting paternal and maternal heterozygous sites as anchor sites based on variant detection results; statistically counting alleles at anchor sites based on bovine embryo trophoblast liveness detection sequence data to form haplotype transmission observations; and calculating haplotype transmission posterior based on haplotype transmission observations and combined with haplotype data.

[0021] In this embodiment, anchor sites are selected from variant detection results and a one-to-one correspondence is established with the paternal and maternal haplotypes in the haplotype data. The selection rules include: the site is heterozygous in the paternal parent and homozygous in the maternal parent; the site is heterozygous in the maternal parent and homozygous in the paternal parent; the site is located within a haplotype segment that has undergone haplotype phasing; the site's coverage in both paternal and maternal whole-genome sequencing data meets a preset minimum coverage requirement; the site is not located in low-complexity or repetitive sequence regions; and there are no consistency gap regions around the site that would affect alignment. Through these selections, the anchor sites are ensured to be discriminative of parental origin, and the impact of alignment ambiguity and allele deletions on transmission inference is reduced.

[0022] In establishing haplotype transmission observations, based on the reference genome alignment results of bovine embryo trophoblast viability sequencing data, the number of supporting reads for reference and alternative alleles is counted at each anchor site. Before counting, base quality filtering, alignment quality filtering, and duplicate read filtering are performed to obtain allele support counts for transmission inference. Subsequently, based on the phase information of the anchor sites in the haplotype data, the allele support counts are mapped to paternal and maternal haplotype support counts, thus forming the haplotype transmission observation.

[0023] In calculating the haplotype transmission posterior, the sequence of anchor sites along the chromosome direction is used as input to establish a transmission state set, which is composed of the paternal and maternal haplotype transmission states. The continuity of transmission states between adjacent anchor sites is used as a prior constraint to jointly solve for the haplotype transmission observations at each anchor site, thus obtaining the haplotype transmission posterior. in, To transmit state, This is the candidate transmission state. The number of single-type support read segments corresponding to the candidate transmission state. This represents the total number of supported read segments at the anchor point after filtering. Let be the prior probability given by the continuity of the transitive state. This represents the observed probability obtained from read support counting in the candidate transmission state. The haplotype transmission posterior is used to provide parental origin constraints for subsequent genome segmentation and event posterior probability calculations, thereby reducing the risk of misjudgment caused by relying solely on read depth observations and allele frequency observations.

[0024] Genome segmentation was performed based on read depth and allele frequency observations of bovine embryo trophoblast viability detection sequence data and culture medium free deoxyribonucleic acid sequencing data. Bayesian inference was used to obtain the posterior probabilities of aneuploidy events, copy number variation events, chimeric events, pathogenic allele homozygous events, and sample contamination events. Read depth and allele frequency observations from bovine embryo trophoblast viability sequencing data and culture medium-free DNA sequencing data were input in chromosomal order. Genome segmentation was performed along a reference genome to obtain a set of continuous genome segments and their boundary positions. Genome segmentation was determined based on the consistency of read depth and allele frequency observations between adjacent windows, with boundaries used to divide the entire genome into several statistically stable regions. For each genome segment, a candidate state set was constructed, including aneuploidy events, copy number variation events, chimeric events, homozygous pathogenic allele events, and sample contamination events. Read depth and allele frequency observations were used as segment evidence, combined with haplotype transmission posterior as parental source constraint, and Bayesian inference was used to calculate the posterior probability of each candidate state. By simultaneously utilizing two sources of evidence—trophoblast viability sequencing and culture medium-free DNA—it is possible to distinguish between true copy number changes and apparent shifts caused by sample contamination at the segment scale, providing a probabilistic basis for subsequent approval, review, or rejection conclusions.

[0025] The acquisition of read depth observations includes: dividing the reference genome into genomic windows; obtaining the window read depth based on the coverage statistics of bovine embryo trophoblast viability sequencing data and culture medium free deoxyribonucleic acid sequencing data on the genomic windows; and performing correction and standardization processing on the window read depth based on the genomic base composition to form read depth observations.

[0026] In this embodiment, read depth observations are generated from the alignment results of the reference genome of bovine embryo trophoblast liveness detection sequencing data and the reference genome of culture medium cell-free DNA sequencing data. Read depth observations characterize the coverage intensity at different locations in the reference genome to support subsequent genome segmentation and posterior probability calculations. First, the reference genome is divided into a set of continuous genome windows according to a preset window length, and chromosome number, window start and end coordinates, and window base composition information are recorded for each genome window. Then, coverage statistics are performed on each genome window in the alignment results of bovine embryo trophoblast liveness detection sequencing data and culture medium cell-free DNA sequencing data, respectively. Coverage statistics use a counting method based on the read start position falling within the window range. Before statistics are performed, consistent filtering rules are applied to the alignment results, including removing duplicate reads, removing low-alignment-quality reads, removing non-unique aligned reads, and shielding windows located in highly repetitive sequence regions to reduce coverage expansion caused by alignment ambiguity. This yields the read depth sequences for the bovine embryo trophoblast window and the culture medium window.

[0027] To reduce systematic coverage bias caused by differences in genomic base composition, the base composition ratio of each genomic window was calculated, and the window read depth was corrected based on this ratio. The correction method involved establishing a regression relationship between the window read depth and the window base composition ratio, and then subtracting systematic bias to ensure statistically comparable coverage levels for windows with different base compositions on the same chromosome. After correction, the window read depth was standardized. Standardization calculated the center position and dispersion on a chromosome-by-chromosome basis, mapping the window read depth to a standardized coverage sequence, thus generating observations of bovine embryo trophoblast read depth and culture medium-free DNA read depth. Read depth and allele frequency observations used the same reference genomic coordinate system and were generated under the same read filtering rules. This ensures that subsequent genomic segmentation allows for consistent comparison of coverage changes in read depth observations with proportional changes in allele frequency observations, thereby reducing interference from uneven coverage and sequencing bias in event identification.

[0028] The acquisition of allele frequency observations includes: counting allele support data from bovine embryo trophoblast viability sequencing data and culture medium free deoxyribonucleic acid sequencing data at anchor sites and calculating the allele frequencies at each site; summarizing the allele frequencies at each site according to the genomic window to form allele frequency observations.

[0029] Allele frequency observation is generated through allele support counts at anchor sites. The inputs are the reference genome alignment results of bovine embryo trophoblast biopsy sequencing data, the reference genome alignment results of culture medium free DNA sequencing data, and the aforementioned set of anchor sites. For each anchor site, reads covering that site are extracted from both types of alignment results, and read filtering is performed sequentially: duplicate reads are removed, non-unique aligned reads are removed, reads with alignment quality below a preset threshold are removed, and reads with site base quality below a preset threshold are removed. When the same read overlaps with the same site, only one count is retained to avoid homologous duplication. Subsequently, based on the definitions of reference and alternative alleles at the anchor sites in the variation detection results, the filtered reads are counted according to their base type at the anchor site into the reference allele support count and the alternative allele support count, respectively, thus obtaining the bovine embryo trophoblast biopsy allele support count and the culture medium free DNA allele support count. The site allele frequency is calculated using the following formula: in, Allele frequencies at loci To replace alleles in supporting counting, Allele count is used as a reference.

[0030] To maintain statistical granularity consistent with read depth observations, anchor sites were grouped into the aforementioned genomic window set according to their reference genomic coordinates. Within each genomic window, the alternative allele support count and the reference allele support count were summed to obtain the window alternative allele support count and the window reference allele support count. Then, the window allele frequency was calculated using the aforementioned formula, resulting in observations of bovine embryo trophoblast biopsy allele frequencies and culture medium-free DNA allele frequencies. By using anchor sites as the statistical basis and summarizing at the window level, allele frequency observations can suppress fluctuations in individual site coverage while maintaining parental discriminability, providing stable proportional evidence for subsequent genomic segmentation and posterior probability inference of events.

[0031] Genome segmentation includes: detecting change points by observing read depth and allele frequency along the chromosome direction to determine segment boundaries, and dividing the reference genome into several genome segments based on the segment boundaries.

[0032] When segmenting the genome along the chromosome direction, the genome window set is used as an index to align the read depth observation sequences corresponding to bovine embryo trophoblast viability sequencing data, the read depth observation sequences corresponding to culture medium free DNA sequencing data, and the allele frequency observation sequences corresponding to each of the two types of data according to the window coordinates, forming a window-by-window joint observation record. To reduce the interference of single-window coverage fluctuations on boundary determination, local smoothing is first performed on the read depth observations and allele frequency observations respectively. The smoothing uses the moving median or moving mean on a window-by-window basis, and missing windows are imputed by adjacent windows or marked as unusable windows to ensure that subsequent calculations do not produce false boundaries due to scattered missing data.

[0033] Change point detection uses "consistent changes in statistical characteristics on both sides of the boundary" as the criterion. For each candidate boundary, a preset number of continuous windows are taken on both sides, and the mean observed read depth and mean observed allele frequency are calculated on both sides. The boundary score of the candidate boundary is then calculated. The boundary score can be represented by the following formula: in, Score the boundary. The mean of the read depth observations in the left window of the candidate boundary. The mean of the read depth observations in the window to the right of the candidate boundary. The mean of allele frequencies observed in the window to the left of the candidate boundary. The allele frequency observation mean is represented by the window to the right of the candidate boundary. Candidate boundaries are only included in the scoring when both bovine embryo trophoblast viability sequencing data and culture medium-free DNA sequencing data meet the requirements for window size and observation validity, thus avoiding misjudgment due to the lack of a single source of evidence.

[0034] When determining segment boundaries, the boundary scores of all candidate boundaries are first calculated across the entire chromosome. Then, segment boundaries are selected sequentially from highest to lowest score, and a minimum segment length constraint is applied to the selected boundaries to ensure that any adjacent segment boundaries contain at least a predetermined number of genomic windows. If the distance between a candidate boundary and a selected boundary is less than the minimum segment length, the candidate boundary is discarded, and the process continues with the next candidate boundary. The initial set of selected segment boundaries undergoes boundary refinement. In this refinement, the boundary score is recalculated within the neighborhood window of each selected boundary, and the position with the highest neighborhood score is selected as the final boundary position. Based on the final boundary positions, the reference genome is divided into several genomic segments. For each genomic segment, the segment start and end coordinates, along with the corresponding read depth observation statistics and allele frequency observation statistics, are output. These are used in the subsequent Bayesian inference stage to calculate the posterior probability of each event at the segment scale. By simultaneously constraining the boundary consistency of read depth observations and allele frequency observations, false segmentation caused by local coverage anomalies can be reduced, thereby improving the stability and interpretability of the segmentation results in subsequent event identification.

[0035] The posterior probabilities of aneuploid events, copy number variation events, chimeric events, pathogenic allele homozygous events, and sample contamination events were obtained using Bayesian inference. This involved using the posterior of haplotype transmission as prior information, reading depth observation and allele frequency observation as observation information, and calculating the posterior probabilities of aneuploid events, copy number variation events, chimeric events, pathogenic allele homozygous events, and sample contamination events on genome segments.

[0036] In this embodiment, after obtaining the genome segments, each genome segment is used as an inference unit. The haplotype transmission posterior is used as the prior information, and the read depth observation and allele frequency observation corresponding to the bovine embryo trophoblast viability detection sequence data and culture medium free DNA sequencing data are used as the observation information. The posterior probabilities of aneuploidy events, copy number variation events, chimeric events, pathogenic allele homozygous events, and sample contamination events are calculated. Specifically, read depth observation and allele frequency observation are first summarized within each genome segment. The read depth observation uses the mean or median of the window read depth observations within the segment as the segment coverage index, and the allele frequency observation uses the median and dispersion of the window allele frequency observations within the segment as the segment ratio index. Two sets of segmentation indices are formed for the bovine embryo trophoblast viability detection sequence data and the culture medium free DNA sequencing data, respectively. Subsequently, an event state set is constructed and event priors are defined: the prior of the pathogenic allele homozygous event is constrained by the haplotype transmission posterior to the transmission combination of the pathogenic alleles of the parents; the prior of the aneuploid event and the copy number variation event is constrained by the haplotype transmission posterior to the parental origin consistency of the segment, in order to reduce copy number misjudgment caused by the coverage fluctuation of a single sample; the prior of the chimeric event and the sample contamination event is given by the degree of continuity change of the haplotype transmission posterior in adjacent segments, in order to distinguish between local anomalies and global shifts. For each event state, a coverage consistency test is constructed based on read depth observations, and a proportion consistency test is constructed based on allele frequency observations: for aneuploidy and copy number variation events, the segment coverage index of bovine embryo trophoblast viability sequencing data and culture medium free DNA sequencing data must show a phase shift in the segment proportion index along with a shift in the same direction; for chimeric events, a mixed shift pattern of segment coverage index and segment proportion index is used for matching; for sample contamination events, a pattern of systematic convergence or deviation of segment proportion index between the two types of data is used for matching, combined with the weak correlation characteristics of segment coverage index for exclusion. The posterior probability is calculated using the following core expression: in, For event status, Candidate event status, For observational information on genome segments, Let be the prior probability of the event derived from the haplotype transitive posterior. The observed probability is derived from the degree of matching between read depth observation and allele frequency observation for the candidate event state. The unnormalized results of each candidate event state within the same genomic segment are normalized to obtain the posterior probabilities of the aneuploid event, copy number variation event, chimeric event, pathogenic allele homozygous event, and sample contamination event, which are then used as the probabilistic basis for subsequent output conclusions and review decisions.

[0037] Based on the posterior probability, output the pass conclusion, the review conclusion, or the elimination conclusion. Under the review conclusion, obtain the review sequencing data to update the posterior probability until the termination condition is met, and output the final conclusion and the genome breeding value.

[0038] After obtaining the posterior probabilities of events corresponding to each genomic segment, a pass conclusion, a review conclusion, or an elimination conclusion is output according to preset decision rules. The decision rules include: outputting an elimination conclusion when the posterior probability of an aneuploid event, copy number variation event, or chimeric event meets the elimination criteria; outputting an elimination conclusion when the posterior probability of a pathogenic allele homozygous event meets the elimination criteria; outputting a review conclusion when the posterior probability of a sample contamination event meets the review criteria; and outputting a pass conclusion when the posterior probabilities of all events meet the pass criteria.

[0039] When outputting the verification conclusion, a verification region is determined. The verification region is the genomic segment with a posterior probability within the verification interval or the region corresponding to the set of anchor sites related to sample contamination events. Verification sequencing data corresponding to the verification region is then acquired. Verification sequencing data is obtained by performing targeted resequencing or enhanced whole-genome sequencing on bovine embryo trophoblast biopsy samples and culture medium samples. Based on the verification sequencing data, read depth observations and allele frequency observations of the verification region are regenerated, and the posterior probability is updated according to the aforementioned Bayesian inference process. When the updated posterior probability meets the pass or elimination criteria, the termination condition is met, and the final conclusion is output. The genomic breeding value is calculated by determining the bovine embryo genotype from haplotype data and haplotype transfer posterior, combined with pre-obtained marker effects, and is used for bovine embryo ranking and breeding decisions under the pass conclusion.

[0040] The process of obtaining updated posterior probabilities from retested sequencing data under the retested conclusion includes: determining the retested region based on the posterior probability, where the retested region is the genomic segment used to update the posterior probability within the genomic segmentation; obtaining bovine embryo trophoblast viability retested sequencing data and culture medium free DNA retested sequencing data corresponding to the retested region; updating read depth observation and allele frequency observation based on the retested sequencing data and re-performing Bayesian inference to update the posterior probability; the termination condition is to output a pass conclusion or a rejection conclusion based on the updated posterior probability; the genomic breeding value is calculated by determining the bovine embryo genotype from haplotype data and haplotype transfer posterior, combined with pre-obtained marker effects.

[0041] In this embodiment, after calculating the posterior probabilities of aneuploid events, copy number variation events, chimeric events, homozygous pathogenic allele events, and sample contamination events based on genome segmentation, if the posterior probability triggers a verification conclusion, an iterative update process driven by verification sequencing data is initiated. The verification region is determined by the posterior probability and is limited to the target genome segment within the genome segmentation. The target genome segment satisfies the verification judgment rules, which are jointly determined by the interval assignment of the event's posterior probability, the consistency of the posterior probabilities of adjacent segments, and the evidence consistency between bovine embryo trophoblast viability sequencing data and culture medium free DNA sequencing data on the same segment. This avoids the diffusion of verification costs caused by indiscriminate additional sequencing across the entire genome.

[0042] The acquisition of re-sequencing data was performed separately for each re-sequencing region: re-sequencing data of bovine embryo trophoblast viability sequences corresponding to the re-sequencing region, and re-sequencing data of cell-free DNA from culture medium corresponding to the re-sequencing region. The re-sequencing data underwent read quality control and reference genome alignment rules consistent with the aforementioned data, ensuring that the re-sequencing data and the initial sequencing data were in the same coordinate system and filtering system. Based on the re-sequencing data, read depth observations and allele frequency observations were regenerated within the re-sequencing regions. Read depth observations still used genomic windows as statistical units and underwent correction and standardization based on genomic base composition. Allele frequency observations still counted allele support at anchor sites and summarized by genomic window to ensure consistency in observation definitions before and after the update.

[0043] When updating the posterior probabilities, the original observation information is replaced with the updated read depth and allele frequency observations within the verification region. The Bayesian inference process, which uses haplotype to transmit the posterior as prior information, is then used to recalculate the posterior probabilities of each event within the verification region. Genomic segments outside the verification region retain their original posterior probabilities, thus limiting the impact of the verification to segments containing the uncertainty. The termination condition is set to output either a pass or fail conclusion based on the updated posterior probabilities; the final conclusion is output after the termination condition is met.

[0044] The genomic breeding value is calculated when the output passes the conclusion. The bovine embryo genotype is determined by haplotype data and haplotype posterior, and corresponds to the pre-obtained marker effect. The genomic breeding value is calculated using the following expression: in, This is the value for genomic breeding. For the first A marker for genotype dosage in bovine embryo genotypes For the first The marking effect of a single marker. This represents the number of markers involved in the calculation. By verifying region-driven observation updates and posterior probability iterations, this embodiment concentrates uncertainty on target genome segments without expanding the scope of whole-genome repeat sequencing, thereby improving the stability and consistency of the final conclusions and genome breeding values.

[0045] In one specific embodiment, the complete execution chain and reproducible computational process of the present invention are demonstrated. This embodiment selects a pair of Simmental cattle parents and one bovine embryo to be evaluated. Whole-genome sequencing data from the paternal and maternal parents are collected, along with trophoblast viability sequencing data from the bovine embryo and cell-free DNA sequencing data from the culture medium. Read quality control and reference genome alignment are performed on the four types of sequencing data, and the alignment results are output, generating read depth observations and allele frequency observations in a unified coordinate system. Quality control includes adapter removal, low-quality read removal, and duplicate read marking; the reference genome uses the bovine reference genome version and adheres to unified chromosome nomenclature rules to ensure consistency between anchor site coordinates and genome window division. In the simulated data, the effective reads after paternal and maternal alignment are approximately 7.6 × 10⁻⁶. 8 With 7.4×10 8 The effective reads after biopsy alignment of bovine embryo trophoblasts were approximately 3.2 × 10⁻⁶. 7 The effective reads after alignment of free deoxyribonucleic acid in the culture medium were approximately 1.8 × 10⁻⁶. 7 The comparison rates were all higher than 0.98.

[0046] When constructing haplotype data, variant detection was performed on both paternal and maternal whole-genome sequencing data. The variant detection results were then subjected to consistency filtering, with filtering conditions including sequencing depth support, variant quality, and population frequency constraints, thus forming a variant set suitable for phasing. Haplotype phasing was performed based on this variant set, yielding paternal and maternal haplotypes. The haplotype data was stored in a "chromosome-locus-phase" format. Further, anchor sites were selected based on the variant detection results. Anchor sites were required to satisfy either "paternal heterozygous and maternal homozygous" or "maternal heterozygous and paternal homozygous," and repetitive sequence regions and regions with alignment ambiguity were removed. In the simulated data, the total number of anchor sites was approximately 2.1 × 10⁻⁶. 6 The anchor site located on target chromosome 21 is approximately 9.4 × 10⁻⁶. 4 .

[0047] When calculating the posterior of haplotype transmission, allele support counts are statistically analyzed at anchor sites based on bovine embryo trophoblast liveness detection data to form haplotype transmission observations. The posterior of haplotype transmission is then calculated in conjunction with the haplotype data. Haplotype transmission observations record reference allele support counts and substitution allele support counts at the anchor site granularity, and map them to paternal and maternal haplotype support according to phase. Prior constraints are imposed on the continuity of transmission status to avoid frequent switching in low-coverage areas. Taking a certain anchor site as an example, bovine embryo trophoblast liveness detection data yields a substitution allele support count of 34 and a reference allele support count of 66 at this site. The allele frequency at this site is: in, Allele frequencies at loci To replace alleles in supporting counting, Alleles are used to support the count for reference. This locus... This is more consistent with the 1 / 3 or 2 / 3 morphology commonly seen in trisomy cases, providing directional evidence for subsequent event inferences; the haplotype transmission posterior combines this evidence along the chromosomal direction, outputting the transmission posterior for each genomic window, which is used as prior information input for subsequent Bayesian inference. For example... Figure 2 As shown, the horizontal axis represents the anchor locus index (range 0–300), and the vertical axis represents the haplotype transmission posterior (range 0–1). The solid line in the figure represents the paternal transmission posterior, and the dashed line represents the maternal transmission posterior. Both curves fluctuate continuously with the anchor locus index, with local increases or decreases reflecting the varying transmission probabilities at adjacent anchor locus sites due to differences in the strength of support for paternal / maternal haplotypes from bovine embryo trophoblast biopsy reads. This transmission posterior provides prior information on the "continuity of transmission status" along chromosome directions and serves as a crucial source of prior information for subsequent Bayesian inference.

[0048] Read depth observations were obtained using genomic window statistics. The reference genome was divided into genomic window sets according to fixed window lengths, and the coverage of bovine embryo trophoblast viability sequencing data and culture medium cell-free DNA sequencing data in each window was statistically analyzed to obtain the window read depth. Correction and normalization based on genomic base composition were performed on the window read depths to form read depth observations. The normalized ratio of window read depths on chromosome 21 showed a continuous increase in the simulated data: within the window index range of 101 to 140, the observed ratio of bovine embryo trophoblast read depth fluctuated approximately 1.41–1.56, and the observed ratio of culture medium cell-free DNA read depth fluctuated approximately 1.38–1.53, consistent with the coverage pattern of trisomy or duplicate copies. Figure 3 As shown, the horizontal axis represents the genome window index (0–220), and the vertical axis represents the normalized read depth. The solid line represents the read depth observation of bovine embryo trophoblast viability sequencing (TE) data, and the dashed line represents the read depth observation of culture medium-free deoxyribonucleic acid (cfDNA). The segment boundaries B1 and B2 are marked by two vertical dashed lines, and the genome segment S1 is marked with a light gray shading between B1 and B2. It can be observed that within Segment S1, the normalized read depths of TE and cfDNA generally rise and form a plateau, while outside the segment, both curves fall back and stabilize at around 1.0, thus visually reflecting the abnormal copy number coverage pattern and boundary position of this segment.

[0049] Allele frequency observations were obtained by counting allele support at anchor sites and calculating the allele frequencies at those sites, then summing the results by genomic window. The summarization method involved summing the alternative allele support counts and the reference allele support counts at the anchor sites within the window, and then calculating the window allele frequencies using the aforementioned formula. This resulted in observations of bovine embryo trophoblast allele frequencies and culture medium-free DNA allele frequencies. Simulated data within the same region of chromosome 21 showed that the window allele frequencies shifted from around 0.50 in conventional diploid conditions to around 0.33 or 0.67, aligning with the positional rise observed in read depth, providing evidence of "coverage-proportion" consistency for change point detection. Figure 4 As shown, the horizontal axis represents the genomic window index (0–220), and the vertical axis represents allele frequency. Solid lines represent window allele frequency observations for TE, and dashed lines represent window allele frequency observations for cfDNA. The figure also uses vertical dashed lines to mark the segment boundaries B1 and B2, and light gray shading to indicate Segment S1. It can be observed that outside the segments, both curves fluctuate slightly around 0.5, while within Segment S1, the allele frequency generally shifts downward and forms a relatively stable plateau (around 0.33). Figure 3The reading depth elevation shown remains consistent across the window position, thus providing consistent observational evidence for determining segment boundaries and event type identification.

[0050] The genome segmentation process involves detecting change points along the chromosome direction based on read depth and allele frequency observations to determine segment boundaries. The reference genome is then divided into several segments based on these boundaries. During implementation, boundary scores are constructed using the statistical differences between windows on either side of the candidate boundary, and a minimum segment length constraint is applied to avoid excessive fragmentation. The simulation results show a significant segmentation on chromosome 21, with the segment's start and end points covering windows 101 to 140. The boundary location coincides with the mutation points observed in both read depth and allele frequency observations.

[0051] In the Bayesian inference phase, the posterior probability of haplotype transmission is used as prior information, and read depth observation and allele frequency observation are used as observational information. The posterior probabilities of aneuploidy events, copy number variation events, chimeric events, homozygous pathogenic allele events, and sample contamination events are calculated on the genome segment. The core calculations employ: in, For event status, Candidate event status, For observational information on genome segments, Let be the prior probability of the event derived from the haplotype transitive posterior. This represents the observation probability derived from the degree of matching between read depth observations and allele frequency observations under this candidate event state. Taking chromosome 21 segment as an example, the mean ratio of segment read depth observations is... The expected ratio is 1.0 when the candidate event is diploid and 1.5 when the candidate event is trisomy. Under the same variance assumption, the observation likelihood ratio for the trisomy to diploid read depth is approximately 2.9 × 10⁻⁶. 7 Meanwhile, the observed allele frequencies in the window cluster around 0.33, causing the posterior probability of trisomy events to approach 1 after normalization, thus outputting a high posterior probability for aneuploid events in this segment. For sample contamination events, the observed allele frequencies of free DNA in the culture medium show a global shift towards the parental or exogenous mixing ratio, which is inconsistent with the copy number morphology observed in the read depth; in this embodiment, a local contamination segment is set in the simulated data, whose posterior probability of contamination events is initially inferred to be 0.42, triggering a review conclusion. Figure 5As shown, the vertical axis represents the posterior probability (ranging from 0 to 1), and the horizontal axis represents the event type, including: Aneuploidy, CNV (Copy Number Variation), Mosaicism, Homozygous pathogenic allele, and Sample contamination. The corresponding posterior probability value is labeled at the top of each bar: 0.25 for aneuploidy, 0.10 for copy number variation, 0.05 for mosaicism, 0.85 for homozygous pathogenic allele, and 0.01 for sample contamination. This graph visually demonstrates how the posterior probabilities of different event types can be compared on the same coordinate system within a single Bayesian inference output, thus providing a quantitative basis for the "pass / review / rejection" decision rules.

[0052] Based on the posterior probability, a pass conclusion, a review conclusion, or a rejection conclusion is output. In this embodiment, "the posterior probability of an aneuploid event is higher than the rejection criterion" is directly judged as a rejection conclusion, "the posterior probability of a sample contamination event is within the review interval" is judged as a review conclusion, and "the posterior probabilities of all major events meet the pass criterion" is judged as a pass conclusion. After outputting the review conclusion for the above contamination segments, the review region is determined based on the posterior probability. The review region is limited to the target genome segment used to update the posterior probability, and the bovine embryo trophoblast viability detection sequence review sequencing data and culture medium free deoxyribonucleic acid review sequencing data corresponding to the review region are obtained. The review sequencing data follows the same read quality control and reference genome alignment rules, updates the read depth observation and allele frequency observation, and re-executes Bayesian inference. After simulated review, the posterior probability of the contamination event decreases from 0.42 to 0.06, and the segment is changed to a pass conclusion; the trisomy segment maintains a high posterior probability and retains the rejection criterion, thus the termination condition is met and the final conclusion is output as a rejection conclusion.

[0053] The genomic breeding value is calculated simultaneously with the output of the final conclusion. The bovine embryo genotype is determined by haplotype data and haplotype posterior, and then combined with pre-obtained marker effects for calculation. The formula for calculating the genomic breeding value is: in, This is the value for genomic breeding. For the first A marker for genotype dosage in bovine embryo genotypes For the first The marking effect of a single marker. This represents the number of markers used in the calculation. In the example, 5 markers are used, with marker effects of 0.12, -0.05, 0.08, 0.03, and -0.02, and bovine embryo genotype doses of 2, 1, 0, 1, and 2, respectively. When the final conclusion is "pass," the genomic breeding value is used for sequencing embryos in the same batch; when the final conclusion is "elimination," the genomic breeding value is only used for result recording and model validation, and is not included in the breeding decision.

[0054] The code below generates simulated observations consistent with the description above, outputting two graphs: a genome window read depth observation curve and an allele frequency observation curve, along with a demonstration of segmentation and posterior probability calculations. Simply replace the simulated array in the code with the measured statistical array for direct use in real data analysis.

[0055] import numpy as np import matplotlib.pyplot as plt # ---- 1) Construct window coordinates ---- np.random.seed(7) n_win = 220 # Number of chromosome windows x = np.arange(n_win) # ---- 2) Simulated diploid baseline + trisomy segmentation (101~140) ---- baseline = 1.0 + np.random.normal(0, 0.03, size=n_win) tri_seg = slice(101, 141) depth_te = baseline.copy() depth_cf = baseline.copy() depth_te[tri_seg] += 0.47 + np.random.normal(0, 0.02, size=tri_seg.stop-tri_seg.start) depth_cf[tri_seg] += 0.45 + np.random.normal(0, 0.02, size=tri_seg.stop-tri_seg.start) # ---- 3) Simulated allele frequencies: diploid around 0.50 + trisomy segment around 0.33 ---- af_te = 0.50 + np.random.normal(0, 0.03, size=n_win) af_cf = 0.50 + np.random.normal(0, 0.03, size=n_win) af_te[tri_seg] = 0.33 + np.random.normal(0, 0.02, size=tri_seg.stop-tri_seg.start) af_cf[tri_seg] = 0.34 + np.random.normal(0, 0.02, size=tri_seg.stop-tri_seg.start) # ---- 4) Simplified Change Point Detection: Boundary Score S ---- def boundary_score(depth, af, k=8): S = np.zeros_like(depth) for i in range(k, len(depth)-k): Dl, Dr = np.mean(depth[ik:i]), np.mean(depth[i:i+k]) AFl, AFr = np.mean(af[ik:i]), np.mean(af[i:i+k]) S[i] = abs(Dl-Dr) + abs(AFl-AFr) return S S_te = boundary_score(depth_te, af_te) b1 = np.argmax(S_te) # Demonstration: Take the maximum score as the boundary candidate # ---- 5) Simplified Bayesian: Comparing only diploid (1.0) and trisomy (1.5) ---- def norm_like(d, r, sigma=0.08): return np.exp(-((dr)**2) / (2*sigma**2)) # Select the mean of the three-body segments as the observation. d_obs = float(np.mean(depth_te[tri_seg])) # Priors are derived from haplotypes and passed down to posteriors: Demonstration settings include a trisomy prior of 0.02 and diploid prior of 0.98. P2, P3 = 0.98, 0.02 L2, L3 = norm_like(d_obs, 1.0), norm_like(d_obs, 1.5) post3 = (L3*P3) / ((L3*P3)+(L2*P2)) post2 = 1.0 - post3 print("The average depth of the three-body segment reading observation d = ", round(d_obs, 3)) print("Diploid posterior=", round(post2, 6), "Trisomy posterior=", round(post3, 6)) # ---- 6) Drawing ---- plt.figure() plt.plot(x, depth_te, label="TE segment depth observation") plt.plot(x, depth_cf, label="Observation of depth of free DNA readout in culture medium") plt.axvspan(tri_seg.start, tri_seg.stop-1, alpha=0.2, label="Three-Body Segmentation") plt.legend() plt.title("Read Segment Depth Observation (Standardized Ratio)") plt.xlabel("Genome Window Index") plt.ylabel("Reading segment depth observation") plt.figure() plt.plot(x, af_te, label="Observation of TE allele frequency") plt.plot(x, af_cf, label="Observation of Free DNA Allele Frequency in Culture Medium") plt.axvspan(tri_seg.start, tri_seg.stop-1, alpha=0.2, label="Three-Body Segmentation") plt.legend() plt.title("Observation of Allele Frequencies") plt.xlabel("Genome Window Index") plt.ylabel("Allele Frequency Observation") plt.show() If you want to include the closed loop of "review conclusion - review sequencing data - update posterior probability" in the same demo code, you can additionally construct "depth and frequency observations corresponding to the review sequencing data" in the above script, replace the observations in the tri_seg interval with them and repeat the posterior calculation, and you can directly obtain a visual evidence chain of "changes in posterior probability before and after review".

[0056] like Figure 6 As shown, a bovine embryo pre-implantation genetic assessment system based on whole-genome sequencing is used to implement the aforementioned bovine embryo pre-implantation genetic assessment method based on whole-genome sequencing. The system includes: The data acquisition module is used to acquire paternal whole-genome sequencing data, maternal whole-genome sequencing data, bovine embryo trophoblast viability detection sequencing data, and culture medium free DNA sequencing data. The hardware of the data acquisition module includes a sequencing data access terminal, a sample identification acquisition unit, a data buffer and interface control unit, and a storage array. The sequencing data access terminal connects to the data output port of the high-throughput sequencer and supports receiving paternal whole-genome sequencing data, maternal whole-genome sequencing data, bovine embryo trophoblast viability detection sequencing data, and culture medium free DNA sequencing data via a wired network interface or high-speed peripheral interface. The sample identification acquisition unit is used to read barcodes or RFID tags and establish a one-to-one mapping with the corresponding sequencing data. The data buffer and interface control unit is used to complete data packetization, verification, and flow control. The storage array is used to persistently store the original read files and metadata for subsequent modules to read and access.

[0057] The haplotype module is used to construct haplotype data based on paternal and maternal whole-genome sequencing data, and to calculate the haplotype transmission posterior based on bovine embryo trophoblast viable sequencing data at anchor sites. The haplotype module consists of a computing unit composed of a general-purpose processor and an accelerated processor, high-speed memory, and local solid-state storage. The computing unit performs variant detection, haplotype phasing, and haplotype data construction based on the paternal and maternal whole-genome sequencing data, and performs statistical allele support counting at anchor sites based on bovine embryo trophoblast viable sequencing data to calculate the haplotype transmission posterior. High-speed memory is used to cache the anchor site set, haplotype fragment index, and counting results, reducing random access overhead. Local solid-state storage is used to store intermediate result files and checkpoint data, thus supporting direct reading by subsequent segmentation inference modules within the same computing node.

[0058] The segmentation inference module is used to segment the genome based on read depth and allele frequency observations of bovine embryo trophoblast viability detection sequence data and culture medium free deoxyribonucleic acid sequencing data. Bayesian inference is used to obtain the posterior probability of aneuploidy events, copy number variation events, chimeric events, pathogenic allele homozygous events, and sample contamination events. The segmentation inference module includes parallel computing nodes, vectorized computing acceleration units, high-throughput storage controllers, and internal interconnect buses in terms of hardware. Parallel computing nodes are used to perform change point detection and genome segmentation on read depth observation and allele frequency observation generated from bovine embryo trophoblast viability detection sequence data and culture medium free deoxyribonucleic acid sequencing data; vectorized computing acceleration units are used to perform batch probability calculations on the segmented observation sequences to support parallel updates of Bayesian inference on each genome segment; high-throughput storage controllers and internal interconnect buses are used to realize the rapid transfer of observation sequences and prior information between computing nodes and storage arrays, thereby outputting the posterior probabilities of aneuploid events, copy number variation events, chimeric events, pathogenic allele homozygous events, and sample contamination events.

[0059] The conclusion verification module outputs a pass conclusion, a verification conclusion, or a rejection conclusion based on the posterior probability. Under the verification conclusion, it acquires verification sequencing data to update the posterior probability until the termination condition is met, and outputs the final conclusion and genome breeding value. The conclusion verification module consists of a policy control processor, a result presentation terminal, a verification data access interface, and an audit storage unit. The policy control processor reads the posterior probability output by the segmented inference module and generates a pass conclusion, a verification conclusion, or a rejection conclusion. The result presentation terminal displays the conclusion and verification region information and triggers the access of verification sequencing data. The verification data access interface receives bovine embryo trophoblast viability detection sequence verification sequencing data and culture medium-free DNA verification sequencing data corresponding to the verification region, and writes the verification data to the storage array to update the posterior probability until the termination condition is met. The audit storage unit records the input data version, posterior probability, and output conclusion for each iteration, and forms a traceable hardware-level log chain when outputting the final conclusion and genome breeding value.

[0060] The above are merely embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A method for pre-implantation genetic assessment of bovine embryos based on whole-genome sequencing, characterized in that, Includes the following steps: Acquire paternal whole genome sequencing data, maternal whole genome sequencing data, bovine embryo trophoblast viability sequencing data, and culture medium free DNA sequencing data; Haplotype data were constructed based on paternal and maternal whole-genome sequencing data, and haplotype transfer posterior was calculated at anchor sites based on bovine embryo trophoblast viability detection sequencing data. Genome segmentation was performed based on read depth and allele frequency observations of bovine embryo trophoblast viability detection sequence data and culture medium free deoxyribonucleic acid sequencing data. Bayesian inference was used to obtain the posterior probabilities of aneuploidy events, copy number variation events, chimeric events, pathogenic allele homozygous events, and sample contamination events. Based on the posterior probability, output the pass conclusion, the review conclusion, or the elimination conclusion. Under the review conclusion, obtain the review sequencing data to update the posterior probability until the termination condition is met, and output the final conclusion and the genome breeding value.

2. The method according to claim 1, characterized in that, After acquiring paternal whole-genome sequencing data, maternal whole-genome sequencing data, bovine embryo trophoblast viability detection sequencing data, and culture medium free DNA sequencing data, read quality control and reference genome alignment were performed on the paternal whole-genome sequencing data, maternal whole-genome sequencing data, bovine embryo trophoblast viability detection sequencing data, and culture medium free DNA sequencing data, respectively. Based on the reference genome alignment results, read depth observation and allele frequency observation were generated.

3. The method according to claim 1, characterized in that, The construction of haplotype data based on paternal and maternal whole-genome sequencing data includes: performing variant detection on paternal and maternal whole-genome sequencing data to obtain variant detection results, and performing haplotype phasing based on the variant detection results to obtain paternal and maternal haplotypes, thus forming haplotype data.

4. The method according to claim 1, characterized in that, The calculation of haplotype transmission posterior based on bovine embryo trophoblast liveness detection sequence data at anchor sites includes: selecting paternal and maternal heterozygous sites as anchor sites based on variant detection results; statistically counting alleles at anchor sites based on bovine embryo trophoblast liveness detection sequence data to form haplotype transmission observations; and calculating haplotype transmission posterior based on haplotype transmission observations and combined with haplotype data.

5. The method according to claim 1, characterized in that, The acquisition of read depth observations includes: dividing the reference genome into genomic windows; obtaining the window read depth based on the coverage statistics of bovine embryo trophoblast viability sequencing data and culture medium free deoxyribonucleic acid sequencing data on the genomic windows; and performing correction and standardization processing on the window read depth based on the genomic base composition to form read depth observations.

6. The method according to claim 1, characterized in that, The acquisition of allele frequency observations includes: counting allele support data from bovine embryo trophoblast viability sequencing data and culture medium free deoxyribonucleic acid sequencing data at anchor sites and calculating the allele frequencies at each site; summarizing the allele frequencies at each site according to the genomic window to form allele frequency observations.

7. The method according to claim 1, characterized in that, Genome segmentation includes: detecting change points by observing read depth and allele frequency along the chromosome direction to determine segment boundaries, and dividing the reference genome into several genome segments based on the segment boundaries.

8. The method according to claim 1, characterized in that, The posterior probabilities of aneuploid events, copy number variation events, chimeric events, pathogenic allele homozygous events, and sample contamination events were obtained using Bayesian inference. This involved using the posterior of haplotype transmission as prior information, reading depth observation and allele frequency observation as observation information, and calculating the posterior probabilities of aneuploid events, copy number variation events, chimeric events, pathogenic allele homozygous events, and sample contamination events on genome segments.

9. The method according to claim 1, characterized in that, The process of obtaining updated posterior probabilities from retested sequencing data under the retested conclusion includes: determining the retested region based on the posterior probability, where the retested region is the genomic segment used to update the posterior probability within the genomic segmentation; obtaining bovine embryo trophoblast viability retested sequencing data and culture medium free DNA retested sequencing data corresponding to the retested region; updating read depth observation and allele frequency observation based on the retested sequencing data and re-performing Bayesian inference to update the posterior probability; the termination condition is to output a pass conclusion or a rejection conclusion based on the updated posterior probability; the genomic breeding value is calculated by determining the bovine embryo genotype from haplotype data and haplotype transfer posterior, combined with pre-obtained marker effects.

10. A bovine embryo pre-implantation genetic assessment system based on whole-genome sequencing, used to implement the bovine embryo pre-implantation genetic assessment method based on whole-genome sequencing as described in any one of claims 1-9, characterized in that, The system includes: The data acquisition module is used to acquire paternal whole genome sequencing data, maternal whole genome sequencing data, bovine embryo trophoblast viability sequencing data, and culture medium free deoxyribonucleic acid sequencing data. The haplotype module is used to construct haplotype data based on paternal and maternal whole-genome sequencing data, and to calculate the haplotype transfer posterior based on bovine embryo trophoblast viability detection sequencing data at anchor sites. The segmentation inference module is used to segment the genome based on read depth and allele frequency observation of bovine embryo trophoblast viability detection sequence data and culture medium free deoxyribonucleic acid sequencing data. Bayesian inference is used to obtain the posterior probability of aneuploidy events, copy number variation events, chimeric events, pathogenic allele homozygous events, and sample contamination events. The conclusion verification module is used to output a pass conclusion, a verification conclusion, or a rejection conclusion based on the posterior probability. Under the verification conclusion, the module obtains verification sequencing data to update the posterior probability until the termination condition is met, and outputs the final conclusion and the genome breeding value.