Construction of the Atlantic salmon pedigree matrix based on multi-source data and the method of selection decision

CN122658408APending Publication Date: 2026-08-28QINGDAO COLORFUL SEED TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610802566.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0002]针对大西洋鲑养殖群体执行亲缘矩阵构建与选配决策是水产基因组育种过程中的关键工序与核心步骤,该流程需要进行多源数据分析,同时对选配方案进行动态管理以保证遗传增益的持续提升以及近交风险的有效控制,确保育种群体的遗传多样性维持和长期选育效率,然而传统选配方法在面对更加复杂的遗传数据结构关联的场景时依然面临着诸多技术瓶颈

Benefits of technology

通过采集大西洋鲑养殖群体的多源遗传数据并进行数据清洗得到标准遗传数据;先后基于该标准遗传数据进行热点区域识别、育种价值矩阵构建、遗传相关系数计算、分层优化搜索、时间偏差识别与个体调整,实现了大西洋鲑亲缘矩阵构建与选配决策;与现有技术相比,通过执行区块边界解析与自适应区段划分,消除了假梯度产生与热点区域位置的偏移;通过引入多代家系计算遗传相关系数,优化了跨代衰减存在的估计偏差;通过配对对象重组凸性检验与分层搜索策略,区分了凸可行域与非凸可行域特征,避免了常规优化算法在联合约束下陷入局部最优或搜索失效的情况;通过育种价值矩阵时效标记与时间偏差触发式更新,实现了育种价值信息与风险监测周期的同步,避免了选配决策基于过时信息导致的调整方向偏差;通过构建关联图谱并执行替补操作,实现了替补个体引入前的连锁风险预判与定向替换,避免了盲目替换引发的大范围配对关系联动调整;提升了选配方案的遗传增益稳定性与近交风险可控性,保证了多代选育过程中群体有效含量的持续维持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122658408A_ABST
    Figure CN122658408A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of biological selection, and discloses a method for constructing a salmonid kinship matrix and making a selection decision based on multi-source data; the method comprises the following steps: collecting multi-source data of a salmonid breeding population and performing data cleaning to obtain standard genetic data; performing linkage block boundary analysis and identifying hot spot regions to output a hot spot region positioning set; performing multi-trait aggregation evaluation and constructing a breeding value priority matrix of individuals in the breeding population; calculating a genetic correlation coefficient and constructing a weighted kinship matrix; performing feasible region feature analysis, executing optimization search to generate a candidate selection scheme; identifying time bias and updating data; constructing a correlation atlas of each candidate individual, executing individual adjustment to output an adjusted selection scheme; and extracting an optimal scheme to generate a selection decision instruction; the method improves the genetic gain stability and inbreeding risk controllability of the selection scheme, and ensures the continuous maintenance of the effective content of the population in the multi-generation breeding process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biological mating technology, and more specifically, to a method for constructing a kinship matrix of Atlantic salmon and making mating decisions based on multi-source data. Background Technology

[0002] Constructing kinship matrices and making mating decisions for Atlantic salmon farmed populations are key procedures and core steps in aquaculture genomics breeding. This process requires multi-source data analysis and dynamic management of mating schemes to ensure continuous improvement of genetic gains and effective control of inbreeding risks, thereby ensuring the maintenance of genetic diversity and long-term breeding efficiency of the breeding population. However, traditional mating methods still face many technical bottlenecks when dealing with more complex genetic data structures and relationships.

[0003] In real-world scenarios, the background recombination rate exhibits significant heterogeneity across different regions of the Atlantic salmon chromosome. The recombination rate in the telomere region is significantly higher than that around the centromere, and there are substantial differences in recombination rates between males and females. However, in traditional selection processes, the microsegmentation of chromosomes often results in misalignment between segment boundaries and the natural boundaries of linkage disequilibrium blocks due to the use of fixed physical distances or fixed SNP counts. This easily leads to false gradients within linkage disequilibrium blocks and shifts in the location of hotspot regions. Furthermore, the long generational intervals in Atlantic salmon make it difficult to obtain sufficiently deep, continuous multi-generational genomic data for verifying the cross-generational IBD decay pattern. Traditional selection methods do not establish a correction relationship between measured cross-generational survival rates and theoretical recombination rates, potentially leading to significant biases when theoretical predictions are directly applied to long-generation species. Additionally, the genetic background of Atlantic salmon breeding populations tends to narrow after multiple generations of artificial selection, making the joint constraints formed by the breeding value priority matrix and the weighted kinship matrix more prone to non-specificity. While convexity is a key feature, traditional mating methods lack the ability to determine the convexity of the feasible region, often leading to local optima or poor solutions. Furthermore, the collection of phenotypic data related to Atlantic salmon is limited by the farming cycle and growth stage, resulting in a lower update frequency of the breeding value matrix compared to the monitoring frequency of the genetic diversity risk index. Traditional mating methods often employ static matrix decision-making logic, lacking continuous monitoring and synchronous updates of the breeding value matrix, leading to adjustments based on outdated information that deviate from the current population state. Moreover, the number of high-value parents in Atlantic salmon farming populations is limited, and their kinship is close. When high-risk pairings require individual replacement, the introduction of a substitute individual may create new high-kinship relationships with other paired individuals in the scheme, triggering a chain replacement demand. However, traditional mating methods often employ local repair replacement logic, lacking assessment of the impact range of the substitute individual's kinship network, leading to single-point replacements disrupting the overall scheme's balance and necessitating large-scale coordinated adjustments.

[0004] In view of this, the present invention proposes a method for constructing and selecting Atlantic salmon kinship matrices based on multi-source data to solve the above problems. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a method for constructing and selecting a mating decision based on a multi-source data Atlantic salmon kinship matrix, comprising: S1. Collect multi-source data from Atlantic salmon farming populations and perform data cleaning to obtain standard genetic data; the standard genetic data includes genome data, target trait data, and chromosome distribution data; S2. Perform linkage block boundary analysis based on chromosome distribution data, adaptively divide chromosomes into segments according to the block boundary analysis results, identify hotspot regions based on the division results, and output a set of hotspot region locations; S3. Based on the target trait data, conduct multi-trait aggregation evaluation, combine the evaluation results to construct a breeding value priority matrix for individuals in the breeding population, and mark the breeding value priority matrix with timeliness. S4. Extract the IBD distribution characteristics from the genomic data, spatially map the IBD distribution characteristics to hotspot regions and calculate the genetic correlation coefficient; construct a weighted kinship matrix based on the genetic correlation coefficient; S5. Perform feasible region feature analysis on the weighted kinship matrix and breeding value priority matrix of candidate individuals. Based on the analysis results, divide the candidate individuals into multiple search levels, perform optimized search on each search level, integrate cross-level search results, and generate candidate matching schemes. S6. Calculate the genetic diversity risk index after the candidate mating scheme is executed, and simultaneously identify time deviation and update data for each breeding value priority matrix; construct the association map of each candidate individual based on the weighted kinship matrix, adjust the individuals executing the candidate mating scheme based on the genetic diversity risk index and the association map, and output the adjusted mating scheme; S7. Perform a comprehensive priority evaluation on the adjusted selection scheme, extract the optimal scheme to generate a selection decision instruction, and send the selection decision instruction to the preset breeding management terminal.

[0006] Furthermore, the method for performing chain block boundary analysis includes: Identify SNP locus information in genomic data, pair SNP loci according to their physical location on chromosomes in chromosome distribution data, and calculate the linkage disequilibrium coefficient for each pair of SNP loci. Set a sliding window and iterate through and count the linkage disequilibrium coefficients of all SNP loci pairs along the direction of the chromosome's physical location, and calculate the mean coefficient of each sliding window; The system identifies the corresponding window position where the mean coefficient transitions between a preset low value range and a preset high value range, and calculates the magnitude of change in the mean coefficient. If the magnitude exceeds a preset threshold, the transition point is used as a candidate boundary point. A fixed distance is set on both sides of each candidate boundary point. The difference between the average coefficients within the fixed distance range on both sides is calculated. Candidate boundary points whose difference is greater than the preset difference threshold are determined as block boundary points. The adaptive segmentation method includes: Using the block boundary point as the segment division position, if there are adjacent segment division positions in the physical position arrangement direction of the same chromosome, then the continuous segment between the adjacent segment division positions is defined as an adaptive segment.

[0007] Furthermore, the method for identifying hotspot areas includes: Extract recombination rate data from the chromosome distribution data within the region corresponding to each adaptive segment, and calculate the mean recombination rate of each adaptive segment; The difference between the average recombination rates of two adjacent adaptive segments is calculated as the recombination rate gradient of the adjacent adaptive segment; the gradient direction is determined based on the sign of the recombination rate gradient, and the adjacent adaptive segments with positive gradient direction are used as the ascending boundary, and the adjacent adaptive segments with negative gradient direction are used as the descending boundary. Calculate the mean overall recombination rate of the chromosome corresponding to each adaptive segment, and calculate the ratio of the absolute value of the recombination rate gradient of each pair of adjacent adaptive segments to the mean overall recombination rate as the gradient ratio. If the gradient ratio is higher than the preset significance threshold, the corresponding rising or falling boundary is determined as the hotspot boundary. At the same time, the segment between adjacent hotspot boundaries is taken as the hotspot region along the physical location arrangement direction of the chromosome. The location information of each hotspot region is integrated to obtain the hotspot region location set.

[0008] Furthermore, the method for conducting multi-trait aggregation evaluation includes: Extract the trait record for each farmed individual from the target trait data and construct a trait record vector; Calculate the correlation coefficient between any two traits in the trait record vector in the breeding population, and construct a trait correlation coefficient matrix based on the correlation coefficient. Eigenvalue decomposition is performed on the trait correlation coefficient matrix, and eigenvectors with eigenvalues ​​greater than a preset eigenvalue threshold are extracted as principal components of the traits. Project the trait record vector of each farmed individual onto the direction of the trait principal component to obtain the individual principal component vector of the corresponding farmed individual; calculate the ratio of the corresponding eigenvalue of each trait principal component to the sum of all remaining eigenvalues ​​to obtain the information contribution rate of the corresponding trait principal component. The joint weight is obtained by multiplying the information contribution rate with the preset breeding target weight of the corresponding trait principal component. The joint weight is then used to sum the weighted components of each individual principal component vector of each cultured individual to obtain the breeding index of the corresponding cultured individual. The methods for constructing the breeding value priority matrix for individuals in the aquaculture population include: The breeding individuals are sorted according to their breeding index to form a breeding value sequence. The number of each breeding individual is used as the row index, and the corresponding position in the breeding value sequence is used as the column index to form a breeding value priority matrix. The construction time of the breeding value priority matrix is ​​recorded, and an effective usage period is set as a time mark.

[0009] Furthermore, the method for performing spatial mapping and calculating genetic correlation coefficients includes: Extract the IBD shared fragment between any two cultured individuals from the genome data, record the location range of the IBD shared fragment, compare the location range with the hot spot region, and calculate the overlap length between the IBD shared fragment and the hot spot region; The overlap lengths of all IBD shared fragments and all hotspot regions for each pair of cultured individuals are summed to obtain the hotspot shared length of the corresponding individual pair; the total length of all IBD shared fragments for each pair of cultured individuals is calculated to obtain the genome shared length of that individual pair. The hotspot coverage rate is obtained by calculating the ratio of the hotspot shared length to the sum of the lengths of all hotspot regions, and the hotspot density is obtained by calculating the ratio of the hotspot shared length to the genome shared length. The sharing enhancement coefficient of the corresponding individual pair is obtained by multiplying the hotspot coverage rate and the hotspot density. Screen for family groups with continuous multi-generational kinship in genomic data and extract IBD shared fragments in the same hotspot region between adjacent generations; The ratio of the length of the IBD shared fragment between offspring and parents in the same hotspot region is calculated to obtain the intergenerational survival rate of the corresponding hotspot region; the genetic segregation decay coefficient of each pair of cultured individuals is calculated based on the intergenerational survival rate and the recombination rate data of the corresponding hotspot region. The genetic correlation coefficient of the corresponding individual pair is obtained by weighting and summing the sharing enhancement coefficient and the genetic segregation decay coefficient according to a preset ratio; The methods for constructing the weighted kinship matrix include: Construct a square matrix using the ID numbers of all cultured individuals as row and column indices, and fill the row and column intersection positions of the corresponding individual pairs with the genetic correlation coefficients of the square matrix to form a weighted kinship matrix.

[0010] Furthermore, the method for performing feasible region feature analysis includes: Identify the breeding index of each candidate individual in the breeding value priority matrix, set a lower limit for the breeding index, and select candidate individuals whose breeding index is higher than the lower limit to form a breeding candidate set; Extract the genetic correlation coefficients between any two candidate individuals in the breeding candidate set from the weighted kinship matrix, and set an upper limit threshold for the genetic correlation coefficients as a kinship constraint condition. Traverse all candidate pairs in the breeding candidate set and select candidate pairs that satisfy the kinship constraint as feasible pairs; Extract several pairs from the feasible pairs and combine them. For each pair combination, perform pair object swapping to output recombined pairs. Calculate the proportion of recombined pairs that satisfy the kinship constraint as the convexity retention rate. If the convexity retention rate is lower than the preset convexity threshold, the feasible region formed by the constraint is determined to be a non-convex feasible region. The proportion of feasible pairs formed by each candidate individual to the total number of possible pairs is used as the pairing feasibility rate, and all candidate individuals are ranked according to the pairing feasibility rate to form the analysis results.

[0011] Furthermore, the method for generating candidate matching schemes includes: The analysis results are divided into levels according to the preset stratification rules. Within each search level, optimization search is performed with the optimization objectives of maximizing the breeding index and minimizing the genetic correlation coefficient to obtain the level matching schemes for each search level. The level matching schemes of all search levels are integrated and the conditions are filtered to obtain candidate matching schemes.

[0012] Furthermore, the methods for identifying time deviations and updating data include: The system acquires the latest monitoring timestamp in real time and calculates the time deviation value by comparing it with the construction time of the breeding value priority matrix. If the time deviation value exceeds the time range corresponding to the time mark, a matrix recalculation instruction is generated; otherwise, the breeding value priority matrix is ​​updated using the data corresponding to the latest monitoring timestamp.

[0013] Furthermore, the method for constructing the association graph of each candidate individual includes: Each candidate individual in the candidate mating scheme is used as a graph node. At the same time, the genetic correlation coefficient of any pair of individuals in the weighted kinship matrix is ​​identified. If the genetic correlation coefficient is higher than the preset association threshold, an association edge is established between the two graph nodes of the individual pair to form an association graph. The number of associated edges for each graph node is counted as the association density; the sum of the genetic correlation coefficients corresponding to all associated edges for each graph node is calculated as the association strength. Identify whether the candidate individuals corresponding to the two graph nodes connected by each associated edge belong to the same pairing in the candidate matching scheme. If they do, mark them as an intra-pair associated edge; otherwise, mark them as an inter-pair associated edge.

[0014] Furthermore, the method of performing individual adjustments includes: Pairs with genetic diversity risk indices all above a preset risk threshold in candidate mating schemes are identified as pairs to be adjusted. Identify the graph nodes in the association graph that correspond to the pairings to be adjusted, and select candidate individuals corresponding to graph nodes with association strength higher than the selectable strength threshold as individuals to be replaced. The candidate individuals in the breeding candidate set whose breeding index difference with the individual to be replaced is less than the index difference threshold and whose association density in the association graph is lower than the preset density threshold constitute the substitute individual set. The number of pairing association edges between the graph node corresponding to each substitute individual in the statistical association graph and the graph nodes corresponding to other paired individuals in the candidate pairing scheme is used as the impact range value. The individual to be replaced is replaced by the substitute individual with the smallest impact range value to generate an adjustment pair. The adjustment pair replaces the pair to be adjusted and the candidate pairing scheme and association map are updated simultaneously. The adjustment process is repeated until the genetic diversity risk index of all pairs is lower than the preset risk threshold. All the finally updated candidate pairing schemes are then integrated to obtain the adjusted pairing scheme.

[0015] The technical effects and advantages of the Atlantic salmon kinship matrix construction and mating decision method based on multi-source data of this invention are as follows: Standard genetic data was obtained by collecting multi-source genetic data from Atlantic salmon farming populations and cleaning the data. Based on this standard genetic data, hotspot region identification, breeding value matrix construction, genetic correlation coefficient calculation, hierarchical optimization search, time deviation identification, and individual adjustment were carried out to realize the construction of Atlantic salmon kinship matrix and mating decision-making. Compared with existing technologies, the generation of false gradients and the offset of hotspot region positions were eliminated by performing block boundary resolution and adaptive segmentation. The estimation bias of cross-generational attenuation was optimized by introducing multi-generational family calculation of genetic correlation coefficients. Convexity test of mating recombinant and hierarchical search strategy were used to distinguish convexity. The row region and non-convex feasible region features prevent conventional optimization algorithms from getting stuck in local optima or failing under joint constraints; by using breeding value matrix time-sensitive marking and time deviation-triggered updates, the breeding value information and risk monitoring cycle are synchronized, avoiding the deviation in adjustment direction caused by outdated information in mating decisions; by constructing association graphs and performing replacement operations, the chain risk prediction and targeted replacement before the introduction of replacement individuals are realized, avoiding large-scale linkage adjustments in pairing relationships caused by blind replacement; the genetic gain stability and inbreeding risk controllability of the mating scheme are improved, ensuring the continuous maintenance of the effective content of the population during multiple generations of breeding. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the Atlantic salmon kinship matrix construction and mating decision-making method based on multi-source data according to the present invention; Figure 2 This is a schematic diagram of the Atlantic salmon kinship matrix construction and selection decision system based on multi-source data according to the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example

[0018] Please see Figure 1 As shown in this embodiment, the Atlantic salmon kinship matrix construction and mating decision-making method based on multi-source data includes: S1. Collect multi-source data from Atlantic salmon farming populations and perform data cleaning to obtain standard genetic data; the standard genetic data includes genome data, target trait data, and chromosome distribution data; S2. Perform linkage block boundary analysis based on chromosome distribution data, adaptively divide chromosomes into segments according to the block boundary analysis results, identify hotspot regions based on the division results, and output a set of hotspot region locations; S3. Based on the target trait data, conduct multi-trait aggregation evaluation, combine the evaluation results to construct a breeding value priority matrix for individuals in the breeding population, and mark the breeding value priority matrix with timeliness. S4. Extract the IBD distribution characteristics from the genomic data, spatially map the IBD distribution characteristics to hotspot regions and calculate the genetic correlation coefficient; construct a weighted kinship matrix based on the genetic correlation coefficient; S5. Perform feasible region feature analysis on the weighted kinship matrix and breeding value priority matrix of candidate individuals. Based on the analysis results, divide the candidate individuals into multiple search levels, perform optimized search on each search level, integrate cross-level search results, and generate candidate matching schemes. S6. Calculate the genetic diversity risk index after the candidate mating scheme is executed, and simultaneously identify time deviation and update data for each breeding value priority matrix; construct the association map of each candidate individual based on the weighted kinship matrix, adjust the individuals executing the candidate mating scheme based on the genetic diversity risk index and the association map, and output the adjusted mating scheme; S7. Perform a comprehensive priority evaluation on the adjusted selection scheme, extract the optimal scheme to generate a selection decision instruction, and send the selection decision instruction to the preset breeding management terminal.

[0019] It should be noted that all technical features included in this embodiment are designed specifically for the aquaculture characteristics of Atlantic salmon. Multi-source data on the biological reproduction process is collected through aquaculture monitoring terminals at Atlantic salmon farms, and data cleaning is achieved through filtering, noise reduction, and missing value imputation. Genomic data includes genotype information for each Atlantic salmon and chromosome-related information for each SNP locus. Target trait data includes measurement records of growth-related traits and observation records of disease resistance-related traits in Atlantic salmon. Chromosome distribution data refers to the physical location information of chromosomes and parameters such as recombination rate.

[0020] Methods for performing chain block boundary analysis include: Identify SNP loci information in genomic data, pair SNP loci according to their physical location on chromosomes in chromosome distribution data, and calculate the linkage disequilibrium coefficient for each pair of SNP loci.

[0021] SNP sites refer to the locations where a single nucleotide in the genome is mutated, which is a common form of genetic variation. In this embodiment, SNP site information includes the chromosome number and the genotype coding information of the corresponding cultured individual at that site. The SNP site information is matched with the physical arrangement order of chromosomes in the chromosome distribution data, and two SNP sites that are physically adjacent on the same chromosome are paired.

[0022] The formula for calculating the cascading imbalance coefficient is as follows: ;in, This represents the linkage disequilibrium coefficient; in a pair of SNP loci, the allele at locus 1 is A / a, and the allele at locus 2 is B / b; where... and This refers to the allele frequency at locus number one, and ; and This refers to the allele frequency at locus 2, and Haplotype frequencies include , , , And the sum is 1; Represents the base coefficient. .

[0023] Set a sliding window and iterate through and count the linkage disequilibrium coefficients of all SNP loci pairs along the direction of the chromosome's physical location, and calculate the mean coefficient of each sliding window.

[0024] The window size of the sliding window is set based on the physical length of the chromosome to ensure that each sliding window can cover a continuous SNP locus pair, and the window size needs to be smaller than the physical length of the chromosome. The sliding window is used to slide along the alignment direction of the chromosome to calculate the mean of the linkage disequilibrium coefficient of all SNP locus pairs contained in each sliding window, which is used to smooth the random fluctuations of individual locus pairs.

[0025] The system identifies the window position where the mean coefficient transitions between a preset low value range and a preset high value range. It also calculates the magnitude of change in the mean coefficient and uses the transition point as a candidate boundary point if it exceeds a preset magnitude threshold.

[0026] The specific ranges of the preset low-value interval and the preset high-value interval are set based on genetic knowledge. A sliding window with the mean value of the identification coefficient is located in the middle range between the preset low-value interval and the preset high-value interval. The change range of the mean value of the coefficient is obtained by calculating the absolute value of the difference between the mean value of the coefficient of the above-mentioned transition sliding window and the mean value of the adjacent sliding window. If it is higher than the preset range threshold set based on the theoretical knowledge of calculating the continuous disequilibrium coefficient in genetics, the SNP site corresponding to the linkage disequilibrium coefficient value in the transition sliding window that is in the preset low-value interval and the preset high-value interval is taken as the transition point, and the transition point is taken as the candidate boundary point.

[0027] A fixed distance is set on both sides of each candidate boundary point. The difference between the average coefficients within the fixed distance range on both sides is calculated. Candidate boundary points whose difference is greater than the preset difference threshold are determined as block boundary points.

[0028] The length of the fixed distance is set based on historical experience related to genetics, ensuring that the range of the fixed distance extending on both sides is sufficient to reflect the trend of the average change of the coefficients on both sides of the candidate boundary point. The absolute value of the difference between the average values ​​of all the chain imbalance coefficients within several sliding windows on both sides of the candidate boundary point is calculated, and this value is compared with a preset difference threshold set based on historical experience, ensuring that it is easier to observe that there is a significant difference in the degree of imbalance on both sides. If the difference is higher than the preset difference threshold, the corresponding candidate boundary point is taken as the block boundary point.

[0029] The adaptive segmentation method includes: Using the block boundary point as the segment division position, if there are adjacent segment division positions in the physical position arrangement direction of the same chromosome, then the continuous segment between the adjacent segment division positions is defined as an adaptive segment.

[0030] The positions at both ends of each adaptive segment are the physical positions of the boundary points of consecutive adjacent blocks on the same chromosome.

[0031] Methods for identifying hotspot areas include: Extract recombination rate data from the chromosome distribution data within the region corresponding to each adaptive segment, and calculate the mean recombination rate of each adaptive segment.

[0032] The recombination rate data of the corresponding segment position in the chromosome distribution data is extracted based on the physical location of each adaptive segment, and the average recombination rate in each adaptive segment is calculated as the data basis for subsequent operations.

[0033] The difference between the mean recombination rates of two adjacent adaptive segments is calculated as the recombination rate gradient of the adjacent adaptive segment, where the recombination rate gradient is used to represent the degree of change in recombination rate in adjacent adaptive segments.

[0034] The gradient direction is determined by the sign of the recombination rate gradient. Adjacent adaptive segments with positive gradient directions are used as the ascending boundary, while those with negative gradient directions are used as the descending boundary.

[0035] The gradient direction is determined by the sign of the recombination rate gradient. A positive sign indicates a positive gradient direction, while a negative sign indicates a negative gradient direction. This is used to distinguish between boundaries where the recombination rate increases from low to high and boundaries where it decreases from high to low, thus giving the boundaries a directional marker.

[0036] Calculate the mean overall recombination rate of the chromosome corresponding to each adaptive segment, and calculate the ratio of the absolute value of the recombination rate gradient of each pair of adjacent adaptive segments to the mean overall recombination rate as the gradient ratio.

[0037] The mean recombination rate refers to the average recombination rate of all sites in the chromosome containing the adaptive segment. The gradient ratio is obtained by calculating the ratio of the absolute value of the recombination rate gradient of adjacent adaptive segments belonging to the same chromosome to the mean recombination rate. The gradient is then normalized using the mean recombination rate to make the gradients of different segments comparable.

[0038] If the gradient ratio is higher than the preset significance threshold, the corresponding rising or falling boundary is determined as the hotspot boundary, and the segment between adjacent hotspot boundaries is taken as the hotspot region along the physical location arrangement direction of the chromosome.

[0039] The preset significant threshold is used to distinguish between normal parameter fluctuations and significant change boundaries. Based on the experience of identifying historical hotspot areas, if the gradient ratio is higher than the preset significant threshold, the corresponding adjacent adaptive segment is taken as the hotspot boundary. The set of continuous adaptive segments between adjacent hotspot boundaries is defined as the hotspot region based on the gradient direction to determine whether it is an ascending or descending boundary.

[0040] The location information of each hotspot area is integrated to obtain a hotspot area location set. The location information includes the specific physical location coordinates of the hotspot boundary of each hotspot area and is bound to the corresponding chromosome number to prevent misidentification.

[0041] Methods for evaluating the aggregation of multiple traits include: Extract the trait record of each farmed individual from the target trait data and construct a trait record vector.

[0042] The trait record includes multiple parameter fields that match the number of each cultured individual, such as trait name label and corresponding physiological data, trait measurement conditions, and disease resistance type trait level fields. The fields in each trait record are combined to form a trait record vector.

[0043] Calculate the correlation coefficient between any two traits in the trait record vector in the breeding population, and construct a trait correlation coefficient matrix based on the correlation coefficient.

[0044] The trait type corresponding to the trait record vector is determined based on the trait name label in the trait record vector. The correlation coefficient in this embodiment is obtained by calculating the Pearson correlation coefficient of the numerical sequence composed of the components of any two traits in the physiological data dimension. All correlation coefficients are filled into a square matrix, and the rows and columns of the square matrix represent the corresponding trait types.

[0045] The trait correlation coefficient matrix is ​​subjected to eigenvalue decomposition, and the eigenvectors with eigenvalues ​​greater than a preset eigenvalue threshold are extracted as principal components of the traits.

[0046] In this embodiment, the eigenvalue decomposition algorithm is used to decompose the trait correlation coefficient matrix, outputting eigenvalues ​​and their corresponding eigenvectors. Based on historical eigenvalue decomposition experience, a preset eigenvalue threshold is set to ensure that eigenvalues ​​significantly higher than the background fluctuation can be filtered out through this threshold. The eigenvectors of the corresponding eigenvalues ​​are used as trait principal components, so that the relevant information in the original trait space is transformed into mutually independent principal component space representations. Each trait principal component represents a trait type.

[0047] Projecting the trait record vector of each cultured individual onto the direction of the trait principal component yields the individual principal component vector of the corresponding cultured individual.

[0048] This involves using vector projection to project the trait record vector corresponding to each cultured individual onto the direction of each trait principal component, thus transforming it into an individual principal component vector corresponding to each trait principal component.

[0049] The information contribution rate of the corresponding trait principal component is obtained by calculating the ratio of the corresponding eigenvalue of each trait principal component to the sum of all remaining eigenvalues.

[0050] The information contribution rate is used to quantify the explanatory power and importance of the corresponding trait to the variation of the population trait. It is obtained by calculating the ratio of the eigenvalue of the principal component of each trait type to the sum of the remaining eigenvalues.

[0051] The joint weight is obtained by multiplying the information contribution rate with the preset breeding target weight of the corresponding trait principal component. The joint weight is then used to sum the weighted components of each individual principal component vector of each cultured individual to obtain the breeding index of the corresponding cultured individual.

[0052] The preset breeding target weights are derived from the values ​​set in the original Atlantic salmon farming plan and are determined based on the relative importance of various traits. The joint weight of the corresponding trait type is obtained by multiplying the information contribution rate of the principal component of each trait with the preset breeding target weight. This joint weight can simultaneously reflect the breeding target preference and the explanatory power of trait variation. The breeding index is obtained by weighted summing the principal component vectors of the corresponding trait type of each farmed individual using the joint weight, which is a comparable breeding evaluation value.

[0053] Methods for constructing a breeding value priority matrix for individuals in aquaculture populations include: Individuals are ranked according to their breeding index to form a breeding value sequence. By constructing the breeding value sequence, the results of multi-trait aggregation evaluation are transformed into a more intuitive priority ranking.

[0054] The number of each cultured individual is used as the row index, and the corresponding position in the breeding value sequence is used as the column index to form a breeding value priority matrix. Each row of the breeding value priority matrix corresponds to a cultured individual, and each column represents the position of the corresponding cultured individual in the breeding value sequence.

[0055] Record the construction time of the breeding value priority matrix and set the effective use period as a time marker.

[0056] The construction time refers to the timestamp of each construction of the breeding value priority matrix; the specific time length of the effective use period is set based on the monitoring period length of the preset breeding management terminal, and the effective use period is added as a time marker to the corresponding breeding value priority matrix.

[0057] Methods for performing spatial mapping and calculating genetic correlation coefficients include: Extract the IBD shared fragment between any two cultured individuals from the genomic data, record the location range of the IBD shared fragment, compare the location range with the hot spot region, and calculate the overlap length between the IBD shared fragment and the hot spot region.

[0058] IBD shared segments refer to continuous segments in the same chromosomal region between two cultured individuals that have a homologous inheritance relationship. The location range refers to the location interval formed by the starting and ending physical positions of the IBD shared segments. The location range of each IBD shared segment is compared with the location information of the hot spot region, and the intersection operation is performed to obtain the overlap length of the spatial overlap between the IBD shared segments and the hot spot region.

[0059] The overlap lengths of all IBD shared segments and all hotspot regions for each pair of cultured individuals are summed to obtain the hotspot shared length for the corresponding pair.

[0060] The summation refers to the summation of the overlap lengths obtained by intersecting all IBD shared fragments with each hotspot region across the entire genome for each pair of cultured individuals, to obtain the hotspot sharing length. This value is used to quantify the overall sharing scale of the corresponding pair of cultured individuals within the hotspot region.

[0061] The total length of all IBD shared fragments for each pair of cultured individuals is calculated to obtain the genome shared length of that pair, where the genome shared length refers to the sum of the physical lengths of all IBD shared fragments across the entire genome of that pair of cultured individuals.

[0062] The hotspot coverage rate is obtained by calculating the ratio of the hotspot shared length to the sum of the lengths of all hotspot regions, and the hotspot density is obtained by calculating the ratio of the hotspot shared length to the genome shared length.

[0063] Hotspot coverage rate represents the degree of coverage of each individual pair with the hotspot area set, while hotspot density represents the degree of concentration of each individual pair's IBD shared fragments in the hotspot area.

[0064] The product of hotspot coverage and hotspot density is used to obtain the sharing enhancement coefficient of the corresponding individual pair. The sharing enhancement coefficient is used to characterize the sharing enhancement effect in the hotspot area. The distribution of IBD shared fragments in the hotspot area is converted into a comparable indicator at the individual pair level through the sharing enhancement coefficient.

[0065] We screened family groups with continuous multi-generational kinship in genomic data and extracted IBD shared fragments in the same hotspot regions of adjacent generations.

[0066] The family group refers to the set of individuals in the pedigree records of the breeding database that can determine the correspondence between parents and offspring, and where individuals in both generations have genomic data. IBD shared fragments between adjacent generations are identified as samples for subsequent operations.

[0067] The intergenerational retention rate of the corresponding hotspot region is obtained by calculating the ratio of the IBD shared fragment length between offspring and parents in the same hotspot region.

[0068] The intergenerational retention rate is used to represent the degree of fragment retention in the corresponding hotspot region during the two-generation transmission process, and is used to correct the theoretical decay of hotspot regions inferred solely from the recombination rate.

[0069] The genetic segregation decay coefficient for each pair of cultured individuals was calculated based on intergenerational retention rate and recombination rate data in corresponding hotspot areas.

[0070] In this embodiment, for each hotspot region, the recombination rate of the hotspot region is converted into genetic distance using the Haldane mapping function. For the pedigree path of an individual pair, the number of meiotic divisions traversed by the path is counted as the generation. The cumulative genetic distance of the hotspot region is obtained by multiplying the genetic distance by the generation, and the segregation attenuation value is calculated. ;in Indicates hotspot areas go through The separation attenuation value of the generation, Indicates hotspot areas The cumulative genetic distance; the segregation decay values ​​of each hotspot region are weighted and summed using the overlap length of the IBD shared fragments of individual pairs in each hotspot region as the weight, to obtain the genetic segregation decay coefficient of the corresponding individual pair; the intergenerational survival rate is used as the correction basis for the segregation decay value based on recombination rate, and is used to correct the theoretical estimation bias between hotspot regions.

[0071] The genetic correlation coefficient of the corresponding individual pair is obtained by weighting and summing the shared enhancement coefficient and the genetic segregation attenuation coefficient according to a preset ratio.

[0072] The preset ratio is used to adjust the contribution of the sharing enhancement coefficient and the genetic segregation decay coefficient to the genetic correlation. Based on the distribution of genetic relationships in historical pedigree data, the spatial sharing enhancement and temporal decay effects are integrated into the genetic correlation coefficient.

[0073] Methods for constructing a weighted kinship matrix include: Construct a square matrix using the ID numbers of all cultured individuals as row and column indices, and fill the row and column intersection positions of the corresponding individual pairs with the genetic correlation coefficients of the square matrix to form a weighted kinship matrix.

[0074] Methods for performing feasible region feature analysis include: Identify the breeding index of each candidate individual in the breeding value priority matrix, set a lower limit for the breeding index, and select candidate individuals whose breeding index is higher than the lower limit to form a breeding candidate set.

[0075] In this embodiment, candidate individuals refer to the breeding individuals set in the original breeding plan for selection and mating. The breeding index corresponding to the breeding individual is searched in the breeding value priority matrix using the number of the breeding individual as an index. A lower limit for the breeding index is set based on the content of the breeding plan, and candidate individuals with a breeding index higher than the lower limit are selected to ensure that the breeding ability of the individuals in the breeding candidate set is not lower than the target level.

[0076] Extract the genetic correlation coefficients of any two candidate individuals in the breeding candidate set from the weighted kinship matrix, and set an upper limit threshold for the genetic correlation coefficients as a kinship constraint.

[0077] The upper limit threshold of the genetic correlation coefficient is set based on the content of the breeding plan, and this threshold is used as a constraint condition for kinship.

[0078] Traverse all candidate pairs in the breeding candidate set, and select candidate pairs that meet the kinship constraint as feasible pairs. Among them, select candidate pairs whose genetic correlation coefficient is not higher than the upper limit threshold of the genetic correlation coefficient as feasible pairs.

[0079] Extract several pairs from the feasible pairs and combine them. For each pair combination, perform the swapping of the paired objects to output the recombined pairs. Calculate the proportion of recombined pairs that satisfy the kinship constraint as the convexity preservation rate.

[0080] In this process, a certain number of feasible pairs without shared individuals are randomly selected from all feasible pairs and combined. The specific number of selections is set based on the breeding plan. In each group of feasible pairs, the pairing objects are interchanged. Each recombined pair is taken as a recombined pair. The proportion of recombined pairs that still satisfy the kinship constraint is counted out of all recombined pairs. This is the convexity retention rate, which is used to measure the connectivity and stability of feasible pairings in the solution space under the pairing and recombination case.

[0081] It should be noted that convexity refers to the fact that in continuous optimization problems, if the new solution obtained by linearly combining any two feasible solutions is still feasible, then the feasible region can be defined as a convex set. If, after swapping the pairs, most of the new pairs obtained by recombining feasible pairs are still feasible, then the set of feasible pairs under joint constraints exhibits connectivity characteristics similar to a convex set in the solution space. Conversely, if most of the recombined pairs no longer satisfy the constraints, then the set of feasible pairs exhibits multiple mutually separate solution clusters in the solution space, i.e., it lacks convexity.

[0082] If the convexity retention rate is lower than the preset convexity threshold, the feasible region formed by the constraints is determined to be a non-convex feasible region.

[0083] The convexity threshold is set based on the optimization requirements of the breeding program for the selection scheme. When the convexity retention rate is lower than the preset convexity threshold, it means that most of the new pairing combinations obtained by pairing and exchanging no longer meet the kinship constraint. The solution space where the pairings selected by the breeding index lower limit and the kinship constraint are located is the feasible region. It does not have convexity characteristics and there are multiple feasible solutions that are separated from each other. Therefore, the feasible region at this time is determined to be a non-convex feasible region.

[0084] The proportion of feasible pairs formed by each candidate individual to the total number of possible pairs is used as the pairing feasibility rate, and all candidate individuals are ranked according to the pairing feasibility rate to form the analysis results.

[0085] All possible pairings refer to the number of all possible pairings preset in the original breeding plan. The pairing feasibility rate is obtained by calculating the proportion of feasible pairs that each candidate individual can form to the total number of possible pairings. This rate is used to reflect the degree of participation of each candidate individual in the non-convex feasible domain.

[0086] The methods for generating candidate configuration schemes include: The analysis results are divided into levels according to the preset stratification rules. Within each search level, optimization search is performed with the optimization objectives of maximizing the breeding index and minimizing the genetic correlation coefficient, so as to obtain the level pairing scheme for each search level.

[0087] The pre-defined stratification rule refers to dividing candidate individuals into several search strata based on the distribution characteristics of pairing feasibility. For example, individuals with high pairing feasibility are included in the priority search strata, while individuals with low pairing feasibility are included in the secondary or final search strata. The specific stratification threshold is set based on stratification experience in historical breeding programs. Within each search strata, maximizing the breeding index and minimizing the genetic correlation coefficient are set as optimization objectives. A multi-objective optimization algorithm is used to optimize the search and obtain the stratification pairing scheme corresponding to each search strata.

[0088] Integrate all search levels' hierarchical matching schemes and perform conditional filtering to obtain candidate matching schemes.

[0089] The conditional screening means that each pairing in the hierarchical pairing scheme must meet two constraints: the breeding index must be higher than the lower limit of the breeding index and the genetic correlation coefficient must not be higher than the upper limit of the genetic correlation coefficient.

[0090] Methods for identifying time deviations and updating data include: In this embodiment, by real-time monitoring of changes in the effective population content and inbreeding coefficient after the implementation of candidate mating schemes, a comprehensive evaluation model based on a neural network model is constructed in conjunction with existing genetic knowledge. Using mating schemes that have been actually implemented in historical breeding data and their corresponding multi-generation population monitoring results as training samples, the actual observed degree of population diversity degradation is used as a label value, and genetically related characteristics are used as inputs to calculate a genetic diversity risk index. This genetic diversity risk index is a parameter that can serve as an overall trigger condition for determining whether to initiate the individual replacement and scheme adjustment process. It can also be used to identify high-risk pairings and local network regions where high-risk individuals are located when combined with association maps, and as a convergence criterion during the adjustment iteration process to determine whether the adjustment has reached the expected risk level, thereby controlling the risk of population diversity within an allowable range while maintaining genetic gains.

[0091] The latest monitoring timestamp is obtained in real time, and the time deviation value is calculated by comparing the latest monitoring timestamp with the construction time of the breeding value priority matrix.

[0092] The latest monitoring timestamp refers to the monitoring time recorded when the genetic diversity risk index is calculated, corresponding to the time when the latest batch of genetic assessment or population monitoring data is generated. The time deviation value is obtained by calculating the difference between the latest monitoring timestamp and the construction time contained in the breeding value priority matrix, which is used to quantify the time lag between the current breeding value priority matrix used for mating decisions and the real-time status of the population.

[0093] If the time deviation value exceeds the time range corresponding to the time mark, a matrix recalculation instruction is generated; otherwise, the breeding value priority matrix is ​​updated using the data corresponding to the latest monitoring timestamp.

[0094] If the time deviation value is higher than the effective usage period corresponding to the time mark, it is considered that the current matrix can no longer accurately reflect the latest population traits and breeding target status. In this case, a matrix recalculation instruction is used to re-start from the target trait data to complete the multi-trait aggregation evaluation and priority matrix construction. If the time deviation value is lower than or equal to the effective usage period, it means that the current matrix still has reference value. At the same time, based on the newly added target trait data corresponding to the latest monitoring timestamp and the latest genetic evaluation results, the breeding index and ranking position fields in the breeding value priority matrix are incrementally updated or adjusted.

[0095] Methods for constructing association graphs for each candidate individual include: Each candidate individual in the candidate mating scheme is used as a graph node. At the same time, the genetic correlation coefficient of any pair of individuals in the weighted kinship matrix is ​​identified. If the genetic correlation coefficient is higher than the preset association threshold, an association edge is established between the two graph nodes of the individual pair to form an association graph.

[0096] The system sets a preset association threshold based on historical experience in constructing genetic association maps. This threshold is used to distinguish between background correlation and high genetic correlation. If the genetic correlation coefficient between a pair of individuals is higher than the preset association threshold, it indicates that there is a high and significant kinship between the pair of individuals. Therefore, an association edge is established between the two individuals. All map nodes and all existing association edges are integrated to form an association map.

[0097] The number of associated edges for each graph node is used as the association density.

[0098] The association density is used to represent the degree of connection of the corresponding candidate individual in a higher kinship network. The larger the value, the higher the genetic correlation between the corresponding candidate individual and more other candidate individuals, and the wider the influence on the overall kinship structure in the mating scheme.

[0099] The sum of the genetic correlation coefficients corresponding to all associated edges of each graph node is used as the association strength.

[0100] The association strength is used to represent the degree of kinship accumulated by the corresponding candidate individuals in a network with a high degree of kinship. Under the same association density, the higher the association strength, the higher the overall genetic correlation coefficient between the corresponding candidate individual and the connected object, and the greater the risk of kinship aggregation in the current mating scheme.

[0101] Identify whether the candidate individuals corresponding to the two graph nodes connected by each associated edge belong to the same pairing in the candidate matching scheme. If they do, mark them as an intra-pair associated edge; otherwise, mark them as an inter-pair associated edge.

[0102] The intra-pair association edge refers to the association edge connecting two candidate individuals in the same selected pair, which is used to characterize the degree of kinship clustering within the selected pair; the inter-pair association edge refers to the association edge connecting candidate individuals between different selected pairs, which is used to characterize the possible latent high kinship between different pair combinations.

[0103] The methods for implementing individual adjustments include: Pairs with genetic diversity risk indices all above a preset risk threshold in the candidate mating schemes are identified as pairs to be adjusted.

[0104] Based on genetic knowledge, a preset risk threshold is set. If the genetic diversity risk index is higher than the preset risk index, it means that the corresponding pair may lead to a decrease in the effective population content or too rapid inbreeding during multiple generations of breeding. Therefore, these pairings are identified as pairings to be adjusted.

[0105] Identify the graph nodes in the association graph that correspond to the pairings to be adjusted, and select candidate individuals corresponding to graph nodes with association strength higher than the selectable strength threshold as the individuals to be replaced.

[0106] Based on the theory of kinship, an optional strength threshold is set, and candidate individuals with a correlation strength higher than the optional strength threshold are selected as objects with a high degree of kinship aggregation. In this embodiment, the pair to be adjusted corresponds to two candidate individuals, and the candidate individual with a higher correlation strength and higher than the optional strength threshold is selected as the individual to be replaced, so that the replacement operation is preferentially applied to the individuals that have a greater impact on the overall network.

[0107] It should be noted that in this embodiment, if the association strength of the two candidate individuals to be paired is not higher than the corresponding selectable strength threshold, it means that although the two individuals in the pair have high risk, their structure is still balanced, and replacement can be temporarily not performed. After other high-risk individuals with higher association strength are processed, the above-mentioned individuals with balanced structure will be processed.

[0108] The replacement individual set consists of candidate individuals whose breeding index difference with the individual to be replaced is less than the index difference threshold and whose association density in the association graph is lower than the preset density threshold.

[0109] The system sets an index difference threshold and a preset density threshold based on historical experience in adjusting the selection scheme. The index difference threshold is used to ensure that the difference between the breeding index of the substitute individual and the individual to be replaced is kept within an acceptable range, so as to minimize the negative impact on the overall genetic gain level during replacement. The association density is lower than the preset density threshold, which means that the corresponding candidate individual has fewer connections in the network of higher kinship and has a lower risk. Candidate individuals that meet the conditions of similar breeding index and low association density are selected as substitute individuals to construct the substitute individual set.

[0110] The number of pairing association edges between the graph node corresponding to each substitute individual in the statistical association graph and the graph nodes corresponding to other paired individuals in the candidate pairing scheme is used as the ripple range value.

[0111] Starting from the graph node corresponding to each substitute individual, the graph nodes corresponding to other paired individuals in the candidate pairing scheme are traversed to count the number of associated edges between pairings. This value is used to represent the range of kinship diffusion that may occur in the network after the corresponding substitute individual is introduced into the pairing scheme. The larger the range value, the more kinship the substitute individual has with more individuals already used in the scheme, and more pairing relationships may need to be adjusted after replacement.

[0112] The individual to be replaced is replaced by the substitute individual with the smallest impact range value to generate an adjustment pair. The adjustment pair replaces the pair to be adjusted and the candidate pairing scheme and association map are updated simultaneously.

[0113] To reduce the impact of adjusting pairing relationships, the substitute individual with the smallest affected range value is selected to replace the individual to be replaced to form an adjusted pairing. After replacing the original pairing to be adjusted with the adjusted pairing, the pairing list in the candidate pairing scheme is updated, and the associated edges related to the corresponding nodes of the individual to be replaced are deleted in the association graph. The associated edges between the corresponding nodes of the substitute individual and the new pairing object that meet the constraints are added. By synchronously updating the candidate pairing scheme and the association graph, the consistency between the pairing relationship and the kinship network structure is maintained.

[0114] The adjustment process is repeated until the genetic diversity risk index of all pairs is lower than the preset risk threshold. All the finally updated candidate pairing schemes are then integrated to obtain the adjusted pairing scheme.

[0115] The repeated adjustment process refers to recalculating the genetic diversity risk index corresponding to the updated scheme after each replacement operation and update of the candidate mating scheme and association map, and identifying the set of pairs with risk indices higher than the preset risk threshold as new pairs to be adjusted. The number of high-risk pairs is gradually reduced. The adjustment process can be terminated when the risk index of all pairs is lower than the preset risk threshold or the upper limit of the number of adjustments set by the breeding plan is reached. The candidate mating scheme at the time of termination is output as the adjusted mating scheme.

[0116] In this embodiment, performing comprehensive priority evaluation refers to constructing a comprehensive evaluation vector for each mating scheme in the adjusted mating scheme set based on multiple evaluation indicators such as the expected value of the population breeding index, the genetic diversity risk index, the number of pairings within the scheme, and the participation level of key families or core groups. The comprehensive evaluation vector is evaluated based on the evaluation criteria set by the Atlantic salmon breeding program. The evaluation results of different mating schemes are ranked using the Pareto dominance-based ranking method to obtain the final priority sequence. One or more mating schemes located at the Pareto optimal front are extracted from the priority sequence as the optimal schemes. The pairing information in the optimal schemes is converted into mating decision instructions, which include the male and female individual numbers of each pair, pairing batch information, execution cycle information, and individual physiological state and trait-related parameters.

[0117] This embodiment obtains standardized genetic data by collecting multi-source genetic data from Atlantic salmon farmed populations and performing data cleaning. Based on this standardized genetic data, it then performs hotspot region identification, breeding value matrix construction, genetic correlation coefficient calculation, hierarchical optimization search, time deviation identification, and individual adjustment, thus realizing the construction of the Atlantic salmon kinship matrix and mating decisions. Compared with existing technologies, by performing block boundary resolution and adaptive segmentation, it eliminates the generation of false gradients and the offset of hotspot region positions; by introducing multi-generational family pedigrees to calculate genetic correlation coefficients, it optimizes the estimation bias of intergenerational attenuation; and by using the convexity test of mating recombinant pairs and a hierarchical search strategy, it distinguishes between... The convex and non-convex feasible region features prevent conventional optimization algorithms from getting stuck in local optima or failing under joint constraints. By using the breeding value matrix time-sensitive marking and time deviation-triggered updates, the breeding value information and risk monitoring cycle are synchronized, avoiding deviations in adjustment direction caused by outdated information in mating decisions. By constructing association graphs and performing replacement operations, the chain risk prediction and targeted replacement before the introduction of replacement individuals are realized, avoiding large-scale linkage adjustments in pairing relationships caused by blind replacement. The genetic gain stability and inbreeding risk controllability of the mating scheme are improved, ensuring the continuous maintenance of the effective content of the population during multiple generations of breeding.

[0118] Example 2 Please see Figure 2 As shown, parts not described in detail in this embodiment are described in Embodiment 1. A system for constructing and selecting Atlantic salmon kinship matrices based on multi-source data is provided, including: The data acquisition module collects multi-source data from Atlantic salmon farming populations and performs data cleaning to obtain standardized genetic data; the standardized genetic data includes genome data, target trait data, and chromosome distribution data. The hotspot identification module performs linkage block boundary analysis based on chromosome distribution data, adaptively divides chromosomes into segments according to the block boundary analysis results, identifies hotspot regions based on the division results, and outputs a set of hotspot region locations. The matrix construction module performs multi-trait aggregation evaluation based on target trait data, constructs a breeding value priority matrix for individuals in the breeding population based on the evaluation results, and marks the breeding value priority matrix with timeliness. The genetic computing module extracts the IBD distribution characteristics from the genomic data, spatially maps these IBD distribution characteristics to hotspot regions, and calculates the genetic correlation coefficient; based on the genetic correlation coefficient, it constructs a weighted kinship matrix. The search module is optimized by performing feasible region feature analysis on the weighted kinship matrix and breeding value priority matrix of candidate individuals. Based on the analysis results, the candidate individuals are divided into multiple search levels. Optimized search is performed on each search level and cross-level search results are integrated to generate candidate mating schemes. The mating adjustment module calculates the genetic diversity risk index after the candidate mating schemes are executed, and simultaneously identifies time deviations and updates data for each breeding value priority matrix; it constructs an association map of each candidate individual based on a weighted kinship matrix, performs individual adjustments to the candidate mating schemes based on the genetic diversity risk index and the association map, and outputs the adjusted mating schemes. The comprehensive evaluation module performs a comprehensive priority evaluation on the adjusted selection scheme, extracts the optimal scheme, generates a selection decision instruction, and sends the selection decision instruction to the preset breeding management terminal; the modules are connected to each other via wired and / or wireless means.

[0119] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0120] All formulas in this manual are dimensionless and calculated numerically. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0121] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for constructing and selecting Atlantic salmon kinship matrices based on multi-source data, characterized in that, include: S1. Collect multi-source data from Atlantic salmon farming populations and perform data cleaning to obtain standard genetic data; the standard genetic data includes genome data, target trait data, and chromosome distribution data; S2. Perform linkage block boundary analysis based on chromosome distribution data, adaptively divide chromosomes into segments according to the block boundary analysis results, identify hotspot regions based on the division results, and output a set of hotspot region locations; S3. Based on the target trait data, conduct multi-trait aggregation evaluation, combine the evaluation results to construct a breeding value priority matrix for individuals in the breeding population, and mark the breeding value priority matrix with timeliness. S4. Extract the IBD distribution characteristics from the genomic data, spatially map the IBD distribution characteristics to hotspot regions, and calculate the genetic correlation coefficient; Construct a weighted kinship matrix based on genetic correlation coefficients; S5. Perform feasible region feature analysis on the weighted kinship matrix and breeding value priority matrix of candidate individuals. Based on the analysis results, divide the candidate individuals into multiple search levels, perform optimized search on each search level, integrate cross-level search results, and generate candidate matching schemes. S6. Calculate the genetic diversity risk index after the candidate mating scheme is executed, and simultaneously identify time deviation and update data for each breeding value priority matrix; construct the association map of each candidate individual based on the weighted kinship matrix, adjust the individuals executing the candidate mating scheme based on the genetic diversity risk index and the association map, and output the adjusted mating scheme; S7. Perform a comprehensive priority evaluation on the adjusted selection scheme, extract the optimal scheme to generate a selection decision instruction, and send the selection decision instruction to the preset breeding management terminal.

2. The method for constructing and selecting Atlantic salmon kinship matrices based on multi-source data according to claim 1, characterized in that, The methods for performing chain block boundary analysis include: Identify SNP locus information in genomic data, pair SNP loci according to their physical location on chromosomes in chromosome distribution data, and calculate the linkage disequilibrium coefficient for each pair of SNP loci. Set a sliding window and iterate through and count the linkage disequilibrium coefficients of all SNP loci pairs along the direction of the chromosome's physical location, and calculate the mean coefficient of each sliding window; The system identifies the corresponding window position where the mean coefficient transitions between a preset low value range and a preset high value range, and calculates the magnitude of change in the mean coefficient. If the magnitude exceeds a preset threshold, the transition point is used as a candidate boundary point. A fixed distance is set on both sides of each candidate boundary point. The difference between the average coefficients within the fixed distance range on both sides is calculated. Candidate boundary points whose difference is greater than the preset difference threshold are determined as block boundary points. The adaptive segmentation method includes: Using the block boundary point as the segment division position, if there are adjacent segment division positions in the physical position arrangement direction of the same chromosome, then the continuous segment between the adjacent segment division positions is defined as an adaptive segment.

3. The method for constructing and selecting Atlantic salmon kinship matrices based on multi-source data according to claim 2, characterized in that, The methods for identifying hotspot areas include: Extract recombination rate data from the chromosome distribution data within the region corresponding to each adaptive segment, and calculate the mean recombination rate of each adaptive segment; The difference between the average recombination rates of two adjacent adaptive segments is calculated as the recombination rate gradient of the adjacent adaptive segment; the gradient direction is determined based on the sign of the recombination rate gradient, and the adjacent adaptive segments with positive gradient direction are used as the ascending boundary, and the adjacent adaptive segments with negative gradient direction are used as the descending boundary. Calculate the mean overall recombination rate of the chromosome corresponding to each adaptive segment, and calculate the ratio of the absolute value of the recombination rate gradient of each pair of adjacent adaptive segments to the mean overall recombination rate as the gradient ratio. If the gradient ratio is higher than the preset significance threshold, the corresponding rising or falling boundary is determined as the hotspot boundary. At the same time, the segment between adjacent hotspot boundaries is taken as the hotspot region along the physical location arrangement direction of the chromosome. The location information of each hotspot region is integrated to obtain the hotspot region location set.

4. The method for constructing and selecting Atlantic salmon kinship matrices based on multi-source data according to claim 3, characterized in that, The methods for conducting multi-trait aggregation evaluation include: Extract the trait record for each farmed individual from the target trait data and construct a trait record vector; Calculate the correlation coefficient between any two traits in the trait record vector in the breeding population, and construct a trait correlation coefficient matrix based on the correlation coefficient. Eigenvalue decomposition is performed on the trait correlation coefficient matrix, and eigenvectors with eigenvalues ​​greater than a preset eigenvalue threshold are extracted as principal components of the traits. Project the trait record vector of each farmed individual onto the direction of the trait principal component to obtain the individual principal component vector of the corresponding farmed individual; calculate the ratio of the corresponding eigenvalue of each trait principal component to the sum of all remaining eigenvalues ​​to obtain the information contribution rate of the corresponding trait principal component. The joint weight is obtained by multiplying the information contribution rate with the preset breeding target weight of the corresponding trait principal component. The joint weight is then used to sum the weighted components of each individual principal component vector of each cultured individual to obtain the breeding index of the corresponding cultured individual. The methods for constructing the breeding value priority matrix for individuals in the aquaculture population include: The breeding individuals are sorted according to their breeding index to form a breeding value sequence. The number of each breeding individual is used as the row index, and the corresponding position in the breeding value sequence is used as the column index to form a breeding value priority matrix. The construction time of the breeding value priority matrix is ​​recorded, and an effective usage period is set as a time mark.

5. The method for constructing and selecting Atlantic salmon kinship matrices based on multi-source data according to claim 4, characterized in that, The methods for performing spatial mapping and calculating genetic correlation coefficients include: Extract the IBD shared fragment between any two cultured individuals from the genome data, record the location range of the IBD shared fragment, compare the location range with the hot spot region, and calculate the overlap length between the IBD shared fragment and the hot spot region; The overlap lengths of all IBD shared fragments and all hotspot regions for each pair of cultured individuals are summed to obtain the hotspot shared length of the corresponding individual pair; the total length of all IBD shared fragments for each pair of cultured individuals is calculated to obtain the genome shared length of that individual pair. The hotspot coverage rate is obtained by calculating the ratio of the hotspot shared length to the sum of the lengths of all hotspot regions, and the hotspot density is obtained by calculating the ratio of the hotspot shared length to the genome shared length. The sharing enhancement coefficient of the corresponding individual pair is obtained by multiplying the hotspot coverage rate and the hotspot density. Screen for family groups with continuous multi-generational kinship in genomic data and extract IBD shared fragments in the same hotspot region between adjacent generations; The ratio of the length of the IBD shared fragment between offspring and parents in the same hotspot region is calculated to obtain the intergenerational survival rate of the corresponding hotspot region; the genetic segregation decay coefficient of each pair of cultured individuals is calculated based on the intergenerational survival rate and the recombination rate data of the corresponding hotspot region. The genetic correlation coefficient of the corresponding individual pair is obtained by weighting and summing the sharing enhancement coefficient and the genetic segregation decay coefficient according to a preset ratio; The methods for constructing the weighted kinship matrix include: Construct a square matrix using the ID numbers of all cultured individuals as row and column indices, and fill the row and column intersection positions of the corresponding individual pairs with the genetic correlation coefficients of the square matrix to form a weighted kinship matrix.

6. The method for constructing and selecting Atlantic salmon kinship matrices based on multi-source data according to claim 5, characterized in that, The methods for performing feasible region feature analysis include: Identify the breeding index of each candidate individual in the breeding value priority matrix, set a lower limit for the breeding index, and select candidate individuals whose breeding index is higher than the lower limit to form a breeding candidate set; Extract the genetic correlation coefficients between any two candidate individuals in the breeding candidate set from the weighted kinship matrix, and set an upper limit threshold for the genetic correlation coefficients as a kinship constraint condition. Traverse all candidate pairs in the breeding candidate set and select candidate pairs that satisfy the kinship constraint as feasible pairs; Extract several pairs from the feasible pairs and combine them. For each pair combination, perform pair object swapping to output recombined pairs. Calculate the proportion of recombined pairs that satisfy the kinship constraint as the convexity retention rate. If the convexity retention rate is lower than the preset convexity threshold, the feasible region formed by the constraint is determined to be a non-convex feasible region. The proportion of feasible pairs formed by each candidate individual to the total number of possible pairs is used as the pairing feasibility rate. All candidate individuals are then ranked according to their pairing feasibility rates to form the analysis results.

7. The method for constructing and selecting Atlantic salmon kinship matrices based on multi-source data according to claim 6, characterized in that, The methods for generating candidate matching schemes include: The analysis results are divided into levels according to the preset stratification rules. Within each search level, optimization search is performed with the optimization objectives of maximizing the breeding index and minimizing the genetic correlation coefficient to obtain the level matching schemes for each search level. The level matching schemes of all search levels are integrated and the conditions are filtered to obtain candidate matching schemes.

8. The method for constructing and selecting Atlantic salmon kinship matrices based on multi-source data according to claim 7, characterized in that, The methods for time deviation identification and data updating include: The system acquires the latest monitoring timestamp in real time and calculates the time deviation value by comparing it with the construction time of the breeding value priority matrix. If the time deviation value exceeds the time range corresponding to the time mark, a matrix recalculation instruction is generated; otherwise, the breeding value priority matrix is ​​updated using the data corresponding to the latest monitoring timestamp.

9. The method for constructing and selecting Atlantic salmon kinship matrices based on multi-source data according to claim 8, characterized in that, The methods for constructing the association graph of each candidate individual include: Each candidate individual in the candidate mating scheme is used as a graph node. At the same time, the genetic correlation coefficient of any pair of individuals in the weighted kinship matrix is ​​identified. If the genetic correlation coefficient is higher than the preset association threshold, an association edge is established between the two graph nodes of the individual pair to form an association graph. The number of associated edges for each graph node is counted as the association density; the sum of the genetic correlation coefficients corresponding to all associated edges for each graph node is calculated as the association strength. Identify whether the candidate individuals corresponding to the two graph nodes connected by each associated edge belong to the same pairing in the candidate matching scheme. If they do, mark them as an intra-pair associated edge; otherwise, mark them as an inter-pair associated edge.

10. The method for constructing and selecting Atlantic salmon kinship matrices based on multi-source data according to claim 9, characterized in that, The methods for implementing individual adjustments include: Pairs with genetic diversity risk indices all above a preset risk threshold in candidate mating schemes are identified as pairs to be adjusted. Identify the graph nodes in the association graph that correspond to the pairings to be adjusted, and select candidate individuals corresponding to graph nodes with association strength higher than the selectable strength threshold as individuals to be replaced. The candidate individuals in the breeding candidate set whose breeding index difference with the individual to be replaced is less than the index difference threshold and whose association density in the association graph is lower than the preset density threshold constitute the substitute individual set. The number of pairing association edges between the graph node corresponding to each substitute individual in the statistical association graph and the graph nodes corresponding to other paired individuals in the candidate pairing scheme is used as the impact range value. The individual to be replaced is replaced by the substitute individual with the smallest impact range value to generate an adjustment pair. The adjustment pair replaces the pair to be adjusted and the candidate pairing scheme and association map are updated simultaneously. The adjustment process is repeated until the genetic diversity risk index of all pairs is lower than the preset risk threshold. All the finally updated candidate pairing schemes are then integrated to obtain the adjusted pairing scheme.