Auxiliary method and system for breeding lutjanus erythropterus based on multi-character collaborative selection

By employing a multi-trait synergistic selection method, the problem of genetic antagonism between traits in redfin snapper breeding was solved, resulting in improved breeding efficiency and genetic gain, and promoting the comprehensive genetic progress and diversity maintenance of redfin snapper populations.

CN120883931APending Publication Date: 2025-11-04SHENZHEN BASE OF SOUTH CHINA SEA FISHERIES RES INST CHINESE ACAD OF FISHERY SCI +2

Patent Information

Application Number
CN202511014309.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

In traditional redfin snapper breeding, single-trait directional selection leads to significant genetic antagonism between traits, limiting the overall rate of genetic progress and exacerbating the risk of population genetic diversity decline and inbreeding depression.

Method used

A multi-trait synergistic selection method was adopted. By collecting phenotypic data and constructing a standardized matrix through gene detection, genetic parameter estimation and breeding value prediction were performed. The trait associations were analyzed, dynamic weights were configured, the genomic selection index was calculated, a parental selection priority list was generated, and mating combination analysis was performed to improve breeding efficiency.

Benefits of technology

It improves the accuracy and balance of genetic gain in redfin snapper breeding, breaks through the limitations of traditional empirical matching, and improves breeding efficiency and environmental adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120883931A_ABST
    Figure CN120883931A_ABST
Patent Text Reader

Abstract

The invention discloses a lutjanus erythropterus breeding auxiliary method and system based on multi-character collaborative selection, and the method comprises the steps: carrying out phenotype data collection and gene detection, and constructing a standardized phenotype matrix and a genotype matrix; performing genetic parameter estimation and breeding value prediction based on the standardized phenotype matrix and the genotype matrix to obtain breeding value prediction information; the method comprises the following steps: analyzing an association relationship between characters of lutjanus rubripes to be bred, constructing a character association network, and carrying out multi-character dynamic weight configuration to obtain a dynamic weight set; performing genome selection index calculation and parent priority analysis according to the breeding value prediction information and the dynamic weight set to generate a parent selection priority list; and generating a candidate parent pool through the parent selection priority list, performing hybridization combination analysis, performing genetic gain prediction on the generated hybridization combination, and generating a hybridization auxiliary scheme. The genetic gain accuracy and balance of breeding of the lutjanus erythropterus are improved, the limitation of traditional experience selection and matching is broken through, and the breeding efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of red drum breeding, and particularly relates to a red drum breeding auxiliary method and system based on multi-trait coordinated selection. BACKGROUND

[0002] As an important marine economic fish, the genetic improvement of red drum has long relied on traditional breeding methods. Conventional breeding methods focus on single-trait directional selection, such as selecting individuals with fast growth rates for breeding. However, in actual production, it is found that continuous reinforcement of single-trait selection can lead to the decline of other important traits, such as the decline of disease resistance and the reduction of feed conversion efficiency in populations that have been excessively selected for growth performance. This genetic antagonism between traits is particularly pronounced in red drum, as growth rate and stress resistance exhibit negative genetic correlation, making it difficult to achieve comprehensive genetic progress through single-trait optimization strategies.

[0003] Traditional breeding strategies have long relied on single-trait directional selection, leading to the highlighting of genetic antagonism between key traits. This single-trait priority mode not only limits the rate of comprehensive genetic progress, but also exacerbates the decline of population genetic diversity by ignoring trait coordination, and there is also a risk of inducing inbreeding depression. Therefore, a red drum breeding auxiliary method and system based on multi-trait coordinated selection is proposed to improve the environmental adaptability and breeding efficiency of red drum breeding. SUMMARY

[0004] The present application overcomes the defects of the prior art and provides a red drum breeding auxiliary method and system based on multi-trait coordinated selection.

[0005] To achieve the above-mentioned purpose, the first aspect of the present application provides a red drum breeding auxiliary method based on multi-trait coordinated selection, comprising:

[0006] Phenotype data collection and gene detection are performed on the red drum population to be bred, and a standardized phenotype matrix and a genotype matrix are constructed based on the collected data;

[0007] Genetic parameter estimation and breeding value prediction are performed based on the standardized phenotype matrix and the genotype matrix, and breeding value prediction information is obtained;

[0008] Based on the breeding value prediction information, the correlation between the traits of the red drum to be bred is analyzed and a trait correlation network is constructed, and multi-trait dynamic weight configuration is performed to obtain a dynamic weight set;

[0009] Genomic selection index calculation and parent priority analysis are performed based on the breeding value prediction information and the dynamic weight set, and a parent selection priority list is generated;

[0010] The candidate parent pool is generated by the parent selection priority list, the mating combination analysis is performed through the candidate parent pool, and the genetic gain prediction is performed on the generated mating combination to generate a mating assistance scheme.

[0011] In the scheme, the phenotype data of the to-be-bred red sea bream population is collected and the gene detection is performed, and the standardized phenotype matrix and genotype matrix are constructed according to the collected data, specifically including:

[0012] The fish body weight of the to-be-bred red sea bream population in the target breeding pool is weighed by using a weighing platform in a preset collection period, and the body length growth rate of the to-be-bred red sea bream population in the preset collection period is calculated by arranging a multi-source sensor array and combining machine vision technology;

[0013] The amount of residual feed in the target breeding pool is obtained by arranging a residual feed collection device, and the amount of residual feed is calculated by subtracting the daily feed amount to obtain the feed intake data, and the feed intake-weight gain ratio is calculated by using the obtained fish body weight data;

[0014] Based on the preset disease-resistant breeding test scheme, the red sea bream in the target breeding pool is subjected to disease-resistant test, and the survival rate in the preset time period is recorded as the disease-resistant trait quantitative value, and finally the original phenotype data set is obtained combined with the fish body identity mark;

[0015] Subsequently, the gene sample of the to-be-bred red sea bream population is collected, and the whole genome typing is performed on the collected gene sample by using the gene detection technology to form the original genotype data set containing the site genotype and physical position;

[0016] According to the original phenotype data set, the minimum and maximum observed values of the current population are calculated for each trait column, and then the original data is mapped and processed by using a linear transformation algorithm to convert the standardization phenotype matrix represented by rows and standardization trait values represented by columns;

[0017] The original gene data in the original genotype data set is subjected to individual quality control and site filtering to obtain qualified sites, and the qualified sites are converted into a two-dimensional numerical matrix by using a genotype encoding rule and adding individual ID and site coordinate index to generate a genotype matrix.

[0018] In the scheme, the genetic parameter estimation and breeding value prediction are performed based on the standardized phenotype matrix and genotype matrix to obtain breeding value prediction information, specifically including:

[0019] The standardized phenotype matrix and genotype matrix are obtained, the weighted cosine similarity algorithm is used to calculate the allele similarity of the whole genome SNP site between individuals according to the genotype matrix, and the genomic relationship matrix is generated;

[0020] The standardized phenotype matrix is taken as a response variable, an individual additive genetic effect is set as a random effect, a multi-trait animal model is constructed in combination with the genomic relationship matrix, and a restricted maximum likelihood method is used for iterative solution;

[0021] The genetic variance and environmental variance parameters are initialized, the variance component estimate values are updated in cycles through an expectation maximization algorithm, the conditional expectation of the observation data is calculated in each iteration, and the likelihood function is maximized until the convergence threshold value, and finally the genetic parameter estimate information is output;

[0022] The variance component estimate values are obtained, the multi-trait best linear unbiased prediction algorithm is introduced, the genomic relationship matrix is taken as the covariance prior of the random effect, the genomic estimated breeding value of each individual for each trait is calculated, and breeding value prediction information is generated.

[0023] In the scheme, the breeding value prediction information is analyzed to analyze the correlation between the traits of the to-be-bred red sea bream and construct a trait correlation network, and multi-trait dynamic weight configuration is performed to obtain a dynamic weight set, specifically including:

[0024] The breeding value prediction information and breeding environment information are obtained, the maximum information coefficient algorithm is introduced, the breeding value prediction information is taken as input to scan all trait combinations, the breeding value distribution interval of each trait is traversed in the form of a sliding window, the nonlinear correlation strength of different trait combinations at the genetic level is calculated, and a trait combination correlation strength matrix is generated;

[0025] Subsequently, the obtained breeding environment information is discretized into several environmental intervals according to a preset division rule, independent repeated correlation strength calculation is performed in each environmental interval, and an environment-specific matrix is generated according to the calculated correlation strength values;

[0026] The trait combination correlation strength matrix and the environment-specific matrix are associated, a directed weighted network is constructed based on the matrix association result, taking traits as nodes, the correlation strength between trait combinations as correlation weights, and the environment-specific matrix corresponding to the traits as an attached feature;

[0027] Based on the directed weighted network, node centrality parameters and environmental sensitivity coefficients are extracted and input into a dynamic weight allocator to analyze the contribution weight of each trait to the comprehensive breeding value;

[0028] The SHAP model and network constraint rules are integrated in the dynamic weight allocator, the conditional expectation Shapley value is calculated through individual whole-genome markers and environmental parameters, and the marginal contribution degree of each trait is obtained using the calculated Shapley value;

[0029] The dynamic weights of multiple traits are configured according to the marginal contribution of each trait to obtain the initial dynamic weights of multiple traits. The initial dynamic weights of multiple traits are then corrected by the network constraint rules to obtain the final dynamic weight set.

[0030] In this scheme, the step of calculating the genomic selection index and analyzing parental priority based on breeding value prediction information and dynamic weight set to generate a parental selection priority list specifically includes:

[0031] Obtain breeding value prediction information and dynamic weight set, generate a genome breeding value matrix based on the breeding value prediction information, perform standardization transformation on the genome breeding value matrix, and calculate the population mean and standard deviation of each trait column by column;

[0032] For each individual trait value, subtract the mean of the corresponding trait and divide by the standard deviation to generate a standardized breeding value matrix. The standardized breeding matrix is ​​then linearly weighted with the dynamic weight set to generate an initial genomic selection index.

[0033] Obtain the genotype matrix, extract the genotype codes of the target individuals in the preset QTL regions through the pre-trained QTL effect value model, multiply the genotype codes of each locus by the corresponding standardized effect value and then sum them to generate the QTL additive genetic contribution value.

[0034] The QTL additive genetic contribution value and the initial genomic selection index are arithmetically summed and fused to obtain the final genomic selection index. The parental priority of each candidate individual is generated based on the final genomic selection index.

[0035] A parent selection priority list is generated based on the parental priority corresponding to each candidate individual, sorted in descending order.

[0036] In this scheme, the process of generating a candidate parent pool through the parent selection priority list, performing mating combination analysis through the candidate parent pool, predicting genetic gain of the generated mating combinations, and generating a mating assistance scheme specifically includes:

[0037] Obtain a parent selection priority list, select several candidate parent individuals within a preset range according to the parent selection priority list to generate a candidate parent pool, extract a subset of the genotype matrix, calculate the kinship coefficient through the genome relationship matrix, and generate a kinship relationship edge weight matrix;

[0038] Based on the kinship edge weight matrix, a graph theory network is constructed with candidate parent individuals as nodes and kinship coefficients as edge weights. The graph theory network is then used as the input of the minimum spanning tree algorithm for breeding combination analysis.

[0039] Initialize an empty tree structure, starting from the node with the highest genome selection index, select the valid edge with the smallest coefficient of kinship connected to the current tree each time, limit the number of connections of each node by a preset pairing threshold, and calculate the inbreeding coefficient of the population after adding a new edge for inbreeding control;

[0040] When all nodes are connected or there is no legal edge to be added, terminate the iteration, output the minimum kinship cost pairing tree, generate a pairing list according to the minimum kinship cost pairing tree, extract the standardized genome estimated breeding value of each trait of the parents for each pairing combination in the pairing list, and simulate the offspring genotype distribution by Mendelian inheritance law;

[0041] The simulation result of the offspring genotype distribution is input into a QTL effect model to calculate the expected breeding value of the offspring for genetic gain prediction, and finally the pairing combination within the preset genetic gain range is selected to generate a pairing assistance scheme for recommendation.

[0042] The second aspect of the application provides a red drum breeding assistance system based on multi-trait collaborative selection, which comprises a memory and a processor, wherein the memory comprises a red drum breeding assistance method program based on multi-trait collaborative selection, and the red drum breeding assistance method program based on multi-trait collaborative selection is executed by the processor to realize the following steps:

[0043] Phenotype data collection and gene detection are performed on the red drum population to be bred, and a standardized phenotype matrix and a genotype matrix are constructed according to the collected data;

[0044] Genetic parameter estimation and breeding value prediction are performed based on the standardized phenotype matrix and the genotype matrix to obtain breeding value prediction information;

[0045] Based on the breeding value prediction information, the correlation between the traits of the red drum to be bred is analyzed and a trait correlation network is constructed, and multi-trait dynamic weight configuration is performed to obtain a dynamic weight set;

[0046] Genome selection index calculation and parent priority analysis are performed according to the breeding value prediction information and the dynamic weight set to generate a parent selection priority list;

[0047] A candidate parent pool is generated through the parent selection priority list, pairing combination analysis is performed through the candidate parent pool, and genetic gain prediction is performed on the generated pairing combination to generate a pairing assistance scheme.

[0048] The application discloses a red snapper breeding auxiliary method and system based on multi-trait collaborative selection, which comprises the following steps: collecting phenotype data and performing gene detection, and constructing a standardized phenotype matrix and a genotype matrix; performing genetic parameter estimation and breeding value prediction based on the standardized phenotype matrix and the genotype matrix to obtain breeding value prediction information; analyzing the correlation between the traits of the red snapper to be bred and constructing a trait correlation network, performing multi-trait dynamic weight configuration, and obtaining a dynamic weight set; performing genome selection index calculation and parent priority analysis according to the breeding value prediction information and the dynamic weight set, and generating a parent selection priority list; generating a candidate parent pool through the parent selection priority list, performing mating combination analysis, and performing genetic gain prediction on the generated mating combination to generate a mating auxiliary scheme. The genetic gain precision and balance of the red snapper breeding are improved, the limitations of traditional experience matching are broken through, and the breeding efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments or examples of the present application, the drawings needed to be used in the embodiments or examples will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0050] Figure 1 A flow chart of a red snapper breeding auxiliary method based on multi-trait collaborative selection is provided for an embodiment of the present application.

[0051] Figure 2 A block diagram of a red snapper breeding auxiliary system based on multi-trait collaborative selection is provided for an embodiment of the present application.

[0052] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0053] In order to more clearly illustrate the technical solutions in the embodiments or examples of the present application, the drawings needed to be used in the embodiments or examples will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0054] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, however, the present application can also be implemented in other ways different from those described herein, therefore, the protection scope of the present application is not limited by the specific embodiments disclosed below.

[0055] Figure 1 A flow chart of a red snapper breeding auxiliary method based on multi-trait collaborative selection is provided for an embodiment of the present application.

[0056] As Figure 1 shown, the present application provides a multi-trait collaborative selection-based red drum breeding assistance method flowchart, comprising:

[0057] S102, phenotype data collection and gene detection are performed on the red drum population to be bred, and a standardized phenotype matrix and a genotype matrix are constructed according to the collected data;

[0058] S104, genetic parameter estimation and breeding value prediction are performed based on the standardized phenotype matrix and the genotype matrix, and breeding value prediction information is obtained;

[0059] S106, based on the breeding value prediction information, the correlation between the traits of the red drum to be bred is analyzed and a trait correlation network is constructed, and multi-trait dynamic weight configuration is performed, and a dynamic weight set is obtained;

[0060] S108, according to the breeding value prediction information and the dynamic weight set, a genomic selection index is calculated and parent priority is analyzed, and a parent selection priority list is generated;

[0061] S110, a candidate parent pool is generated through the parent selection priority list, mating combination analysis is performed through the candidate parent pool, and genetic gain prediction is performed on the generated mating combination, and a mating assistance scheme is generated.

[0062] Further, in a preferred embodiment of the present application, the phenotype data collection and gene detection are performed on the red drum population to be bred, and the standardized phenotype matrix and the genotype matrix are constructed according to the collected data, specifically comprising:

[0063] The target breeding pool is weighed by a weighing platform within a preset collection period, and the body length growth rate of the red drum population to be bred is calculated by arranging a multi-source sensor array and combining machine vision technology;

[0064] The amount of residual feed in the target breeding pool is obtained by arranging a residual feed collection device, and the amount of food intake is calculated by difference calculation with the daily feed amount, and the food intake-weight gain ratio is calculated using the obtained fish weight data;

[0065] Based on a preset disease-resistant breeding test scheme, the red drum in the target breeding pool is subjected to disease-resistant test, and the survival rate within a preset time period is recorded as a disease-resistant trait quantitative value, and finally the original phenotype data set is obtained in combination with the fish identity identification;

[0066] Subsequently, the red drum population to be bred is collected, and whole genome typing is performed on the collected gene samples by gene detection technology to form an original genotype data set containing locus genotype and physical position.

[0067] According to the original phenotype data set, the minimum and maximum observed values of each trait column of the current population are independently calculated, and the original data is mapped and processed by a linear transformation algorithm to convert the standard phenotype matrix represented by rows and columns into individual and standardized trait values.

[0068] The original gene data in the original genotype data set is subjected to individual quality control and site filtering to obtain qualified sites, and the qualified sites are converted into a two-dimensional numerical matrix by genotype encoding rules and individual ID and site coordinate index are added to generate a genotype matrix.

[0069] It should be noted that the integrated dynamic weighing platform is used to weigh the red drum population in the target breeding pool at a preset collection period (such as every 7 days), and the underwater multi-source sensor array (including a depth camera and a laser scanner) is used to track the individual length change trajectory through machine vision algorithm, and the length growth rate of each fish in the period is calculated based on time series analysis. The residual bait collection device is used to recover residual bait through the bottom siphon pipe every day, and the accurate food intake is obtained by subtracting the residual bait weight from the recorded bait amount of the automatic feeding system after drying and weighing, and then the feed conversion efficiency (food intake-gain weight ratio) is calculated. The resistance to disease is quantified according to the standard pathogen challenge test procedure: randomly divided groups are tested, and the survival state is recorded within 96 hours of observation, and the individual survival rate is calculated as the resistance phenotype value. After associating the phenotype data with the fish identity (such as PIT tag code), the original phenotype data set is formed, covering the original observation values of three key traits of growth, feed efficiency and disease resistance. At the same time, fin samples of the red drum population are collected for whole genome SNP typing, and an original genotype data set containing several site genotypes (AA / AB / BB) and physical location information is generated. Subsequently, individual-level quality control is performed to remove invalid individuals with a detection rate of less than 95%; simultaneously, site-level filtering is performed to remove low-quality sites with a detection rate of less than 98% or a minor allele frequency of less than 0.05; finally, the qualified sites are uniformly represented as 0 / 1 / 2 numeric matrix (AA=0, AB=1, BB=2) through encoding conversion, and the structured genotype matrix is formed by adding individual ID index and SNP physical location annotation. Finally, the original phenotype data set is independently processed for each trait column: the minimum and maximum observation values of the length growth rate, the feed conversion rate and the disease survival rate are extracted respectively, and the original data is mapped to the [0, 1] interval through linear transformation formula, for example, the normalized value of the growth trait = (actual gain weight-minimum gain weight) / (maximum gain weight-minimum gain weight), the generated standardized phenotype matrix, the row corresponds to the individual ID, and the column bears the trait value after dimensionless processing, provides unbiased input for subsequent genetic evaluation. The genotype matrix retains the effective site information of genetic variation through quality control filtering and encoding conversion, and shares the individual ID index system with the phenotype matrix, and together forms the standardized data foundation for breeding analysis.

[0070] Further, in a preferred embodiment of the present application, the genetic parameter estimation and breeding value prediction based on the standardized phenotype matrix and genotype matrix are performed to obtain breeding value prediction information, specifically comprising:

[0071] Obtaining a standardized phenotype matrix and a genotype matrix, calculating the allele similarity of whole genome SNP sites between individuals using a weighted cosine similarity algorithm based on the genotype matrix, and generating a genomic relationship matrix;

[0072] The standardized phenotype matrix is taken as a response variable, an individual additive genetic effect is set as a random effect, a multi-trait animal model is constructed in combination with the genomic relationship matrix, and a restricted maximum likelihood method is used for iterative solution;

[0073] The genetic variance and environmental variance parameters are initialized, the variance component estimates are cyclically updated through an expectation maximization algorithm, the conditional expectation of the observation data is calculated in each round of iteration, and the likelihood function is maximized until the convergence threshold value, and finally the genetic parameter estimation information is output;

[0074] The variance component estimates are obtained, a multi-trait best linear unbiased prediction algorithm is introduced, the genomic relationship matrix is taken as the covariance prior of the random effect, the genomic estimated breeding value of each individual for each trait is calculated, and breeding value prediction information is generated.

[0075] It should be noted that firstly, the weighted cosine similarity algorithm is used to calculate the whole genome genetic similarity between individuals. By applying weights to each SNP site (information weight is calculated based on allele frequency), the matching degree of alleles of individuals at tens of thousands of sites is quantified through the cosine similarity formula, and finally a symmetric genomic relationship matrix (GRM) is generated, and the element value represents the molecular kinship intensity between individuals (range -1 to 1, positive value indicates genetic similarity). Then, a multi-trait animal model is constructed for genetic parameter estimation: taking the standardized phenotype matrix as the response variable, setting the individual additive genetic effect as the random effect, and defining the covariance structure by the genomic relationship matrix. The restricted maximum likelihood method is used to start the iterative solution: the genetic variance and environmental variance components are initialized, the parameters are cyclically updated through the expectation maximization algorithm, the conditional expectation of the observation data is calculated in each round of iteration, the restricted likelihood function is maximized, and the convergence is determined when the change amplitude of the genetic variance estimate is less than the preset threshold value, and the key genetic parameter information is output, including the genetic correlation coefficient between traits, the variance component estimates and the genetic correlation coefficient between traits. Based on the converged variance component results, a multi-trait best linear unbiased prediction algorithm is introduced to calculate the genomic breeding value. The genomic relationship matrix is taken as the covariance prior of the random effect, and a mixed model equation system is constructed: the fixed effect part contains the environmental factor design matrix, and the random effect part constrains the covariance structure by the GRM. The sparse matrix solver is used to calculate the genomic estimated breeding value (GEBV) of each individual for each trait, and finally the breeding value prediction information forms a structured data table (row = individual ID, column = three-trait GEBV value), which also carries the precision index of the genetic parameter estimation, providing a quantitative genetic potential basis for subsequent selection decisions.

[0076] Further, in a preferred embodiment of the present application, the breeding value prediction information is used to analyze the correlation between the traits of the to-be-bred red sea bream and construct a trait correlation network, and a multi-trait dynamic weight configuration is performed to obtain a dynamic weight set, specifically including:

[0077] The breeding value prediction information and the breeding environment information are obtained, the maximum information coefficient algorithm is introduced, the breeding value prediction information is taken as input to scan all trait combinations, the breeding value distribution interval of each trait is traversed in the form of a sliding window, the nonlinear correlation strength of different trait combinations at the genetic level is calculated, and a trait combination correlation strength matrix is generated;

[0078] Subsequently, the obtained breeding environment information is discretized into a plurality of environment intervals according to a preset division rule, independent repeated correlation strength calculation is performed in each environment interval, and an environment-specific matrix is generated according to the calculated correlation strength values;

[0079] The trait combination correlation strength matrix and the environment-specific matrix are associated, a directed weighted network is constructed with traits as nodes, the correlation strength between trait combinations as correlation weights, and the environment-specific matrix corresponding to the traits as an attached feature based on the matrix association result;

[0080] Based on the directed weighted network, a node centrality parameter and an environment sensitivity coefficient are extracted and input into a dynamic weight allocator to analyze the contribution weight of each trait to the comprehensive breeding value;

[0081] The SHAP model and the network constraint rule are integrated in the dynamic weight allocator, the conditional expected Shapley value is calculated based on the individual whole genome marker and the environment parameter, and the marginal contribution degree of each trait is obtained by using the calculated Shapley value;

[0082] The marginal contribution degrees corresponding to the traits are used to configure the dynamic weights of multiple traits, an initial dynamic weight of multiple traits is obtained, the initial dynamic weight of multiple traits is modified by using the network constraint rule, and finally a dynamic weight set is obtained.

[0083] It should be noted that the genomic breeding value (GEBV) prediction information and the timing data of the environmental parameters (water temperature, salinity, dissolved oxygen) are obtained, and a maximum information coefficient (MIC) algorithm is introduced for multi-dimensional correlation analysis. Taking the GEBV matrix as input, all trait combinations are scanned through a sliding window mechanism, and the non-linear correlation strength is calculated at the genetic level: the window slides along the trait value distribution interval, the mutual information entropy of each trait pair in the sub-box is counted, and the MIC value (0-1 interval, the larger the value, the stronger the correlation) is output after normalization, and a symmetric trait combination correlation strength matrix is generated. Process the environmental data: discretize the continuous environmental parameters into mutually exclusive intervals according to the preset threshold. Independently repeat the MIC calculation in each environmental interval subset to generate an environment-specific correlation matrix. Align the genetic level correlation matrix and the environmental matrix in time and space, and construct an undirected weighted network based on the alignment result, with traits as nodes (node size proportional to cross-environmental correlation stability), MIC value as edge weight, and environment-specific matrix corresponding to the trait as an attached feature. Subsequently, key parameters are extracted using the undirected weighted network: calculate the centrality through the node weighted degree, and fit the sensitivity coefficient through the edge weight environmental gradient. Input the centrality and sensitivity coefficient into the dynamic weight distributor, which integrates a dual-channel processing engine: SHAP analysis channel: taking individual whole-genome SNP data and real-time environmental parameters as input, the marginal contribution of each trait is calculated through conditional expectation Shapley value. Network constraint channel: apply the centrality to set the weight lower limit (centrality>0.8 forces weight≥0.3), and dynamically modulate the contribution degree using the sensitivity coefficient (e.g. disease resistance contribution degree x 1.5 under high temperature). The initial multi-trait dynamic weight is corrected through the network constraint rule. Finally, a dynamic weight set of environmental index is generated.

[0084] Further, in a preferred embodiment of the present application, the genomic selection index calculation and parent priority analysis according to the breeding value prediction information and the dynamic weight set are performed to generate a parent selection priority list, specifically comprising:

[0085] The breeding value prediction information and the dynamic weight set are obtained, a genomic breeding value matrix is generated according to the breeding value prediction information, the genomic breeding value matrix is standardized and converted, and the population mean and standard deviation of each trait are calculated column by column;

[0086] The standardized breeding value matrix is generated by subtracting the trait value of each individual from the corresponding trait mean and dividing by the standard deviation, the standardized breeding matrix is linearly weighted with the dynamic weight set, and the initial genomic selection index is generated;

[0087] The genotype matrix is obtained, the genotype code of the target individual in the preset QTL region is extracted through the pre-trained QTL effect value model, the genotype code of each locus is multiplied by the corresponding standardized effect value, and then accumulated to generate the QTL additive genetic contribution value;

[0088] The QTL additive genetic contribution value is arithmetically added to the initial genome selection index to obtain a final genome selection index, and the parent priority of each candidate individual is generated based on the final genome selection index;

[0089] The parent priority list is generated in descending order based on the parent priority corresponding to each candidate individual.

[0090] It should be noted that a structured genome breeding value matrix is first constructed: the rows correspond to individual IDs, and the columns store the original breeding values of the growth rate, disease resistance, and feed efficiency of three traits, respectively. Standardization conversion is performed on the matrix, and the population mean and standard deviation of each trait are calculated column by column. The trait value of each individual is subtracted from the corresponding mean and then divided by the standard deviation to generate a Z-score standardized matrix. Thus, the dimensional differences between traits are eliminated, ensuring the fairness of index synthesis. Next, the standardized breeding value matrix is linearly weighted and fused with the dynamic weight set: the three trait standardized values of each individual are multiplied by the dynamic weight coefficient under the current environment, and the product is accumulated to generate an initial genome selection index, forming an initial index list carrying individual ID and environment label, representing the comprehensive genetic potential under dynamic weight adjustment. The pre-trained QTL effect value model is called simultaneously, and the genotype code of the target individual in the preset core QTL region is extracted from the genotype matrix. The effect value conversion is performed on each site, and the genotype code is multiplied by the corresponding standardized effect value to accumulate and generate the QTL additive genetic contribution value. Then, the QTL additive genetic contribution value is arithmetically added to the initial genome selection index to generate the final genome selection index. The parent priority of the candidate individual is generated based on the index value (the larger the value, the higher the priority), and the parent priority list is output in descending order.

[0091] Further, in a preferred embodiment of the present application, the candidate parent pool is generated through the parent selection priority list, the mating combination analysis is performed through the candidate parent pool, the genetic gain prediction is performed on the generated mating combination, and the mating assistance scheme is generated, specifically including:

[0092] The parent selection priority list is obtained, and a number of candidate parent individuals within a preset range interval are selected from the parent selection priority list to generate a candidate parent pool and extract a genotype matrix subset. The kinship coefficient is calculated through the genome relationship matrix, and a kinship edge weight matrix is generated;

[0093] The kinship edge weight matrix is used to construct a graph theory network with the candidate parent individuals as nodes and the kinship coefficient as edge weight, and the graph theory network is used as the input of the minimum spanning tree algorithm for mating combination analysis;

[0094] initializing an empty tree structure, traversing from the node with the highest index in the genome, selecting each time the valid edge with the smallest coefficient of relationship connecting to the current tree, limiting the number of connections for each node by a preset mating threshold, and calculating the inbreeding coefficient of the population after adding a new edge for inbreeding control;

[0095] terminating the iteration when all nodes are connected or there is no legal edge that can be added, outputting the minimum cost mating tree, generating a mating list according to the minimum cost mating tree, extracting the standardized genomic estimated breeding value of each trait of the parents for each mating combination in the mating list, and simulating the genotype distribution of offspring by Mendelian inheritance law;

[0096] inputting the simulation result of the genotype distribution of the offspring into a QTL effect model to calculate the expected breeding value of the offspring for genetic gain prediction, and finally selecting a mating combination within a preset genetic gain range to generate a mating assistance scheme for recommendation.

[0097] It should be noted that the parent selection priority list is obtained, the candidate parent pool is formed by screening the candidate individuals ranked within a preset range (such as the top 20%), and the genotype matrix subset of the corresponding individual is synchronously extracted. The coefficient of kinship between two individuals is calculated through the genomic relationship matrix (the calculation formula is the standardized value of the shared allele frequency), and a symmetric kinship edge weight matrix is generated (the matrix element value ranges from 0 to 1, and the smaller the value, the farther the kinship). Subsequently, a graph theory network is constructed based on the edge weight matrix: the candidate parent individuals are nodes (the node size maps the high and low of the genomic selection index), and the coefficient of kinship is the edge weight (only the edges with a coefficient of less than 0.3 are retained to avoid the risk of close relatives). The network is input into the minimum spanning tree algorithm: the empty tree structure is initialized, and the node with the highest GSI is selected as the starting root node, and the effective edge with the smallest coefficient of kinship is preferentially selected when the tree is iteratively expanded (that is, the farthest legal pairing). Each time a new edge is added, the number of connections of each node is limited by a preset pairing threshold (such as ≤2 times), the average inbreeding coefficient ΔF of the population after the new edge is added is calculated in real time (the calculation formula is ΔF = 0.5 × the coefficient of the new edge), and it is ensured that ΔF is always <0.01. When all nodes are connected or there is no legal edge that can be added, the iteration is terminated, and the minimum kinship cost mating tree is output (the tree edge represents the optimal mating combination). Then, the mating list is generated according to the mating tree structure, for each mating combination in the list, the standardized genomic estimated breeding value (growth / disease resistance / forage efficiency three-trait Z-score) of the parents is extracted, the genotype distribution of the offspring is simulated through Mendelian inheritance law: for each SNP site, independent allele random separation calculation is performed (for example, when the parent is AB type and the mother is AA type, the genotype probability of the offspring is 50% AA + 50% AB). The genotype probability distribution of the offspring is input into the pre-trained QTL effect model to calculate the expected breeding value of the offspring. Specifically, for each site in the target QTL interval, multiply the genotype probability of the offspring by the standardized effect value and accumulate to generate the predicted value of the three-trait breeding value. The genetic gain is calculated based on the difference between the average breeding value of the offspring and the parent population (ΔG = offspring mean - parent mean), and finally the mating combination with ΔG exceeding the preset threshold is screened to generate the mating assistance scheme (including the recommended mating pair list, genetic gain heat map, and inbreeding risk rating) to guide the fish mating scheduling and family establishment plan, and realize the precise control of genetic progress.

[0098] Figure 2 A red snapper breeding assistance system 2 based on multi-trait collaborative selection is provided for an embodiment of the present application, which comprises a memory 21 and a processor 22. The memory 21 contains a red snapper breeding assistance method program based on multi-trait collaborative selection. When the red snapper breeding assistance method program based on multi-trait collaborative selection is executed by the processor 22, the following steps are implemented:

[0099] Phenotype data of the to-be-bred red sea bream population is collected and gene detection is performed, and a standardized phenotype matrix and a genotype matrix are constructed according to the collected data;

[0100] Genetic parameter estimation and breeding value prediction are performed based on the standardized phenotype matrix and the genotype matrix, to obtain breeding value prediction information;

[0101] Based on the breeding value prediction information, a correlation between traits of the to-be-bred red sea bream is analyzed, a trait correlation network is constructed, and multi-trait dynamic weight configuration is performed, to obtain a dynamic weight set;

[0102] Genomic selection index calculation and parent priority analysis are performed according to the breeding value prediction information and the dynamic weight set, to generate a parent selection priority list;

[0103] A candidate parent pool is generated through the parent selection priority list, cross combination analysis is performed through the candidate parent pool, and genetic gain prediction is performed on the generated cross combination, to generate a cross assistance scheme.

[0104] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division mode, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed components can be through some interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0105] The units described above as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units; they can be located in one place or distributed on multiple network units; part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0106] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be a unit alone, or two or more units can be integrated in one unit; the integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.

[0107] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware, the foregoing program can be stored in a computer readable storage medium, and the program executes the steps of the method embodiments when executed; and the foregoing storage medium includes a mobile storage device, a read-only memory (ROM), a random access memory (RAM), a magnetic disc or an optical disc, and various storage medium capable of storing program codes.

[0108] Alternatively, the integrated unit of the present application can be stored in a computer readable storage medium if it is realized in the form of a software function module and sold or used as an independent product. Based on such understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes a mobile storage device, a ROM, a RAM, a magnetic disc or an optical disc, and various storage medium capable of storing program codes.

[0109] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A breeding assistance method for redfin snapper based on multi-trait synergistic selection, characterized in that, include: Phenotypic data and gene testing were conducted on the redfin snapper population to be bred, and a standardized phenotypic matrix and genotype matrix were constructed based on the collected data. Genetic parameters are estimated and breeding values ​​are predicted based on the standardized phenotypic matrix and genotype matrix to obtain breeding value prediction information. Based on the breeding value prediction information, the correlation between traits of the redfin snapper to be bred is analyzed and a trait association network is constructed. Then, the dynamic weights of multiple traits are configured to obtain a dynamic weight set. Based on the breeding value prediction information and dynamic weight set, the genomic selection index is calculated and the parental priority is analyzed to generate a parental selection priority list. A candidate parent pool is generated by the parent selection priority list, and mating combination analysis is performed on the candidate parent pool. Genetic gain prediction is then performed on the generated mating combinations to generate mating assistance schemes.

2. The breeding assistance method for redfin snapper based on multi-trait synergistic selection according to claim 1, characterized in that, The process involves collecting phenotypic data and performing gene testing on the redfin snapper population to be bred, and constructing standardized phenotypic and genotypic matrices based on the collected data. Specifically, this includes: The weighing platform was used to weigh the red snapper population to be bred in the target breeding pond within a preset collection period. The body length growth rate of the red snapper population to be bred within the preset collection period was calculated by deploying a multi-source sensor array and combining machine vision technology. The amount of uneaten feed in the target breeding pond is obtained by deploying uneaten feed collection devices. The difference between the uneaten feed and the daily feed amount is calculated to obtain the feed intake data. The feed intake-weight gain ratio is calculated using the obtained fish body weight data. Based on a pre-set disease resistance breeding test plan, a disease resistance test was conducted on redfin snapper in the target breeding pond. The survival rate within the pre-set time period was recorded as the quantitative value of the disease resistance trait. Finally, the original phenotypic dataset was obtained by combining the fish body identification. Subsequently, gene samples were collected from the redfin snapper population to be bred. Based on the collected gene samples, whole-genome typing was performed using gene detection technology to form an original genotype dataset containing locus genotypes and physical locations. Based on the original phenotypic dataset, the minimum and maximum observed values ​​of the current population are independently calculated for each trait column. Then, the original data is mapped and transformed into a standardized phenotypic matrix with rows representing individuals and columns representing standardized trait values ​​through a linear transformation algorithm. Individual quality control and site filtering are performed on the original gene data in the original genotype dataset to obtain qualified sites. The qualified sites are converted into a two-dimensional numerical matrix according to the genotype encoding rules and individual IDs and site coordinate indices are added to generate the genotype matrix.

3. The breeding assistance method for redfin snapper based on multi-trait synergistic selection according to claim 1, characterized in that, The process of estimating genetic parameters and predicting breeding values ​​based on the standardized phenotypic matrix and genotype matrix to obtain breeding value prediction information specifically includes: Obtain standardized phenotypic and genotype matrices. Based on the genotype matrix, use a weighted cosine similarity algorithm to calculate the allele similarity of whole-genome SNP sites among individuals, and generate a genome relationship matrix. The standardized phenotypic matrix was used as the response variable, and the individual additive genetic effect was set as a random effect. A multi-trait animal model was constructed by combining the genomic relationship matrix, and the restricted maximum likelihood method was used for iterative solution. Initialize the genetic and environmental variance parameters, update the variance component estimates iteratively using the expectation-maximization algorithm, calculate the conditional expectation of the observed data in each iteration, and maximize the likelihood function until the convergence threshold is reached, finally outputting the genetic parameter estimates. By obtaining variance component estimates, a multi-trait optimal linear unbiased prediction algorithm is introduced. The genomic relationship matrix is ​​used as the covariance prior of random effects to calculate the genomic estimated breeding value of each trait for each individual, and breeding value prediction information is generated.

4. The breeding assistance method for redfin snapper based on multi-trait synergistic selection according to claim 1, characterized in that, The process involves analyzing the correlation between traits of the redfin snapper to be bred based on the breeding value prediction information, constructing a trait association network, and configuring dynamic weights for multiple traits to obtain a dynamic weight set. Specifically, this includes: The breeding value prediction information and breeding environment information are obtained. The maximum information coefficient algorithm is introduced to scan all trait combinations with the breeding value prediction information as input. The breeding value distribution interval of each trait is traversed in the form of a sliding window. The nonlinear association strength of different trait combinations at the genetic level is calculated, and the trait combination association strength matrix is ​​generated. Subsequently, the acquired breeding environment information is discretized into several environmental intervals according to a preset division rule. The correlation strength is calculated independently and repeatedly in each environmental interval, and an environment-specific matrix is ​​generated based on the calculated correlation strength values. The trait combination association strength matrix is ​​associated with the environment specificity matrix. Based on the matrix association result, an undirected weighted network is constructed with traits as nodes, the association strength between trait combinations as association weights, and the environment specificity matrix corresponding to the traits as auxiliary features. Based on the undirected weighted network, the node centrality parameter and environmental sensitivity coefficient are extracted and input into the dynamic weight allocator to analyze the contribution weight of each trait to the comprehensive breeding value. The dynamic weight allocator integrates the SHAP model and network constraint rules, calculates the conditional expectation Shapley value through individual whole genome markers and environmental parameters, and uses the calculated Shapley value to obtain the marginal contribution of each trait. The dynamic weights of multiple traits are configured according to the marginal contribution of each trait to obtain the initial dynamic weights of multiple traits. The initial dynamic weights of multiple traits are then corrected by the network constraint rules to obtain the final dynamic weight set.

5. The breeding assistance method for redfin snapper based on multi-trait synergistic selection according to claim 1, characterized in that, The step of calculating the genomic selection index and analyzing parental priority based on breeding value prediction information and dynamic weight set to generate a parental selection priority list specifically includes: Obtain breeding value prediction information and dynamic weight set, generate a genome breeding value matrix based on the breeding value prediction information, perform standardization transformation on the genome breeding value matrix, and calculate the population mean and standard deviation of each trait column by column; For each individual trait value, subtract the mean of the corresponding trait and divide by the standard deviation to generate a standardized breeding value matrix. The standardized breeding matrix is ​​then linearly weighted with the dynamic weight set to generate an initial genomic selection index. Obtain the genotype matrix, extract the genotype codes of the target individuals in the preset QTL regions through the pre-trained QTL effect value model, multiply the genotype codes of each locus by the corresponding standardized effect value and then sum them to generate the QTL additive genetic contribution value. The QTL additive genetic contribution value and the initial genomic selection index are arithmetically summed and fused to obtain the final genomic selection index. The parental priority of each candidate individual is generated based on the final genomic selection index. A parent selection priority list is generated based on the parental priority corresponding to each candidate individual, sorted in descending order.

6. The breeding assistance method for redfin snapper based on multi-trait synergistic selection according to claim 1, characterized in that, The process of generating a candidate parent pool from the parent selection priority list, performing mating combination analysis on the candidate parent pool, predicting genetic gain from the generated mating combinations, and generating a mating assistance plan specifically includes: Obtain a parent selection priority list, select several candidate parent individuals within a preset range according to the parent selection priority list to generate a candidate parent pool, extract a subset of the genotype matrix, calculate the kinship coefficient through the genome relationship matrix, and generate a kinship relationship edge weight matrix; Based on the kinship edge weight matrix, a graph theory network is constructed with candidate parent individuals as nodes and kinship coefficients as edge weights. The graph theory network is then used as the input of the minimum spanning tree algorithm for breeding combination analysis. Initialize an empty tree structure, start traversing from the node with the highest genomic selection index, and select the effective edge that connects to the current tree and has the smallest affinity coefficient each time. Limit the number of mating connections of each node by a preset pairing threshold, and calculate the inbreeding coefficient of the population after adding new edges to control inbreeding. The iteration terminates when all nodes are connected or no legal edges can be added, and the minimum kinship cost mating tree is output. A mating list is generated based on the minimum kinship cost mating tree. For each mating combination in the mating list, the standardized genomic estimated breeding value of each trait of the parents is extracted, and the genotype distribution of the offspring is simulated through Mendel's laws of inheritance. The simulation results of offspring genotype distribution are input into the QTL effect model to calculate the expected breeding value of offspring and predict genetic gain. Finally, mating combinations within the preset genetic gain range are selected to generate mating assistance schemes for recommendation.

7. A breeding assistance system for redfin snapper based on multi-trait synergistic selection, characterized in that, The system includes a memory and a processor. The memory contains a program for a breeding assistance method for red snapper based on multi-trait synergistic selection. When the processor executes the program for the breeding assistance method for red snapper based on multi-trait synergistic selection, it performs the following steps: Phenotypic data and gene testing were conducted on the redfin snapper population to be bred, and a standardized phenotypic matrix and genotype matrix were constructed based on the collected data. Genetic parameters are estimated and breeding values ​​are predicted based on the standardized phenotypic matrix and genotype matrix to obtain breeding value prediction information. Based on the breeding value prediction information, the correlation between traits of the redfin snapper to be bred is analyzed and a trait association network is constructed. Then, the dynamic weights of multiple traits are configured to obtain a dynamic weight set. Based on the breeding value prediction information and dynamic weight set, the genomic selection index is calculated and the parental priority is analyzed to generate a parental selection priority list. A candidate parent pool is generated by the parent selection priority list, and mating combination analysis is performed on the candidate parent pool. Genetic gain prediction is then performed on the generated mating combinations to generate mating assistance schemes.

Citation Information

Patent Citations

  • Multi-character selection breeding method of fish and shrimp

    CN102823528A

  • Genome selection method for complex character genetic improvement

    CN113223606A

  • Method for evaluating comprehensive breeding value of growth and resistance characters of lateolabrax japonicus and application

    CN116064846A

  • Chicken intramuscular fat and abdominal fat genome balanced breeding method integrating prior information of SNP point set and biological gene chip of chicken intramuscular fat and abdominal fat genome balanced breeding method

    CN118136103A

  • Marker assisted best linear unbiased prediction (ma-blup): software adaptions for large breeding populations in farm animal species

    US20070105107A1

Cited By

  • Method for screening cross parent combination based on genome functional site information

    CN122337325A