Laying hen genetic disease molecular marker screening system based on data fusion and AI prediction
The molecular marker screening system, which combines data fusion and AI prediction, solves the problems of multi-omics data integration and dynamic correlation capture, achieving efficient and accurate molecular marker screening and supporting disease-resistant breeding of laying hens.
Patent Information
- Application Number
- CN202511554956.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-02-03
AI Technical Summary
Existing technologies lack the ability to integrate multi-omics data in screening molecular markers for genetic diseases in laying hens, and cannot effectively capture the dynamic correlation of molecular characteristics, resulting in insufficient screening efficiency and accuracy, making it difficult to meet the needs of large-scale and precision breeding.
A molecular marker screening system based on data fusion and AI prediction is adopted, including multi-omics data acquisition, FineDataLink data fusion, dynamic temporal graph neural network processing, attention-enhanced deep forest analysis, and federated variational autoencoder modeling. Through feature alignment, association mapping, temporal feature capture, feature importance evaluation, and distributed training, a molecular marker probability distribution model is generated.
It achieves deep integration and efficient screening of multi-omics data, improves the accuracy and efficiency of molecular marker screening, provides reliable key molecular markers for disease-resistant breeding of laying hens, and meets the needs of large-scale and precision breeding.
Smart Images

Figure CN121459946A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of laying hen genetic diseases technology, and in particular to a molecular marker screening system for laying hen genetic diseases based on data fusion and AI prediction. Background Technology
[0002] In the egg-laying hen farming industry, the prevention and control of genetic diseases plays a crucial role in ensuring farming efficiency and egg quality. Screening for molecular markers strongly associated with genetic diseases is a core prerequisite for disease-resistant breeding. Currently, with the development of molecular biology techniques, multi-omics data such as SNP locus data and transcriptome gene expression data of laying hens can be obtained through professional collection methods. However, how to efficiently extract key information from massive amounts of multi-omics data and accurately locate molecular markers associated with genetic diseases has become a significant problem restricting the progress of disease-resistant breeding. Traditional molecular marker screening methods often rely on single-omics data, making it difficult to fully explore the association information between different omics data. Furthermore, they have low processing efficiency when dealing with complex data, failing to meet the needs of large-scale, precision breeding. Therefore, there is an urgent need to leverage data fusion technology and AI algorithms to build an efficient and accurate molecular marker screening system, integrating multi-dimensional data resources to improve the scientific rigor and timeliness of molecular marker screening.
[0003] Existing technologies for screening molecular markers for genetic diseases in laying hens have two significant drawbacks: First, in the data processing stage, existing technologies lack a professional fusion mechanism that can effectively integrate multi-omics data. This makes it impossible to achieve deep association mapping and feature alignment between different types of data, such as genomics and transcriptomics, leading to the omission of potential association information contained in the data. Consequently, it is difficult to form a comprehensive data foundation that fully reflects molecular characteristics, affecting the accuracy of subsequent molecular marker screening. Second, in the model application stage, the analysis models used in existing technologies mostly lack the ability to process time-series features and distributed training architecture. They cannot effectively capture the dynamic association relationships of molecular features over time, nor can they achieve multi-node collaborative training to optimize model performance while ensuring data security. This results in insufficient efficiency and accuracy of the model in screening molecular markers, making it difficult to stably output key molecular markers that meet the needs of disease-resistant breeding. Summary of the Invention
[0004] To overcome the shortcomings and deficiencies of existing technologies, this invention provides a molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction.
[0005] The technical solution adopted in this invention is a molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction, comprising: a multi-omics data acquisition module, a FineDataLink data fusion module, a dynamic temporal graph neural network processing module, an attention-enhanced deep forest analysis module, a federated variational autoencoder modeling module, and a molecular marker screening output module. The multi-omics data acquisition module acquires SNP site data from the laying hen genome and gene expression data from the transcriptome, and transmits the acquired data to the FineDataLink data fusion module. The FineDataLink data fusion module performs feature alignment and association mapping on the received genomic SNP site data and transcriptome gene expression data to generate a fusion feature matrix, which is then transmitted to the dynamic temporal graph neural network processing module. The dynamic temporal graph neural network processing module uses the fusion feature matrix as input to construct... A temporal correlation graph of molecular features related to genetic diseases in laying hens is generated. Through graph node feature updates and edge weight optimization, a temporal feature vector is output and transmitted to the attention-enhanced deep forest analysis module. This module evaluates the feature importance of the temporal feature vector and assigns hierarchical attention weights. Combining multi-granularity scanning and cascaded forest iterative learning, it outputs feature selection results, which are then transmitted to the federated variational autoencoder modeling module. Based on the feature selection results, the federated variational autoencoder modeling module constructs a federated training framework. Through variational inference and distributed parameter updates, it generates a molecular marker probability distribution model, which is transmitted to the molecular marker selection output module. The molecular marker selection output module, based on the molecular marker probability distribution model, extracts molecular markers related to genetic diseases in laying hens with probability values higher than a set threshold and outputs the selected preset molecular markers.
[0006] Furthermore, the dynamic temporal graph neural network model used in the dynamic temporal graph neural network processing module satisfies the formula: ;in, The graph structure matrix of the temporal association graph of molecular features related to genetic diseases in laying hens at time t is shown. It is the Sigmoid activation function. This is the feature weight matrix of the data from the previous time step. for Time-series multi-omics fusion feature matrix This is the weight matrix of the previous time-map structure. for Time-map structure matrix, It is the bias vector; Let be the feature matrix of the graph nodes at time t. It is a linear rectified activation function. Let be the graph adjacency matrix at time t. The feature transformation weight matrix for graph nodes. This is the feature transformation bias vector; simultaneously, the latent variable mean of the federated variational autoencoder is introduced. Constraints are applied to the graph structure update to satisfy... This is the constraint coefficient.
[0007] Furthermore, the attention-enhanced deep forest model used in the attention-enhanced deep forest analysis module satisfies the formula: ;in, For the feature attention weight vector, This is the attention weight matrix. Given the input time-series feature vector, For the number of features, Let i be the i-th eigenvector; This is the vector of feature filtering results. The number of decision trees in a cascaded forest. Let j be the attention weights corresponding to the j-th decision tree. Let be the predicted output of the j-th decision tree for the input features, and Softmax be the normalization function; combined with the dimension of the fused feature matrix output by the FineDataLink data fusion module. Set the decision tree node splitting threshold .
[0008] Furthermore, the federated variational autoencoder model used in the federated variational autoencoder modeling module satisfies the following formula: , ;in, For variational posterior distribution, For the set of variational parameters, As a latent variable, The feature selection results output by the attention-enhanced deep forest analysis module. It follows a normal distribution. The vector of latent variable means. The variance vector of the latent variables; For the federated training loss function, To generate the model parameter set, For expectation operator, To generate the log-likelihood of the model, Let KL divergence be the KL divergence. The prior distribution is based on the number of samples of genetic diseases in laying hens. Set the federated training batch size. .
[0009] Furthermore, the FineDataLink data fusion module satisfies the following formula when performing data fusion: ;in, To fuse the feature matrix, To integrate the weighting coefficients, This is a matrix of SNP loci data from the laying hen genome. For the genomic data weight matrix, This is a transcriptome gene expression data matrix. This is the transcriptome data weight matrix; To fuse the feature similarity matrix, This is the transpose of the fused feature matrix. This involves matrix trace operations; and considering the number of graph nodes in the dynamic temporal graph neural network processing module. Set the fusion feature dimension .
[0010] Furthermore, the molecular marker screening output module combines the molecular marker probability distribution model output by the federated variational autoencoder modeling module with the feature importance output by the attention-enhanced deep forest analysis module, satisfying the formula: Score Among them, the score is the overall score of molecular markers. This is the scoring weighting coefficient. The molecular marker probability value output by the federated variational autoencoder. To enhance attention, the feature importance index of the deep forest output is used; the marker is a pre-set set of molecular markers for the selected laying hen genetic diseases. For candidate molecule markers, The scoring threshold; the number of SNP loci acquired based on the multi-omics data acquisition module. Set the number of candidate molecular markers .
[0011] Furthermore, the dynamic temporal graph neural network processing module includes a temporal feature extraction unit, a graph structure construction unit, a graph node update unit, and an edge weight optimization unit. The temporal feature extraction unit receives the fused feature matrix output by the FineDataLink data fusion module, segments the fused feature matrix according to the time sequence, extracts the changing trends of molecular features related to laying hen genetic diseases in different time periods, and generates a temporal feature sequence. The graph structure construction unit uses the temporal feature sequence as a basis, takes the molecular features at each time point as graph nodes, calculates the connection strength between nodes based on the correlation between features, and constructs an initial molecular feature temporal association graph. The graph node update unit uses the gradient descent algorithm to iteratively update the node features in the initial association graph, and adjusts the feature value of the current node in combination with the feature information of adjacent nodes, so that the node features can better reflect the relationship between molecules. The edge weight optimization unit calculates the similarity between nodes based on the update results of the node features, adjusts the edge weight values according to the similarity, deletes edges with weights lower than the set value, and optimizes the structure of the temporal association graph.
[0012] Furthermore, the FineDataLink data fusion module includes a data format conversion unit, a feature alignment unit, an association mapping unit, and a fusion matrix generation unit. The data format conversion unit receives genomic SNP site data and transcriptome gene expression data transmitted from the multi-omics data acquisition module, converting the two types of data into a unified data format to ensure consistency in data dimensions and data types. The feature alignment unit performs feature matching on the converted data, matching the gene features corresponding to the genomic SNP site data and transcriptome gene expression data one-to-one based on the gene location information in the laying hen genetic information, thus performing feature-level alignment. The association mapping unit uses the Pearson correlation coefficient to calculate the correlation between aligned features, establishes an association model between genomic features and transcriptome features, and quantifies the association strength between features. The fusion matrix generation unit performs weighted fusion of genomic data and transcriptome data based on the association strength in the association model, generating a fusion feature matrix that includes multi-omics information.
[0013] Furthermore, the federated variational autoencoder modeling module includes a federated node initialization unit, a distributed training unit, a variational inference unit, and a parameter update unit. The federated node initialization unit distributes the genetic disease sample data of laying hens to multiple federated nodes, initializes the variational autoencoder model parameters of each federated node, and sets the initial learning rate and number of training iterations. The distributed training unit, on each federated node, uses the feature selection results output by the attention-enhanced deep forest analysis module as input to perform local model training, calculates the local training loss value, and saves the local model parameters. Within each federated node, the variational inference unit approximates the true posterior distribution through a variational posterior distribution, calculates the mean and variance of the latent variables, generates reconstructed data based on the latent variable distribution, and calculates the reconstruction error. The parameter update unit collects the local model parameters and loss values of each federated node, updates the global model parameters using a federated averaging algorithm, feeds the updated parameters back to each federated node, and repeats the local training and parameter update process until the model converges.
[0014] A molecular marker screening system for genetic diseases in laying hens, based on data fusion and AI prediction, operates in the following steps: First, a multi-omics data acquisition module extracts genomic SNP locus data and transcriptomic gene expression data from laying hen samples, and transmits the extracted data to the FineDataLink data fusion module. Second, the FineDataLink data fusion module converts the received data to a unified format, aligns features based on gene location information, calculates the correlation strength between features, constructs a fused feature matrix based on the correlation strength, and transmits the fused feature matrix to a dynamic temporal graph neural network processing module. Third, the dynamic temporal graph neural network processing module segments the fused feature matrix according to time series, extracts temporal feature sequences, constructs an initial correlation graph using temporal features as nodes, updates node features using a gradient descent algorithm, optimizes edge weights based on node similarity, and generates an optimized temporal correlation graph. The first step involves outputting the temporal feature vector to the attention-enhanced deep forest analysis module. The second step involves the attention-enhanced deep forest analysis module calculating feature attention weights on the temporal feature vector, assigning importance weights to different features, and inputting the weighted features into a cascaded forest for multi-granularity scanning and iterative learning. The feature selection results are then output to the federated variational autoencoder modeling module. The third step involves the federated variational autoencoder modeling module distributing the feature selection results to each federated node, initializing the model parameters for each node, performing local model training and calculating the loss value, obtaining the latent variable distribution through variational inference, updating the global parameters using a federated averaging algorithm, and repeatedly training until the model converges. This generates a molecular marker probability distribution model, which is then transmitted to the molecular marker selection output module. The fourth step involves the molecular marker selection output module combining the probability values in the molecular marker probability distribution model with the feature importance output by the attention-enhanced deep forest to calculate a comprehensive molecular marker score, selecting molecular markers with scores higher than a set threshold and outputting them.
[0015] Beneficial Effects: This invention proposes a molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction. In the data processing stage, a professional multi-omics data fusion mechanism is constructed through the FineDataLink data fusion module. This mechanism performs feature alignment, association mapping, and weighted fusion of genomic and transcriptomic data, fully mining potential association information between data to form a fusion data foundation that comprehensively reflects molecular characteristics. This solves the problems of insufficient data integration capabilities and missing association information in existing technologies, providing accurate data support for subsequent screening. In the model application stage, the dynamic time-series graph neural network processing module can capture the dynamic relationship between molecular features and changes over time. The federated deep forest analysis module enhances feature selection accuracy through feature importance assessment and hierarchical attention allocation. The federated variational autoencoder modeling module achieves multi-node collaborative training with a distributed training architecture, balancing data security and model performance optimization, and solving the problems of existing technologies lacking time-series processing capabilities and difficulty in collaborative training. At the same time, the modules are seamlessly connected to form a complete screening process, with efficient linkage from multi-omics data collection to molecular marker output, which greatly improves the efficiency and accuracy of molecular marker screening, provides reliable key molecular markers for disease-resistant breeding of laying hens, helps accelerate the disease-resistant breeding process, and meets the needs of large-scale and precision breeding. Attached Figure Description
[0016] Figure 1 This is a diagram showing the system module composition of the present invention;
[0017] Figure 2 This is a flowchart of the system operation steps of the present invention. Detailed Implementation
[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] like Figure 1 As shown, the molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction includes: a multi-omics data acquisition module, a FineDataLink data fusion module, a dynamic temporal graph neural network processing module, an attention-enhanced deep forest analysis module, a federated variational autoencoder modeling module, and a molecular marker screening output module.
[0020] The multi-omics data acquisition module acquires SNP site data of the laying hen genome and gene expression data of the transcriptome, and transmits the acquired data to the FineDataLink data fusion module;
[0021] Specifically, the implementation process of the multi-omics data acquisition module is as follows: Genomic and transcriptomic data are collected from 300-500 laying hen samples of different breeds using gene sequencing technology. Genomic data acquisition targets 28,000-32,000 SNP loci per hen. High-throughput sequencers are used to acquire raw sequencing data at a sequencing depth of 10-15×. Data with sequencing error rates higher than 0.01 are removed through quality control, ultimately retaining 25,000-28,000 valid SNP loci. Transcriptomic data acquisition selects tissues closely related to genetic diseases, such as the liver and spleen of laying hens. RNA sequencing technology is used to acquire the expression levels of 18,000-22,000 genes per hen. Genes with expression levels below 1 are filtered out, retaining 15,000-18,000 valid gene expression levels. This module ensures high accuracy and completeness of the acquired genomic SNP locus data and transcriptome gene expression data by precisely controlling sequencing depth and data filtering standards, providing reliable raw data support for subsequent data fusion and analysis. The number of samples collected and the settings for the number of loci and genes can meet the data scale and coverage requirements for molecular marker screening of genetic diseases in laying hens.
[0022] The FineDataLink data fusion module performs feature alignment and association mapping on the received genomic SNP site data and transcriptome gene expression data to generate a fusion feature matrix, which is then transmitted to the dynamic time series graph neural network processing module.
[0023] Specifically, the implementation process of the FineDataLink data fusion module is as follows: After receiving 25,000-28,000 SNP locus data and 15,000-18,000 gene expression level data transmitted from the multi-omics data acquisition module, the two types of data are first converted into a unified matrix format, where rows represent 300-500 laying hen samples and columns represent SNP loci and genes, respectively; then, based on the gene location information in the laying hen genome reference sequence, the SNP loci and corresponding genes are aligned to ensure that each SNP locus matches its corresponding gene sequence, forming 12,000-15,000 aligned feature pairs; subsequently, the Pearson correlation coefficient is used to calculate the association strength between each pair of features, and feature pairs with an absolute correlation coefficient higher than 0.3 are set as effective association pairs, retaining 10,000-12,000 effective association pairs; finally, the SNP locus data and gene expression level data in the effective association pairs are weighted according to a fusion weight coefficient of 0.4-0.6 to generate a fusion feature matrix with 300-500 rows and 10,000-12,000 columns. This module achieves deep fusion of multi-omics data through unified data format, precise feature alignment, strict association screening, and reasonable weight allocation. It fully explores the potential associations between data and provides comprehensive feature data for subsequent AI model processing. The association strength threshold and fusion weight coefficient set by the module can balance the contribution of the two types of data and ensure the effectiveness of the fusion feature matrix.
[0024] The dynamic temporal graph neural network processing module takes the fused feature matrix as input to construct a temporal association graph of molecular features related to genetic diseases in laying hens. Through graph node feature updates and edge weight optimization, it outputs a temporal feature vector, which is then transmitted to the attention-enhanced deep forest analysis module.
[0025] Specifically, the implementation process of the dynamic temporal graph neural network processing module is as follows: After receiving the fusion feature matrix output by the FineDataLink data fusion module, the fusion feature data of 300-500 laying hens is divided into 8-12 time segments according to the time series. Each time segment includes data from 25-40 laying hens. The changing trend of the fusion features within each time segment is extracted to form 8-12 temporal feature sequences. Using the features in each temporal feature sequence as graph nodes, a total of 10,000-12,000 graph nodes are set. The feature similarity between any two nodes is calculated. A connection is established between nodes with a similarity greater than 0.5 to construct an initial temporal correlation graph of molecular features. A gradient descent algorithm is used to iteratively update the node features in the initial correlation graph, with 50-80 iterations and a learning rate of 0.001-0.003. In each iteration, the feature value of the current node is adjusted based on the feature information of adjacent nodes. The similarity between nodes is recalculated based on the updated node features, edge weights are adjusted, edges with weights below 0.2 are deleted, and the temporal correlation graph structure is optimized. The final output is a temporal feature vector with 300-500 rows and 10,000-12,000 columns. This module effectively captures the dynamic correlation of molecular features through time segmentation, graph structure construction, node updates, and edge weight optimization, improving the temporality and correlation of feature data and providing more valuable temporal feature vectors for subsequent feature analysis. The set number of time segments, iteration parameters, and weight thresholds ensure the efficiency and accuracy of graph neural network processing.
[0026] The attention-enhanced deep forest analysis module evaluates the importance of features and assigns hierarchical attention weights to temporal feature vectors. Combining multi-granularity scanning and cascaded forest iterative learning, it outputs feature selection results and transmits the feature selection results to the federated variational autoencoder modeling module.
[0027] Specifically, the implementation process of the attention-enhanced deep forest analysis module is as follows: After receiving the temporal feature vector output by the dynamic temporal graph neural network processing module, the importance of 10,000-12,000 features in the temporal feature vector is initially assessed, and the correlation coefficient between each feature and the phenotypic data of genetic diseases in laying hens is calculated. Features with an absolute correlation coefficient higher than 0.2 (8,000-10,000 features) are retained. Next, a hierarchical attention mechanism is constructed, setting 3-5 attention levels. Each level assigns different attention weights to the input features, with weight values ranging from 0 to 1. This is achieved through backpropagation. The algorithm optimizes weight allocation through 30-50 iterations, giving higher weights to features closely related to the disease. The weighted features are then input into a cascaded forest model, which consists of 4-6 forest layers, each composed of 8-12 decision trees. Features are scanned at multiple granularities with a scanning window size of 5-10, obtaining feature representations at different granularities. Through iterative learning of the cascaded forest, the feature selection criteria are adjusted based on the output of the previous layer in each iteration, with 20-30 iterations. The final output includes 5,000-7,000 selected features. This module accurately identifies features strongly correlated with genetic diseases in laying hens through preliminary feature selection, hierarchical attention weight allocation, multi-granularity scanning, and cascaded forest iterative learning, improving the accuracy and targeting of feature selection and providing high-quality feature data for subsequent modeling. The attention hierarchy, forest structure, and iteration parameters ensure the depth and efficiency of feature analysis.
[0028] The federated variational autoencoder modeling module constructs a federated training framework based on the feature selection results. Through variational inference and distributed parameter updates, it generates a molecular marker probability distribution model and transmits the molecular marker probability distribution model to the molecular marker selection output module.
[0029] Specifically, the implementation process of the federated variational autoencoder modeling module is as follows: After receiving the feature selection results output by the attention-enhanced deep forest analysis module, 5,000-7,000 feature data are distributed to 5-8 federated nodes, with each federated node responsible for processing feature data from 60-80 laying hen samples; the variational autoencoder model parameters of each federated node are initialized, with the number of hidden layers in both the encoder and decoder set to 3-4 layers, the number of neurons in each layer set to 200-300, and the initial learning rate set to 0.002-0.005; local model training is performed on each federated node, with a training batch size set to 16-32. The training rounds are set to 40-60 rounds. In each round, the loss value of the local model is calculated, and the backpropagation algorithm is used to adjust the local model parameters. The mean and variance of the latent variables are calculated through variational inference, with the dimension of the latent variables set to 50-80. Reconstructed data is generated based on the distribution of the latent variables, and the reconstruction error is calculated to evaluate the model performance. The local model parameters and loss values of each federated node are collected, and the global model parameters are updated using a federated averaging algorithm 15-25 times. The updated parameters are fed back to each federated node, and the local training and parameter update process is repeated until the model loss value stabilizes in the range of 0.01-0.03, finally generating a molecular marker probability distribution model. This module achieves efficient model training and generates an accurate molecular marker probability distribution model while ensuring data security through federated node allocation, model initialization, local training, variational inference, and global parameter updates. The number of federated nodes, model structure, and training parameters set can balance the efficiency and accuracy of model training.
[0030] The molecular marker screening output module extracts molecular markers related to genetic diseases in laying hens with probability values higher than a set threshold based on the molecular marker probability distribution model, and outputs the preset molecular markers obtained through screening.
[0031] Specifically, the implementation process of the molecular marker screening output module is as follows: After receiving the molecular marker probability distribution model output by the federated variational autoencoder modeling module, the probability value corresponding to each candidate molecular marker in the model is extracted. The number of candidate molecular markers is 5,000-7,000. At the same time, the feature importance index output by the attention-enhanced deep forest analysis module is obtained, with the importance index ranging from 0 to 1. The probability value and importance index are weighted according to a weight coefficient of 0.5-0.7 to obtain the comprehensive score of each candidate molecular marker. The score calculation formula is: Comprehensive score = Weight coefficient × The algorithm calculates the molecular markers by adding a probability value to (1 - weighting coefficient) and then multiplying it by an importance index. A comprehensive scoring threshold of 0.6-0.8 is set, and 2,000-3,000 candidate molecular markers with comprehensive scores above the threshold are selected. Redundancy detection is performed on the selected molecular markers, calculating the similarity between any two molecular markers and deleting redundant molecular markers with similarity scores higher than 0.7, ultimately retaining 1,500-2,000 non-redundant molecular markers. The retained molecular markers are then sorted from highest to lowest comprehensive score, generating a list of molecular marker screening results, which is output to a terminal device for researchers to view and use. This module accurately extracts key molecular markers strongly correlated with genetic diseases in laying hens through comprehensive score calculation, threshold screening, redundancy detection, and result sorting, ensuring the accuracy and practicality of the output results. The weighting coefficient, scoring threshold, and similarity threshold effectively control the quality and quantity of molecular markers, meeting the needs of disease-resistant breeding for key molecular markers.
[0032] Preferably, the dynamic temporal graph neural network model used in the dynamic temporal graph neural network processing module satisfies the formula: ;in, The graph structure matrix of the temporal association graph of molecular features related to genetic diseases in laying hens at time t is shown. It is the Sigmoid activation function. The feature weight matrix of the data at the previous time step (dimension 1) , For feature dimension, (for input data dimensions) for Time-series multi-omics fusion feature matrix The previous time-map structure weight matrix (dimension 1) (where n is the number of nodes in the graph). for Time-map structure matrix, The bias vector (dimension 1) ); Let be the feature matrix of the graph nodes at time t. It is a linear rectified activation function. Let be the adjacency matrix of the graph at time t (dimension ). ), Transform the graph node features into a weight matrix (dimension 1). ), The feature transformation bias vector (dimension 1) Simultaneously, the latent variable mean of the federated variational autoencoder is introduced. Constraints are applied to the graph structure update to satisfy... This is the constraint coefficient (range 0.1-0.5).
[0033] Specifically, during the execution of the dynamic temporal graph neural network processing module, a dynamic update mechanism for the temporal association graph is first constructed. Targeting the temporal variation characteristics of molecular features of genetic diseases in laying hens, a matching relationship is set between the feature dimension and the input data dimension. The feature dimension is determined based on the number of columns in the fused feature matrix, ranging from 10,000 to 12,000. The input data dimension is consistent with the multi-omics data dimension for each laying hen, ranging from 25,000 to 28,000. The number of graph nodes is consistent with the number of columns in the fused feature matrix, ranging from 10,000 to 12,000. The bias vector dimension matches the number of graph nodes, ranging from 10,000 to 12,000. When calculating the graph structure matrix, the Sigmoid activation function is used to perform a non-linear transformation on the weighted result of the previous time step's data features and graph structure. The dimension of the weight matrix is set based on the input data dimension and the number of graph nodes to ensure the validity of matrix operations. When calculating the graph node feature matrix, a linear rectified activation function is used to process the product of the adjacency matrix, the graph structure matrix, and the transformation weight matrix. The dimension of the transformation weight matrix is set based on the number of graph nodes and the feature dimension. Simultaneously, the mean of the latent variables output by the federated variational autoencoder is introduced to constrain the graph structure update, with the constraint coefficient set to 0.1-0.5. This coefficient is adjusted to balance temporal correlation and model stability. Through the parameter settings and computational logic of the dynamic temporal graph neural network, the module's ability to capture temporal correlations of molecular features is improved, ensuring the accuracy of the output temporal feature vector and providing a reliable data foundation for subsequent feature analysis.
[0034] Preferably, the attention-enhanced deep forest model used in the attention-enhanced deep forest analysis module satisfies the formula: ;in, For the feature attention weight vector, The attention weight matrix (dimension 1) (where q is the attention dimension and q is the feature dimension). Given the input time-series feature vector, For the number of features, Let i be the i-th eigenvector; This is the vector of feature filtering results. The number of decision trees in a cascaded forest. Let j be the attention weights corresponding to the j-th decision tree. Let be the predicted output of the j-th decision tree for the input features, and Softmax be the normalization function; combined with the dimension of the fused feature matrix output by the FineDataLink data fusion module. Set the decision tree node splitting threshold .
[0035] Specifically, in the attention-enhanced deep forest analysis module, when constructing the feature attention weight calculation mechanism, the attention dimension is set to 10,000-12,000 based on the number of features in the temporal feature vector, ensuring consistency between the feature dimension and the attention dimension. The number of features is also 10,000-12,000, the same as the number of columns in the temporal feature vector, ensuring that each feature receives a corresponding attention weight. When calculating the attention weight vector, exponential functions and summation normalization are used to ensure the weight vector satisfies probability distribution characteristics, with weight values controlled within the range of 0-1. When calculating the feature selection result vector, the number of decision trees in the cascaded forest is set to 8-12, ensuring that the prediction results from multiple decision trees comprehensively reflect feature importance. Simultaneously, the Softmax normalization function is used to process the prediction results of the attention-weighted decision trees, ensuring the output conforms to a probability distribution. Combining the dimension of the fused feature matrix output by the FineDataLink data fusion module, the decision tree node splitting threshold is set to 0.3 times the dimension of the fused feature matrix, i.e., 3000-3600, ensuring that node splitting effectively distinguishes features of different importance. By enhancing the parameter settings and calculation rules of attention-based deep forests, the module improves the accuracy of its assessment of the importance of features related to genetic diseases in laying hens, giving key features higher weights and ensuring the relevance and reliability of feature selection results.
[0036] Preferably, the federated variational autoencoder model used in the federated variational autoencoder modeling module satisfies the formula: , ;in, For variational posterior distribution, For the set of variational parameters, Latent variables (dimension: (where I is the dimension of the latent variable). The feature selection results output by the attention-enhanced deep forest analysis module. It follows a normal distribution. The latent variable mean vector (dimension 1) ), The variance vector of latent variables (dimension 1) ); For the federated training loss function, To generate the model parameter set, For expectation operator, To generate the log-likelihood of the model, Let KL divergence be the KL divergence. Let the prior distribution be (denoted as standard normal distribution) Based on the number of samples of genetic diseases in laying hens Set the federated training batch size. .
[0037] Specifically, during the running of the federated variational autoencoder modeling module, the dimension of the latent variables is first determined, set to 50-80 based on the number of potential features of molecular markers for laying hen genetic diseases, ensuring that the latent variables can fully represent the core information of molecular features. The variational parameter set and the generator model parameter set include the weights and bias parameters of the encoder and decoder. The dimension of the weight parameters is set according to the number of input features and the dimension of the latent variables. The number of input features is 5,000-7,000, and the dimension of the bias parameters is consistent with the number of neurons in the corresponding layer. When calculating the variational posterior distribution, a normal distribution is used to model the latent variables, and the dimensions of the latent variable mean vector and variance vector are consistent with the dimension of the latent variables, which is 50-80. When calculating the federated training loss function, the mean of the log-likelihood of the generator model is calculated through the expectation operator, and the KL divergence between the variational posterior distribution and the standard normal prior distribution is calculated to balance the reconstruction accuracy and generalization ability of the model. Based on a sample size of 300-500 laying hen genetic diseases, the batch size for federated training is set to 1 / 10 of the sample size, i.e., 30-50, to ensure that batch training fully utilizes the sample data and avoids memory overflow. Through parameter settings and loss calculation logic of the federated variational autoencoder, the model performance of the module under the distributed training architecture is guaranteed, improving the accuracy of the molecular marker probability distribution model and providing reliable model support for subsequent molecular marker screening.
[0038] Preferably, the FineDataLink data fusion module satisfies the following formula when performing data fusion: ;in, To fuse the feature matrix, The fusion weighting coefficient (value range 0.4-0.6) A matrix of SNP loci in the laying hen genome (dimension 1) (where p is the number of samples and p is the number of SNP sites). The genomic data weight matrix (dimension 1) (for the fused feature dimensions) Transcriptome gene expression data matrix (dimension 1) (q is the number of genes). Transcriptome data weight matrix (dimension 1) ); To fuse the feature similarity matrix, This is the transpose of the fused feature matrix. This involves matrix trace operations; and considering the number of graph nodes in the dynamic temporal graph neural network processing module. Set the fusion feature dimension .
[0039] Specifically, when performing data fusion in the FineDataLink data fusion module, the fusion weight coefficient is first set to 0.4-0.6. This coefficient is adjusted to balance the contributions of genomic and transcriptomic data during the fusion process. A higher value is used when the genetic diseases of laying hens are more closely associated with genomic variations, and a lower value is used when they are more closely associated with gene expression. The sample size is kept consistent with the sample size in the multi-omics data acquisition module, at 300-500, the number of SNP sites is 25,000-28,000, and the number of genes is 15,000-18,000. The dimension of the fused feature matrix is set to 20,000-24,000 based on the subsequent model processing requirements, ensuring that the fused feature matrix balances data integrity and processing efficiency. When calculating the fused feature matrix, the dimension of the genomic data weight matrix is set to 25,000-28,000 × 20,000-24,000 based on the number of SNP sites and the dimension of the fused feature matrix, and the dimension of the transcriptomic data weight matrix is set to 15,000-18,000 × 20,000-24,000 based on the number of genes and the dimension of the fused feature matrix, ensuring the legality of the matrix operations. When calculating the fusion feature similarity matrix, covariance information between features is obtained through matrix transpose and matrix multiplication, followed by normalization through matrix trace operation to control the similarity matrix values within the range of 0-1. Considering the 10,000-12,000 graph nodes in the dynamic temporal graph neural network processing module, the fusion feature dimension is set to twice the number of graph nodes to ensure that the fusion features can fully support graph structure construction. Through parameter settings and calculation rules for data fusion, the module's ability to integrate multi-omics data is improved, ensuring the effectiveness of the fusion feature matrix and providing a comprehensive data foundation for subsequent model processing.
[0040] Preferably, the molecular marker probability distribution model output by the molecular marker screening output module combined with the feature importance output by the federated variational autoencoder modeling module and the attention-enhanced deep forest analysis module satisfies the formula: Score Among them, the score is the overall score of molecular markers. This is the scoring weighting coefficient (range: 0.5-0.7). The molecular marker probability value is the output of the federated variational autoencoder. To enhance the attention of the deep forest output, a feature importance index is used; the marker is a pre-defined set of molecular markers for the selected laying hen genetic diseases. For candidate molecule markers, The scoring threshold (range 0.6-0.8); the number of SNP loci obtained based on the multi-omics data acquisition module. Set the number of candidate molecular markers .
[0041] Specifically, during the molecular marker screening output module operation, the scoring weight coefficient is first set to 0.5-0.7. This coefficient is adjusted based on the reliability of the federated variational autoencoder model and the accuracy of the attention-enhanced deep forest model, with a higher value used when the model reliability is high, ensuring that the comprehensive score prioritizes the results of the high-reliability model. The molecular marker probability values output by the federated variational autoencoder range from 0 to 1, and the feature importance index output by the attention-enhanced deep forest also ranges from 0 to 1. Weighted calculations are used to control the comprehensive score range within 0-1. The scoring threshold is set to 0.6-0.8, adjusted according to the screening accuracy requirements for molecular markers of laying hen genetic diseases, with a higher value used when the accuracy requirement is high, ensuring that the screened molecular markers have strong correlations. The number of candidate molecular markers is consistent with the number of feature screening results output by the attention-enhanced deep forest, at 5,000-7,000. Based on the 25,000-28,000 SNP sites obtained from the multi-omics data acquisition module, the number of candidate molecular markers is set to 1 / 5 of the number of SNP sites, i.e., 5,000-5,600, ensuring that the candidate set size is moderate. When screening key molecular markers, a comprehensive score is compared with a threshold, retaining markers with scores above the threshold. Redundancy detection is then used for further screening, ultimately outputting key molecular markers that meet the requirements. By adjusting the parameters for molecular marker scoring and screening, the module's accuracy in identifying key molecular markers is improved, ensuring the practicality of the output results and providing reliable molecular marker resources for disease-resistant breeding of laying hens.
[0042] Preferably, the dynamic temporal graph neural network processing module includes a temporal feature extraction unit, a graph structure construction unit, a graph node update unit, and an edge weight optimization unit. The temporal feature extraction unit receives the fusion feature matrix output by the FineDataLink data fusion module, segments the fusion feature matrix according to the time sequence, extracts the changing trends of molecular features related to laying hen genetic diseases in different time periods, and generates a temporal feature sequence. The graph structure construction unit uses the temporal feature sequence as a basis, takes the molecular features at each time point as graph nodes, calculates the connection strength between nodes based on the correlation between features, and constructs an initial molecular feature temporal association graph. The graph node update unit uses the gradient descent algorithm to iteratively update the node features in the initial association graph, and adjusts the feature value of the current node in combination with the feature information of adjacent nodes, so that the node features can better reflect the relationship between molecules. The edge weight optimization unit calculates the similarity between nodes based on the update results of the node features, adjusts the edge weight values according to the similarity, deletes edges with weights lower than the set value, and optimizes the structure of the temporal association graph.
[0043] Specifically, the dynamic temporal graph neural network processing module includes a temporal feature extraction unit. This unit receives a fused feature matrix of 300-500 rows and 10,000-12,000 columns from the FineDataLink data fusion module. It then divides the data into 8-12 time segments, each containing feature data from 25-40 laying hen samples. By calculating the mean and variance of the features within each segment, it extracts the trend of feature changes over time, generating 8-12 sets of temporal feature sequences. The graph structure construction unit uses 10,000-12,000 features from the temporal feature sequences as graph nodes. It employs a cosine similarity algorithm to calculate the connection strength between any two nodes, setting a similarity threshold of 0.5. Nodes with similarity values above this threshold are connected, forming an initial molecular feature temporal association graph. The graph node update unit uses a gradient descent algorithm with a learning rate of 0.001-0.003 and 50-80 iterations. During each iteration, it adjusts the current node's feature value based on the mean of features from neighboring nodes, making the node features more closely match the molecular association patterns. The edge weight optimization unit recalculates the similarity based on the updated node features, updates the edge weights to the corresponding similarity values, sets the weight threshold to 0.2, and deletes edges below this threshold to optimize the graph structure and retain key associations. Through the parameter settings and operational logic of each unit, the module ensures that it can accurately capture the dynamic associations of molecular features, providing high-quality temporal feature vectors for subsequent analysis.
[0044] Preferably, the FineDataLink data fusion module includes a data format conversion unit, a feature alignment unit, an association mapping unit, and a fusion matrix generation unit. The data format conversion unit receives genomic SNP site data and transcriptome gene expression data transmitted from the multi-omics data acquisition module, converts the two types of data into a unified data format, and ensures that the data dimensions and data types are consistent. The feature alignment unit performs feature matching on the converted data, and according to the gene location information in the laying hen genetic information, it matches the gene features corresponding to the genomic SNP site data and the transcriptome gene expression data one-to-one, and performs feature-level alignment. The association mapping unit uses the Pearson correlation coefficient to calculate the correlation between the aligned features, establishes an association model between genomic features and transcriptome features, and quantifies the association strength between features. The fusion matrix generation unit performs weighted fusion of genomic data and transcriptome data according to the association strength in the association model, and generates a fusion feature matrix that includes multi-omics information.
[0045] Specifically, the FineDataLink data fusion module includes a data format conversion unit that receives 25,000-28,000 SNP loci and 15,000-18,000 gene expression level data from the multi-omics data acquisition module. It converts both types of data into a matrix format where rows represent 300-500 laying hen samples and columns represent features. The data type is uniformly floating-point, with precision retained to six decimal places to ensure consistency in data dimension and type. The feature alignment unit locates the gene position corresponding to each SNP locus based on the laying hen genome reference sequence. It matches the SNP locus data with the gene expression level data at that position, forming 12,000-15,000 feature pairs, with matching errors controlled within ±1 base nucleotide of the gene position. The association mapping unit uses the Pearson correlation coefficient to calculate the association degree between feature pairs. With a significance level of 0.05, it selects 10,000-12,000 valid association pairs with an absolute correlation coefficient higher than 0.3, quantifying the association strength between features. The fusion matrix generation unit assigns weights based on the strength of the association; higher association strength results in greater weights. Weight values range from 0.1 to 0.9. The data from effective association pairs are weighted and summed to generate a fusion feature matrix with 300-500 rows and 10,000-12,000 columns. By standardizing the operational criteria and parameters of each unit, deep fusion of multi-omics data is achieved, providing comprehensive feature support for subsequent model processing.
[0046] Preferably, the federated variational autoencoder modeling module includes a federated node initialization unit, a distributed training unit, a variational inference unit, and a parameter update unit. The federated node initialization unit distributes the genetic disease sample data of laying hens to multiple federated nodes, initializes the variational autoencoder model parameters of each federated node, and sets the initial learning rate and the number of training iterations. The distributed training unit trains the local model on each federated node using the feature selection results output by the attention-enhanced deep forest analysis module as input, calculates the local training loss value, and saves the local model parameters. Within each federated node, the variational inference unit approximates the true posterior distribution through a variational posterior distribution, calculates the mean and variance of the latent variables, generates reconstructed data based on the latent variable distribution, and calculates the reconstruction error. The parameter update unit collects the local model parameters and loss values of each federated node, updates the global model parameters using a federated averaging algorithm, feeds the updated parameters back to each federated node, and repeats the local training and parameter update process until the model converges.
[0047] Specifically, the federated variational autoencoder modeling module includes a federated node initialization unit. This unit distributes 5,000-7,000 feature data points output from the attention-enhanced deep forest analysis module evenly across 5-8 federated nodes, with each node processing 60-80 laying hen samples. The variational autoencoder model parameters for each node are initialized. Both the encoder and decoder have 3-4 hidden layers, with 200-300 neurons per layer. The initial learning rate is set to 0.002-0.005, and the number of iterations is set to 40-60. Within each node, the distributed training unit uses the feature data as input and performs local training in batches of 16-32. The model loss value is calculated for each batch, and the parameters are adjusted using backpropagation. The loss value convergence threshold is set to 0.01-0.03. The variational inference unit calculates the mean and variance of latent variables using Bayesian inference, with the latent variable dimension set between 50 and 80. It generates latent variable samples based on the mean and variance, and then uses a decoder to generate reconstructed data. The mean squared error between the reconstructed data and the original data is calculated as the reconstruction error. The parameter update unit collects model parameters and loss values from each node, and uses a federated averaging algorithm to weight the global parameters. The weights are proportional to the number of samples at each node. After updating, the parameters are fed back to each node, and training and updating are repeated until the global loss value stabilizes. Through the training and parameter calculation of each unit, model performance is optimized while ensuring data security, generating an accurate molecular marker probability distribution model.
[0048] The dynamic temporal graph neural network is the core model in this invention for processing the temporal features of multi-omics fusion data. Essentially, it captures the changing patterns of molecular characteristics of laying hen genetic diseases over time by constructing a dynamic association graph. In implementation, it first receives a fusion feature matrix of 300-500 rows and 10,000-12,000 columns output from the FineDataLink data fusion module. This matrix is divided into 8-12 segments according to the time series, with each segment containing data from 25-40 laying hens. The mean and variance of features within each segment are calculated to extract the temporal trend. Then, using 10,000-12,000 features as graph nodes, an initial association graph is constructed using a cosine similarity algorithm (threshold 0.5). Subsequently, the node features are updated using a gradient descent algorithm (learning rate 0.001-0.003, iterations 50-80), and the current node value is adjusted based on the features of adjacent nodes. Finally, the similarity is recalculated, redundant edges are removed according to a weight threshold of 0.2 to optimize the graph structure, and the temporal feature vector is output. This model transforms static fusion data into dynamic correlation features, accurately capturing temporal correlations between molecules. It breaks through the limitations of traditional static analysis, providing a dynamic data foundation that is more in line with actual physiological changes for subsequent feature screening, and improving the timeliness and accuracy of molecular marker screening.
[0049] The attention-enhanced deep forest model is the key model in this invention for feature importance assessment and selection. Combining the attention mechanism and the advantages of cascaded forests, it achieves accurate identification of features related to genetic diseases in laying hens. In implementation, the system first receives the temporal feature vector output from a dynamic temporal graph neural network, calculates the correlation coefficient between the feature and the disease phenotype (threshold 0.2), and initially selects 8,000-10,000 effective features. Then, it constructs 3-5 attention layers, assigning weights in the range of 0-1 through backpropagation (30-50 iterations) to give key features higher weights. Next, the weighted features are input into a cascaded forest containing 4-6 forest layers (8-12 decision trees per layer), performing multi-granularity scanning with a scanning window of 5-10. After 20-30 iterations of learning and adjusting the selection criteria, it finally outputs 5,000-7,000 feature results. This method accurately screens features strongly correlated with genetic diseases in laying hens from massive temporal features, solving the problem of insufficient targeting in traditional feature screening. It highlights key features through attention mechanisms and ensures screening accuracy through cascaded forests, providing high-quality features for subsequent modeling and laying the foundation for accurate molecular marker screening.
[0050] The federated variational autoencoder is a distributed training model used in this invention to construct a molecular marker probabilistic model. It combines the advantages of data privacy protection and accurate model modeling, and is suitable for multi-node collaborative processing of laying hen genetic data. In implementation, the 5,000-7,000 feature data output from the attention-enhanced deep forest are first evenly distributed to 5-8 federated nodes (each node processes 60-80 laying hen samples). The model parameters are initialized (3-4 hidden layers each for the encoder / decoder, 200-300 neurons per layer, learning rate 0.002-0.005). Each node undergoes local training in batches of 16-32, iterating for 40-60 rounds. Backpropagation is used to adjust parameters and calculate the loss (convergence threshold 0.01-0.03). Bayesian inference is used to calculate the mean and variance of 50-80 dimensional latent variables, generating reconstructed data and calculating the mean squared error. Finally, a federated averaging algorithm (weights proportional to the sample size) is used to collect node parameters and update the global model until the loss stabilizes, generating a molecular marker probability distribution model. While ensuring data privacy, we construct accurate molecular marker probabilistic models to solve the security challenges of multi-source data sharing, achieve distributed and efficient training, provide reliable probabilistic basis for molecular marker screening, and balance data security and modeling accuracy.
[0051] The FineDataLink data fusion platform is the core module for multi-omics data integration in this invention. Specifically designed for the characteristics of laying hen genome and transcriptome data, it achieves deep correlation and unification between the two types of data. In implementation, it first receives 25,000-28,000 SNP loci and 15,000-18,000 gene expression data from the multi-omics data acquisition module. This data is converted into a floating-point matrix (6 decimal places precision) with rows representing 300-500 laying hen samples and columns representing features, ensuring a unified data format. Then, based on the laying hen genome reference sequence, the gene positions corresponding to the SNP loci are located, and the two types of data are matched to form 12,000-15,000 feature pairs (error ±1 base). Next, the association strength is calculated using the Pearson correlation coefficient (significance 0.05), and 10,000-12,000 effective association pairs with an absolute correlation coefficient higher than 0.3 are selected. Finally, weights ranging from 0.1 to 0.9 are assigned according to the association strength, and the effective association pairs are weighted and summed to generate a 300-500 row, 10,000-12,000 column fusion feature matrix. Breaking down the barriers between genomic and transcriptomic data, achieving deep integration of multi-dimensional data, and solving the problem of one-sided information in traditional single-mathematical analysis, provides comprehensive and relevant feature data for subsequent AI models, which is a key data foundation for improving the accuracy of molecular marker screening.
[0052] like Figure 2As shown, the molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction includes the following steps: First, genomic SNP locus data and transcriptome gene expression data are extracted from laying hen samples using a multi-omics data acquisition module, and the extracted data is transmitted to the FineDataLink data fusion module. Second, the FineDataLink data fusion module converts the received data to a unified format, aligns features based on gene location information, calculates the correlation strength between features, constructs a fusion feature matrix based on the correlation strength, and transmits the fusion feature matrix to the dynamic temporal graph neural network processing module. Third, the dynamic temporal graph neural network processing module segments the fusion feature matrix according to the time series, extracts temporal feature sequences, constructs an initial correlation graph using temporal features as nodes, updates node features using a gradient descent algorithm, optimizes edge weights based on node similarity, generates an optimized temporal correlation graph, and outputs it to the dynamic temporal graph processing module. The process involves six steps: First, the temporal feature vector is fed to the attention-enhanced deep forest analysis module. Second, the attention-enhanced deep forest analysis module calculates the feature attention weights for the temporal feature vector, assigns importance weights to different features, inputs the weighted features into the cascaded forest for multi-granularity scanning and iterative learning, and outputs the feature selection results to the federated variational autoencoder modeling module. Third, the federated variational autoencoder modeling module distributes the feature selection results to each federated node, initializes the model parameters of each node, performs local model training and calculates the loss value, obtains the latent variable distribution through variational inference, updates the global parameters using the federated averaging algorithm, and repeatedly trains until the model converges, generating a molecular marker probability distribution model, which is then transmitted to the molecular marker selection output module. Fourth, the molecular marker selection output module combines the probability values in the molecular marker probability distribution model with the feature importance output by the attention-enhanced deep forest to calculate the comprehensive score of the molecular markers, selects molecular markers with scores higher than a set threshold, and outputs them.
[0053] This molecular marker screening system for genetic diseases in laying hens, based on data fusion and AI prediction, utilizes a FineDataLink data fusion module to construct a specialized fusion mechanism for hen genomic and transcriptomic data. First, the two types of data are formatted and features aligned. Then, feature association mappings are established based on gene location information, and weighted fusion is performed based on association strength. This fully explores the potential associations between different omics data, forming a comprehensive and accurate fusion feature matrix. This process effectively solves the problems of existing technologies that rely solely on single omics data, fail to capture data associations, and lead to the omission of key information. It provides a high-quality data foundation for subsequent molecular marker screening, ensuring the accuracy of the screening process.
[0054] At the model application and overall process level, the dynamic temporal graph neural network processing module can capture the dynamic correlation of molecular features over time. The attention-enhanced deep forest analysis module improves feature selection accuracy through feature importance assessment and hierarchical attention allocation. The federated variational autoencoder modeling module achieves multi-node collaborative training with a distributed training architecture, balancing data security and model optimization. Simultaneously, modules for multi-omics data acquisition, data fusion, model processing, and molecular marker output are seamlessly integrated, forming a complete and efficient screening process, with full linkage from data input to result output. This not only solves the problems of existing technologies lacking temporal processing capabilities and difficulty in collaborative training, but also significantly improves the efficiency and accuracy of molecular marker screening, providing reliable key markers for disease-resistant breeding of laying hens and meeting the needs of large-scale breeding.
[0055] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0056] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction, characterized in that, include: Multi-omics data acquisition module, FineDataLink data fusion module, dynamic temporal graph neural network processing module, attention-enhanced deep forest analysis module, federated variational autoencoder modeling module, molecular marker screening output module; The multi-omics data acquisition module acquires SNP locus data from the genomic genome and gene expression data from the transcriptome of laying hens, and transmits the acquired data to the FineDataLink data fusion module. The FineDataLink data fusion module performs feature alignment and association mapping on the received genomic SNP locus data and transcriptome gene expression data to generate a fused feature matrix, which is then transmitted to the dynamic temporal graph neural network processing module. The dynamic temporal graph neural network processing module uses the fused feature matrix as input to construct a temporal association graph of molecular features related to genetic diseases in laying hens. Through graph node feature updates and edge weight optimization, it outputs a temporal feature vector, which is then transmitted to the attention-enhanced deep forest analysis module. The attention-enhanced deep forest analysis module evaluates the feature importance of the temporal feature vector and assigns hierarchical attention weights. Combining multi-granularity scanning and cascaded forest iterative learning, it outputs feature selection results, which are then transmitted to the federated variational autoencoder modeling module. The federated variational autoencoder modeling module constructs a federated training framework based on the feature selection results. Through variational inference and distributed parameter updates, it generates a molecular marker probability distribution model and transmits the molecular marker probability distribution model to the molecular marker selection output module. The molecular marker selection output module extracts molecular markers related to laying hen genetic diseases with probability values higher than a set threshold based on the molecular marker probability distribution model and outputs the selected preset molecular markers.
2. The molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction as described in claim 1, characterized in that, The dynamic temporal graph neural network model used in the dynamic temporal graph neural network processing module satisfies the following formula: ;in, The graph structure matrix of the time-series correlation graph of molecular features related to genetic diseases in laying hens at time t is shown. It is the Sigmoid activation function. This is the feature weight matrix of the data from the previous time step. for Time-series multi-omics fusion feature matrix This is the weight matrix of the previous time-map structure. for Time-map structure matrix, It is the bias vector; Let be the feature matrix of the graph nodes at time t. It is a linear rectified activation function. Let be the graph adjacency matrix at time t. The feature transformation weight matrix for graph nodes. This is the feature transformation bias vector; simultaneously, the latent variable mean of the federated variational autoencoder is introduced. Constraints are applied to the graph structure update to satisfy... This is the constraint coefficient.
3. The molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction as described in claim 1, characterized in that, The attention-enhanced deep forest analysis module uses an attention-enhanced deep forest model that satisfies the following formula: ;in, For the feature attention weight vector, Here is the attention weight matrix. Given the input time-series feature vector, For the number of features, Let i be the i-th eigenvector; This is the vector of feature filtering results. The number of decision trees in a cascaded forest. Let j be the attention weights corresponding to the j-th decision tree. Let be the predicted output of the j-th decision tree for the input features, and Softmax be the normalization function; combined with the dimension of the fused feature matrix output by the FineDataLink data fusion module. Set the decision tree node splitting threshold .
4. The molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction according to claim 1, characterized in that, The federated variational autoencoder model used in the federated variational autoencoder modeling module satisfies the following formula: , ;in, For variational posterior distribution, For the set of variational parameters, As a latent variable, The feature selection results output by the attention-enhanced deep forest analysis module. It follows a normal distribution. The latent variable mean vector, The variance vector of the latent variables; For the federated training loss function, To generate the model parameter set, For expectation operator, To generate the log-likelihood of the model, Let KL divergence be the KL divergence. The prior distribution is based on the number of samples of genetic diseases in laying hens. Set the federated training batch size. .
5. The molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction according to claim 1, characterized in that, The FineDataLink data fusion module satisfies the following formula when performing data fusion: ;in, To fuse the feature matrix, To integrate the weighting coefficients, This is a matrix of SNP loci data from the laying hen genome. For the genomic data weight matrix, This is a transcriptome gene expression level data matrix. This is the transcriptome data weight matrix; To fuse the feature similarity matrix, This is the transpose of the fused feature matrix. This involves matrix trace operations; and considering the number of graph nodes in the dynamic temporal graph neural network processing module. Set the fusion feature dimension .
6. The molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction according to claim 1, characterized in that, The molecular marker screening output module, combined with the molecular marker probability distribution model output by the federated variational autoencoder modeling module and the feature importance output by the attention-enhanced deep forest analysis module, satisfies the formula: Score Among them, the score is the overall score of molecular markers. This is the scoring weighting coefficient. The molecular marker probability value is the output of the federated variational autoencoder. To enhance the attention of the deep forest output, a feature importance index is used; the marker is a pre-defined set of molecular markers for the selected laying hen genetic diseases. For candidate molecule markers, The scoring threshold; the number of SNP loci acquired based on the multi-omics data acquisition module. Set the number of candidate molecular markers .
7. The molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction according to claim 1, characterized in that, The dynamic temporal graph neural network processing module includes a temporal feature extraction unit, a graph structure construction unit, a graph node update unit, and an edge weight optimization unit. The temporal feature extraction unit receives the fusion feature matrix output by the FineDataLink data fusion module, segments the fusion feature matrix according to the time sequence, extracts the changing trends of molecular features related to genetic diseases of laying hens in different time periods, and generates a temporal feature sequence. The graph structure construction unit is based on the temporal feature sequence. It uses the molecular features at each time point as graph nodes and calculates the connection strength between nodes based on the correlation between features to construct an initial molecular feature temporal association graph. The graph node update unit uses the gradient descent algorithm to iteratively update the node features in the initial association graph. It adjusts the feature value of the current node by combining the feature information of adjacent nodes so that the node features can better reflect the relationship between molecules. The edge weight optimization unit calculates the similarity between nodes based on the updated node features, adjusts the edge weight values according to the similarity, deletes edges with weights lower than the set value, and optimizes the structure of the temporal association graph.
8. The molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction according to claim 1, characterized in that, The FineDataLink data fusion module includes a data format conversion unit, a feature alignment unit, an association mapping unit, and a fusion matrix generation unit. The data format conversion unit receives genomic SNP site data and transcriptome gene expression data transmitted from the multi-omics data acquisition module and converts the two types of data into a unified data format to ensure consistency in data dimensions and data types. The feature alignment unit performs feature matching on the converted data. Based on the gene location information in the laying hen genetic information, it matches the genomic SNP site data with the corresponding gene features in the transcriptome gene expression data one-to-one, thus aligning them at the feature level. The association mapping unit uses the Pearson correlation coefficient to calculate the correlation between aligned features, establishes an association model between genomic features and transcriptome features, and quantifies the association strength between features. The fusion matrix generation unit performs weighted fusion of genomic data and transcriptome data based on the association strength in the association model, generating a fusion feature matrix that includes multi-omics information.
9. The molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction according to claim 1, characterized in that, The federated variational autoencoder modeling module includes a federated node initialization unit, a distributed training unit, a variational inference unit, and a parameter update unit. The federated node initialization unit distributes the genetic disease sample data of laying hens to multiple federated nodes, initializes the variational autoencoder model parameters for each federated node, and sets the initial learning rate and number of training iterations. The distributed training unit, on each federated node, uses the feature selection results output by the attention-enhanced deep forest analysis module as input to perform local model training, calculates the local training loss value, and saves the local model parameters. Within each federated node, the variational inference unit approximates the true posterior distribution using a variational posterior distribution, calculates the mean and variance of latent variables, generates reconstructed data based on the latent variable distribution, and calculates the reconstruction error. The parameter update unit collects the local model parameters and loss values from each federated node, updates the global model parameters using a federated averaging algorithm, feeds the updated parameters back to each federated node, and repeats the local training and parameter update process until the model converges.
10. The molecular marker screening system for genetic diseases in laying hens based on data fusion and AI prediction according to any one of claims 1-9, characterized in that, The system operation includes the following steps: The first step involves extracting genomic SNP locus data and transcriptome gene expression data from laying hen samples using a multi-omics data acquisition module, and then transmitting the extracted data to the FineDataLink data fusion module. The second step is to convert the received data into a format so that the two types of data meet the same format requirements. Then, the feature is aligned according to the gene location information, the correlation strength between features is calculated, a fusion feature matrix is constructed based on the correlation strength, and the fusion feature matrix is transmitted to the dynamic time series graph neural network processing module. The third step involves the dynamic temporal graph neural network processing module segmenting the fused feature matrix according to the time series, extracting the temporal feature sequence, constructing an initial association graph using the temporal features as nodes, updating the node features through the gradient descent algorithm, optimizing the edge weights based on node similarity, generating an optimized temporal association graph, and outputting the temporal feature vector to the attention-enhanced deep forest analysis module. The fourth step involves the attention-enhanced deep forest analysis module calculating feature attention weights for the temporal feature vectors, assigning importance weights to different features, inputting the weighted features into the cascaded forest for multi-granularity scanning and iterative learning, and outputting the feature selection results to the federated variational autoencoder modeling module. The fifth step involves the federated variational autoencoder modeling module distributing the feature selection results to each federated node, initializing the model parameters of each node, performing local model training and calculating the loss value, obtaining the latent variable distribution through variational inference, updating the global parameters using the federated averaging algorithm, and repeatedly training until the model converges, generating a molecular marker probability distribution model, and transmitting it to the molecular marker selection output module. The sixth step involves the molecular marker screening and output module combining the probability values in the molecular marker probability distribution model with the feature importance of the attention-enhanced deep forest output to calculate a comprehensive molecular marker score, and then screening out and outputting molecular markers with scores higher than a set threshold.
Citation Information
Cited By
Law case recommendation method and system based on large model
CN121901408A