Laying hen genetic disease knowledge graph construction and intelligent decision support system

By constructing a knowledge graph system for genetic diseases in laying hens, we have achieved collaborative integration and in-depth analysis of multi-source data, solving the problems of insufficient data integration and intelligent decision support in existing technologies, and improving the accuracy of disease prevention and control and the level of intelligent management.

CN121457593APending Publication Date: 2026-02-03CHINA AGRI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511625023.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve collaborative integration and in-depth analysis of multi-source data on genetic diseases in laying hens, and lack systematic knowledge integration and intelligent decision support, resulting in inaccurate and unintelligent disease diagnosis and management.

Method used

A knowledge graph system for genetic diseases in laying hens is constructed, including a weakly supervised graph extraction module for gene diseases, a multi-factor diagnosis module for diseases, a population health risk clustering module, a multi-omics graph parsing module, and an intelligent decision support module. Scientific decision-making schemes are generated through multi-dimensional data fusion and various algorithms.

Benefits of technology

It enables collaborative association and in-depth analysis of multi-source data, improves the accuracy of disease prevention and control and the level of intelligent breeding management, reduces the impact of diseases on the growth status and production performance of laying hens, and ensures the economic benefits of breeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121457593A_ABST
    Figure CN121457593A_ABST
Patent Text Reader

Abstract

The invention discloses a laying hen genetic disease knowledge graph construction and intelligent decision support system. The system comprises a genetic disease weak supervision graph extraction module, a disease multi-factor diagnosis module, a group health risk clustering module, a multi-omics graph analysis module, a genetic disease knowledge graph construction module and an intelligent decision support module. According to the system, laying hen genes, physiological indexes, breeding environments and multi-omics data are collected through gene sequencing and the like, a gene disease association graph is generated through weak supervision graph extraction, a disease diagnosis result is generated through multi-factor diagnosis, health risk grades are divided through risk clustering, and an association graph is generated through multi-omics analysis; and a genetic disease knowledge graph is constructed, and an intelligent decision scheme is generated in combination with real-time data and an algorithm. The system improves the accuracy of laying hen genetic disease prevention and control and the intelligent level of breeding management, guarantees the economic benefits of breeding, and is suitable for large-scale breeding scenes of laying hens.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of genetic diseases in laying hens, and more particularly to a knowledge graph construction and intelligent decision support system for genetic diseases in laying hens. Background Technology

[0002] In large-scale egg-laying hen farming, the occurrence of genetic diseases significantly impacts the growth status, production performance, and economic benefits of laying hens, making precise prevention and control, as well as scientific management, critical industry needs. Currently, the diagnosis and risk management of genetic diseases in laying hens rely on multi-dimensional information such as gene sequences, physiological indicators, farming environment, and multi-omics data. However, traditional technologies struggle to effectively correlate and deeply analyze various types of data and lack systematic integration and efficient utilization of disease-related knowledge. Furthermore, with the expansion of farming scale and the increasing complexity of disease types, relying solely on human experience or single algorithms for disease diagnosis, risk assessment, and decision-making is no longer sufficient to meet the requirements of precise and intelligent farming management. There is an urgent need to construct an integrated system that can integrate multi-source data, fuse multiple algorithms, and achieve knowledge graph construction and intelligent decision support to improve the accuracy and efficiency of genetic disease prevention and control in laying hens.

[0003] Existing technologies related to genetic diseases in laying hens have two significant drawbacks: First, they lack data integration and analysis capabilities. Most technologies can only process single types of data (such as gene data or physiological indicator data), failing to achieve collaborative integration and in-depth correlation analysis of multi-source data such as gene-disease association data, multi-omics data, and breeding environment data. This results in an incomplete understanding of disease mechanisms and difficulty in accurately identifying the intrinsic connections between diseases and various influencing factors. Second, they lack knowledge integration and intelligent decision support capabilities. Existing technologies lack a systematic review and mapping of knowledge related to genetic diseases in laying hens, failing to provide comprehensive knowledge support for disease diagnosis and management. Furthermore, the decision-making process often relies on a single algorithm or human experience, making it difficult to combine multi-dimensional data and multiple algorithms to generate scientific and optimized decision-making solutions. This fails to effectively meet the actual needs of precise prevention and intelligent management of genetic diseases in laying hens. Summary of the Invention

[0004] To overcome the shortcomings and deficiencies of existing technologies, this invention provides a knowledge graph construction and intelligent decision support system for genetic diseases in laying hens.

[0005] The technical solution adopted in this invention is a knowledge graph construction and intelligent decision support system for laying hen genetic diseases, comprising: a gene disease weakly supervised graph extraction module, which employs a weakly supervised learning mechanism based on laying hen gene sequence fragment feature matching, extracts image representation data of gene locus mutation associations, compares it with a preset gene disease feature library in multiple dimensions, and outputs gene disease association graph structure data; a disease multi-factor diagnosis module, which receives the gene disease association graph structure data, combines it with laying hen physiological indicator detection data and environmental parameter data, and performs data fusion through a multi-factor weight allocation mechanism to generate intermediate disease diagnosis results; and a population health risk clustering module, which obtains the intermediate disease diagnosis results, uses a hierarchical clustering algorithm to perform cluster analysis on the health data of different batches of laying hen populations, and divides them into health risk level clusters. The system comprises several modules: a multi-omics graph analysis module, which receives cluster data on population health risk levels, integrates data from the genomics, transcriptomics, proteomics, metabolomics, microbiome, and epigenetics of laying hens, analyzes the regulatory relationships between these omics data through a multi-omics data association analysis mechanism, and outputs a multi-omics association graph; a genetic disease knowledge graph construction module, which receives the multi-omics association graph and constructs a genetic disease knowledge graph for laying hens using a technical process of knowledge entity extraction, relationship mining, and attribute definition; and an intelligent decision support module, which receives the genetic disease knowledge graph for laying hens, combines it with real-time monitoring data on the growth status of laying hens and the breeding environment, and outputs intelligent decision-making solutions for the prevention and control of genetic diseases in laying hens and the management of breeding through decision rule reasoning and scheme optimization ranking. Different modules transmit and interact with each other through a data interface.

[0006] Furthermore, the expression for the weak supervised graph extraction algorithm for laying hen gene diseases used in the gene disease weak supervised graph extraction module is as follows: This represents the structure data of the genetic disease association graph in laying hens. Indicates the number of gene sequence fragments in laying hens. Indicates the first The weighting coefficients of each gene sequence fragment, Match( This represents a matching function between gene sequence fragment features and a predefined gene disease feature database. Indicates the first Feature vectors of gene sequence fragments This indicates the first [item] in the preset gene disease feature library. Vectors of class features This represents the graph convolution operation function. Indicates the first The adjacency matrix corresponding to each gene sequence fragment Indicates the first The node feature matrix of a gene sequence fragment Attention (represents the weight coefficients of the attention mechanism) ) represents the attention calculation function. They represent the first The feature matrix of a gene sequence fragment after graph convolution operation.

[0007] Furthermore, the multi-factor disease diagnosis module uses the following expression for the multi-factor diagnosis algorithm for laying hen diseases: This indicates intermediate results in the diagnosis of diseases in laying hens. Indicates the number of disease diagnostic factors. Indicates the first The weight values ​​of each diagnostic factor, Indicates the first Confidence coefficient of each physiological indicator test data This represents a function for standardizing environmental parameter data. Indicates the first Data on aquaculture environmental parameters, Indicates the first The influence coefficient of gene association data, Indicates the first Gene disease association graph structural data, Sigmod ( () represents the Sigmoid activation function. Indicates the first The results of a linear combination of diagnostic factors.

[0008] Furthermore, the expression for the egg-laying hen population health risk clustering algorithm used in the population health risk clustering module is as follows: KMeans represents cluster data indicating the health risk level of laying hen populations. () represents the K-means clustering algorithm function. Indicates the number of breeding batches. Indicates the first Weighting coefficients of intermediate results of disease diagnosis in each batch. Indicates the first Interim results of disease diagnosis for each batch of laying hens. Indicates the first Weighting coefficients for each batch of growth status data Indicates the first Growth status characteristics data of each batch of laying hens Indicates the number of clusters. This represents the function for calculating distances between clusters. They represent the first A cluster of health risk levels.

[0009] Furthermore, the expression for the multi-omics data association analysis model used by the multi-omics map parsing module is as follows: , This represents a multi-omics association map of laying hens. This represents a function for correlation analysis of multi-omics data. , These represent data from the genomics, transcriptome, proteome, metabolome, microbiome, and epigenetics of laying hens, respectively. This indicates the number of levels of regulation in multi-omics data. Indicates the first The weighting coefficients of each regulatory level This represents the function for analyzing the regulatory relationship. They represent the first , Multi-omics data feature matrix at each regulatory level.

[0010] Furthermore, the decision scheme generation model expression adopted by the intelligent decision support module is as follows: , RuleInfer represents an intelligent decision-making solution for the prevention and control of genetic diseases in laying hens and for their breeding management. ) represents a rule-based reasoning function based on a knowledge graph. This represents knowledge graph data on genetic diseases in laying hens. This represents a pre-defined base of decision rules. This represents the ranking function for optimizing decision-making schemes. Represents the set of candidate decision-making options. This represents the weight matrix of evaluation indicators for decision-making schemes. This represents the weighting coefficient for the fusion of historical schemes and real-time data, Update( This represents a scheme update function based on historical schemes and real-time data. This represents a database of historical decision-making schemes. This represents real-time monitoring data on the growth status of laying hens and the breeding environment.

[0011] Furthermore, the group health risk clustering module includes: a data preprocessing unit, which receives intermediate disease diagnosis results, performs preliminary processing of health data of different batches of laying hens by filtering outliers and filling in missing data to form a standardized group health dataset, removes non-compliant data through data verification, and supplements missing data using interpolation; a clustering parameter setting unit, which determines the number of clusters, distance calculation method, and number of iterations based on the scale of laying hen farming, breed characteristics, and differences in the farming area environment, compares the clustering effects of different parameter combinations, and selects the optimal parameter configuration; a clustering operation unit, which takes the standardized group health dataset as input, calculates sample similarity according to preset parameters, determines cluster centers, iteratively updates the results, divides the laying hen group into health risk level clusters, and outputs the results; and a result verification unit, which receives the cluster results, calculates the silhouette coefficient, adjusts the Rand index to verify the rationality, and if it does not meet the preset threshold, returns to the clustering parameter setting unit for readjustment until it meets the standard.

[0012] Furthermore, the multi-omics graph analysis module includes: a multi-omics data acquisition unit, which collects multi-omics data from laying hens using gene sequencing equipment, protein detection instruments, and metabolic analysis devices, converts it into a standardized digital format for storage, ensures consistent acquisition time through data synchronization, and reduces storage space through data compression; a data association analysis unit, which receives standardized omics data, analyzes omics data associations using Pearson correlation coefficient, partial correlation analysis, and mutual information calculation, constructs an association matrix, and screens strong association pairs to determine the calibration relationship; a regulatory relationship analysis unit, which combines prior knowledge of laying hen physiological metabolic pathways and gene regulatory networks, analyzes the regulatory direction, intensity, and path of omics data through path analysis and network construction, forming a regulatory relationship network; and a graph generation unit, which uses graphical modeling to present omics entities, associations, and regulatory paths as nodes, edges, and attribute labels, generating a multi-omics association graph, and optimizing the node layout and edge connection methods through graph optimization.

[0013] Furthermore, the genetic disease knowledge graph construction module includes: a knowledge entity extraction unit, which receives multi-omics association graph data, extracts entities related to laying hen genetic diseases using named entity recognition, classifies and standardizes names to form an entity set, distinguishes entities with the same name through entity disambiguation, and links entities to an existing knowledge base; a relationship mining unit, which, based on the entity set, mines entity relationships from multi-omics association data, literature, and experimental data using relationship extraction, defines types and quantifies strengths to form a relationship set, eliminates false relationships through relationship verification, and integrates multi-source data through relationship fusion; an attribute definition unit, which, combined with multi-omics association graph attribute data, detection data, and clinical data, defines basic attributes, characteristic attributes, and associated attributes of entities, standardizes attribute values ​​and defines ranges, forming entity-attribute-value triples; and a graph construction unit, which receives the entity set, relationship set, and triples, constructs a knowledge graph topology structure with entities as nodes, relationships as edges, and attributes as labels, integrates multi-source knowledge through knowledge fusion, detects contradictions and errors through knowledge verification, and forms a complete and accurate knowledge graph.

[0014] A knowledge graph construction and intelligent decision support system for genetic diseases in laying hens is proposed. The system operates as follows: First, laying hen gene sequence data is collected via gene sequencing equipment and transmitted to a weakly supervised gene disease extraction module. This module parses gene sequence features according to preset rules, compares them multiple times with a preset gene disease feature database, generates a gene disease association graph structure, and transmits it to a multi-factor disease diagnosis module. Second, the multi-factor disease diagnosis module receives the association graph structure data and simultaneously collects physiological indicator data and breeding environment data from laying hens. It then weights and fuses these data according to multi-factor weight allocation rules to generate intermediate disease diagnosis results, which are transmitted to a population health risk clustering module. Third, the population health risk clustering module receives the intermediate diagnosis results, collects growth status data from different batches of laying hens, calculates data similarity using a hierarchical clustering algorithm, and divides the data into clusters. The health risk level cluster data is transmitted to the multi-omics graph analysis module; in the fourth step, the multi-omics graph analysis module receives the cluster data, collects multi-omics data of laying hens, analyzes the regulatory relationships through multi-omics data association analysis, generates a multi-omics association graph, and transmits it to the genetic disease knowledge graph construction module; in the fifth step, the genetic disease knowledge graph construction module receives the association graph, extracts entities, mines relationships, defines attributes, constructs a genetic disease knowledge graph of laying hens, and transmits it to the intelligent decision support module; in the sixth step, the intelligent decision support module receives the knowledge graph, collects real-time data on the growth status and breeding environment of laying hens, generates candidate solutions through rule reasoning in combination with a preset decision rule base, evaluates and ranks the solutions using a solution optimization and ranking algorithm, and outputs an intelligent decision-making solution for the prevention and control of genetic diseases in laying hens and the management of breeding.

[0015] Beneficial Effects: This invention proposes a knowledge graph construction and intelligent decision support system for laying hen genetic diseases. It effectively integrates multi-source data such as laying hen gene sequences, physiological indicators, breeding environment, and multi-omics data, achieving collaborative association and in-depth analysis of the data. Simultaneously, the genetic disease knowledge graph constructed by the system can systematically organize disease-related knowledge and generate scientific decision-making solutions by combining multiple algorithms, significantly improving the accuracy of laying hen genetic disease prevention and control and the level of intelligent breeding management, reducing the impact of diseases on the growth status and production performance of laying hens, and ensuring the economic benefits of breeding. Addressing the problem of insufficient data integration and analysis capabilities, the system integrates multi-dimensional data and analyzes the regulatory relationships of various omics data through a multi-omics graph analysis module. Combined with a multi-factor disease diagnosis module, it achieves weighted fusion of multi-source data, comprehensively analyzing the disease occurrence mechanism and the correlation of influencing factors. Addressing the lack of knowledge integration and intelligent decision support capabilities, the system relies on a genetic disease knowledge graph construction module to complete the construction of a disease knowledge graph, providing comprehensive knowledge support for diagnosis and management. The intelligent decision support module combines multiple algorithms and real-time data to generate optimized decision-making solutions, meeting the needs of precise prevention and control and intelligent management of laying hen genetic diseases. Attached Figure Description

[0016] Figure 1 This is a diagram showing the system module composition of the present invention;

[0017] Figure 2 This is a flowchart of the system operation steps of the present invention. Detailed Implementation

[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] like Figure 1 As shown, the knowledge graph construction and intelligent decision support system for genetic diseases in laying hens includes:

[0020] The gene disease weakly supervised graph extraction module adopts a weakly supervised learning mechanism based on the feature matching of laying hen gene sequence fragments. It extracts the image representation data of gene site mutation associations, compares them with the preset gene disease feature library in multiple dimensions, and outputs gene disease association graph structure data.

[0021] Specifically, in the implementation of the weakly supervised graph extraction module for gene diseases, firstly, laying hen gene sequence data is collected using gene sequencing equipment. The gene sequence fragment length is set to 500-1000 base pairs. From this, 200-300 gene locus mutation-related image representations are extracted, including features such as base mutation type, mutation site location, and mutation frequency. A pre-set gene disease feature library contains feature vectors for 100-150 common laying hen genetic diseases, with each feature vector covering 50-80 key gene loci. The module employs a multi-dimensional comparison mechanism based on weakly supervised learning, setting a comparison threshold of 0.85. That is, when the matching degree between the extracted gene sequence fragment features and the feature vector of a certain disease in the feature library exceeds 0.85, it is determined that an association exists. The comparison process is conducted in three rounds. The first round is a coarse screening of feature vectors with a matching degree exceeding 0.7. The second round is an optimization screening of feature vectors with a matching degree exceeding 0.8. The final round is a precise screening of feature vectors with a matching degree exceeding 0.85. The final output is a gene-disease association graph structure data containing information such as disease type, associated gene loci, and matching degree. This module provides accurate gene-level data support for subsequent disease diagnosis through precise extraction and multiple rounds of comparison, reducing diagnostic errors caused by inaccurate gene data analysis.

[0022] The multi-factor disease diagnosis module receives gene-disease association graph structure data, combines it with laying hen physiological index detection data and environmental parameter data, and performs data fusion through a multi-factor weight allocation mechanism to generate intermediate disease diagnosis results.

[0023] Specifically, after receiving the gene-disease association graph structure data, the multi-factor disease diagnosis module simultaneously collects physiological indicator data and breeding environment data of laying hens. The physiological indicator data includes laying hen body temperature (detection range 38.5-40.5℃, detection accuracy ±0.1℃), weight (detection range 1.5-2.5kg, detection accuracy ±0.05kg), and egg production rate (statistical period 24 hours, calculation accuracy ±1%). The breeding environment data includes breeding house temperature (detection range 18-28℃, detection accuracy ±0.5℃), humidity (detection range 50%-70%, detection accuracy ±2%), and light duration (control range 14-16 hours / day, timing accuracy ±10 minutes). The module sets up a multi-factor weight allocation mechanism, with gene-disease association data accounting for 40% of the weight, physiological indicator data accounting for 35%, and environmental data accounting for 25%. The three types of data are then fused and calculated using a weighted summation formula. During the fusion process, various types of data are first standardized and mapped to the 0-1 range. Then, a comprehensive score is calculated according to the weight allocation ratio. When the comprehensive score exceeds 0.6, it is determined that there is a disease risk. An intermediate disease diagnosis result is generated, which includes the risk level (divided into three levels: low, medium, and high), key influencing factors, and risk probability (calculation accuracy ±3%). This module integrates multi-level data from genes, physiology, and environment through multi-factor fusion analysis, avoiding the limitations of single-data diagnosis and improving the comprehensiveness and accuracy of disease diagnosis.

[0024] The population health risk clustering module obtains intermediate results of disease diagnosis and uses a hierarchical clustering algorithm to perform cluster analysis on the health data of different batches of laying hens, dividing them into health risk level clusters.

[0025] Specifically, after obtaining intermediate disease diagnosis results, the group health risk clustering module collects health data from different batches of laying hens. The batch classification standard is set at 1000-2000 laying hens per batch, with a data collection cycle of 7 days per instance. The collected data includes intermediate disease diagnosis results and growth status data for each batch (daily weight gain detection range 15-25g / day, detection accuracy ±1g; feed conversion ratio detection range 2.0-2.5:1, calculation accuracy ±0.05). The module employs a hierarchical clustering algorithm, setting the number of clusters to 3-5 (corresponding to low, medium, high, and extremely high risk levels). Euclidean distance is used for distance calculation, and the number of iterations is set to 10-15 times to ensure stable clustering results. During the clustering process, the similarity between the health data of each batch of laying hens is first calculated. The similarity calculation is based on standardized indicators such as disease risk probability, daily weight gain, and feed conversion rate. Then, the initial cluster centers are determined according to the similarity. The cluster centers are updated iteratively until the change rate of the cluster centers between two adjacent iterations is less than 5%. The iteration stops and the health risk level clusters are divided. The output includes the batch number, risk level, proportion of laying hens in the cluster, and average values ​​of key health indicators for each cluster. This module achieves accurate classification of the health risks of different batches of laying hens through population clustering analysis, providing a basis for subsequent targeted prevention and control.

[0026] The multi-omics mapping module receives cluster data on population health risk levels, integrates data from the laying hen genome, transcriptome, proteome, metabolome, microbiome, and epigenetics, analyzes the regulatory relationships between various omics data through a multi-omics data association analysis mechanism, and outputs a multi-omics association map.

[0027] Specifically, the multi-omics mapping module receives cluster data on population health risk levels and integrates multi-omics data from laying hens. Genomic data is obtained through whole-genome sequencing with a sequencing depth of 30×-50× and a coverage requirement exceeding 98%. Transcriptome data is obtained through RNA sequencing with a sequencing volume of 10-20 Gb / sample. Proteome data is obtained through liquid chromatography-mass spectrometry (LC-MS / MS) with a detection sensitivity of 0.1 ng / mL. Metabolome data is obtained through gas chromatography-mass spectrometry (GC-MS / MS) with a detection resolution of 0.001 Da. Microbiome data is obtained through 16S rRNA gene sequencing with a sequencing depth of 10,000 genes / sample. Epigenetics data is obtained through methylation sequencing with a detection accuracy down to the single-base level. The module employs a multi-omics data association analysis mechanism, setting the significance threshold for association analysis to P<0.05. By calculating the correlation coefficients between various omics data (such as Pearson correlation coefficient and Spearman correlation coefficient), it filters omics data pairs with an absolute correlation coefficient exceeding 0.7, analyzes the regulatory relationships between various omics data, including the regulation of protein synthesis by gene expression and the impact of metabolite changes on the microbial community, and finally outputs a multi-omics association map containing omics data types, association relationship types, regulatory strength (quantification range 0-1), and significance levels. Through multi-omics integration and analysis, this module reveals the molecular mechanisms of genetic diseases in laying hens, providing in-depth data support for knowledge graph construction.

[0028] The genetic disease knowledge graph construction module receives multi-omics association graphs and uses a technical process of knowledge entity extraction, relationship mining, and attribute definition to construct a genetic disease knowledge graph for laying hens.

[0029] Specifically, after receiving the multi-omics association graph, the genetic disease knowledge graph construction module initiates the knowledge entity extraction process. Using named entity recognition technology, it extracts entities related to laying hen genetic diseases from the text descriptions and data tags of the multi-omics association graph. Entity types include disease types (e.g., leukemia, Marek's disease), gene names (e.g., chicken leukemia virus gene, Marek's disease virus gene), symptom characteristics (e.g., feather loss, sudden drop in egg production), detection indicators (e.g., viral antibody levels, gene mutation frequency), and prevention and control measures (e.g., vaccination, environmental disinfection). The entity extraction accuracy is required to exceed 95%. Subsequently, relationship mining is performed. Based on the association data of the multi-omics association graph and literature, relationships between entities are mined, such as "leukemia - associated gene - chicken leukemia virus gene" and "Marek's disease - symptom characteristics - feather loss." The relationship mining coverage is required to exceed 90%. Next, attribute definitions are performed for each entity. For example, the attributes of a disease entity include disease number (using a 10-digit numerical code), age of onset (e.g., 30-60 days old), and mortality rate (statistical range 0-100%, calculation accuracy ±2%). The completeness of attribute definitions must exceed 92%. Finally, a knowledge graph topology is constructed, with entities as nodes, relationships as edges, and attributes as node labels. A graph database is used for storage, with a data storage capacity supporting 100,000 entities and millions of relationships. After construction, knowledge verification technology is used to detect graph contradictions, with a contradiction detection rate required to exceed 98%. Ultimately, a complete and accurate knowledge graph of laying hen genetic diseases is formed. This module integrates scattered disease knowledge through systematic knowledge construction, providing comprehensive knowledge support for intelligent decision-making.

[0030] The intelligent decision support module receives a knowledge graph of genetic diseases in laying hens and combines it with real-time monitoring data on the growth status of laying hens and the breeding environment. Through decision rule reasoning and scheme optimization and ranking techniques, it outputs intelligent decision-making schemes for the prevention and control of genetic diseases in laying hens and breeding management. Different modules transmit and interact with each other through data interfaces.

[0031] Specifically, after receiving the knowledge graph of genetic diseases in laying hens, the intelligent decision support module collects real-time data on the growth status of the hens and the breeding environment. The data collection frequency is set to once per hour. The growth status data includes the body temperature, weight, and egg production rate of the hens, while the breeding environment data includes temperature, humidity, and light duration. The data transmission latency is required to be less than 10 seconds. The module initiates the decision rule reasoning process. The preset decision rule library contains 500-800 rules, such as "When the body temperature of a laying hen exceeds 40℃ and the egg production rate decreases by more than 10%, combined with the 'high temperature - disease risk - medium risk' rule in the knowledge graph, it is determined to be a medium risk disease." The rule reasoning response time is required to be less than 5 seconds. Subsequently, the module optimizes and sorts the solutions. The evaluation index weights are set according to the decision objectives (such as reducing mortality and increasing egg production rate), with mortality accounting for 40%, egg production rate accounting for 30%, and prevention and control costs accounting for 30%. The generated candidate decision solutions (such as "vaccination + environmental disinfection" and "isolation of sick chickens + feed drug addition") are comprehensively scored from 0 to 100 points. The solutions are sorted from highest to lowest score, and the top 3 optimal solutions are output. The final intelligent decision-making solution output includes the solution name, implementation steps (e.g., the dosage of vaccination is 0.5mL / bird, and the vaccination time is at the early stage of disease), expected effect (e.g., a 10%-15% reduction in mortality and a 5%-8% increase in egg production rate), and implementation period (e.g., 7-10 days). This module generates scientifically optimized decision-making solutions by combining real-time data with knowledge graph reasoning, guiding breeding production practices and improving the efficiency of genetic disease prevention and control and breeding management level of laying hens.

[0032] Preferably, the weak supervised graph extraction algorithm for laying hen gene diseases used in the gene disease weak supervised graph extraction module is expressed as follows: This represents the structure data of the genetic disease association graph in laying hens. Indicates the number of gene sequence fragments in laying hens. Indicates the first The weighting coefficients of each gene sequence fragment, Match( This represents a matching function between gene sequence fragment features and a predefined gene disease feature database. Indicates the first Feature vectors of gene sequence fragments This indicates the first [item] in the preset gene disease feature library. Vectors of class features This represents the graph convolution operation function. Indicates the first The adjacency matrix corresponding to each gene sequence fragment Indicates the first The node feature matrix of a gene sequence fragment Attention (represents the weight coefficients of the attention mechanism) ) represents the attention calculation function. They represent the first The feature matrix of a gene sequence fragment after graph convolution operation.

[0033] Specifically, in implementing the weakly supervised graph extraction algorithm for laying hen gene diseases, the number of laying hen gene sequence fragments is first determined, set to 50-80 based on differences in laying hen breeds. The weight coefficient of each fragment is allocated according to the degree of influence of gene site mutations on the disease, with the weight coefficient for core mutation sites set to 0.8-0.9 and the weight coefficient for secondary mutation sites set to 0.3-0.5. During the matching process between gene sequence fragment features and the preset gene disease feature library, a sliding window comparison method is used, with the window size set to 20-30 base pairs. After each comparison, a matching score is calculated, and a score exceeding 0.85 is considered a valid match. During graph convolution operations, the adjacency matrix dimension is set to 50×50 to 80×80 based on the number of gene sequence fragments. The node feature matrix includes 8-10 feature dimensions such as base type and mutation frequency. The ReLU activation function is used to optimize feature extraction during the operation. The attention mechanism weight coefficient is set to 0.6-0.7, and attention weights are assigned by calculating the cosine similarity between the feature matrices of different gene sequence fragments, with fragment pairs with a similarity exceeding 0.75 given higher weights. The final generated gene disease association graph structure data contains 50-80 nodes and 100-150 edges. The weight values ​​of the edges correspond to the association strength between segments. This algorithm improves the accuracy of gene disease association graph construction through multi-parameter collaborative optimization, providing a reliable gene-level data foundation for subsequent disease diagnosis.

[0034] Preferably, the multi-factor disease diagnosis algorithm for laying hens used in the disease multi-factor diagnosis module is expressed as follows: This indicates intermediate results in the diagnosis of diseases in laying hens. Indicates the number of disease diagnostic factors. Indicates the first The weight values ​​of each diagnostic factor, Indicates the first Confidence coefficient of each physiological indicator test data This represents a function for standardizing environmental parameter data. Indicates the first Data on aquaculture environmental parameters, Indicates the first The influence coefficient of gene association data, Indicates the first Gene disease association graph structural data, Sigmod ( () represents the Sigmoid activation function. Indicates the first The results of a linear combination of diagnostic factors.

[0035] Specifically, when implementing the multi-factor diagnostic algorithm for laying hen diseases, the number of disease diagnostic factors is first determined, covering 15-20 factors across three major categories: genetic, physiological, and environmental. The weight of each factor is set based on its contribution to disease diagnosis. Specifically, factors related to gene-related data have a weight of 0.3-0.4, factors related to physiological indicators have a weight of 0.25-0.35, and factors related to environmental parameters have a weight of 0.2-0.3. The confidence coefficient for physiological indicator detection data is determined based on the accuracy of the detection equipment. For high-precision equipment (detection error less than 0.1), the confidence coefficient is set to 0.9-0.95, and for ordinary-precision equipment (detection error 0.1-0.2), the confidence coefficient is set to 0.7-0.8. Environmental parameter data standardization uses the Min-Max standardization method, compressing the data to the 0-1 range, retaining four decimal places during processing. The influence coefficients of gene association data are set according to the strength of the association between genes and diseases. Strongly associated genes (match degree exceeding 0.9) have an influence coefficient of 0.8-0.9, while moderately associated genes (match degree 0.7-0.9) have an influence coefficient of 0.5-0.7. The linear combination results of the Sigmod activation function input values ​​are calculated through weighted summation, with weights consistent with the diagnostic factor weights. A function output value exceeding 0.6 is considered indicative of disease risk. The final intermediate disease diagnosis results include risk level, contribution of each factor, and other information. This algorithm improves the comprehensiveness and accuracy of disease diagnosis through multi-factor collaborative calculation, reducing the bias of single-factor diagnosis.

[0036] Preferably, the health risk clustering algorithm expression for laying hen populations used in the population health risk clustering module is as follows: KMeans represents cluster data indicating the health risk level of laying hen populations. () represents the K-means clustering algorithm function. Indicates the number of breeding batches. Indicates the first Weighting coefficients of intermediate results of disease diagnosis in each batch. Indicates the first Interim results of disease diagnosis for each batch of laying hens. Indicates the first Weighting coefficients for each batch of growth status data Indicates the first Growth status characteristics data of each batch of laying hens Indicates the number of clusters. This represents the function for calculating distances between clusters. They represent the first A cluster of health risk levels.

[0037] Specifically, when implementing the health risk clustering algorithm for laying hen populations, the number of breeding batches is first determined, typically 10-20 batches based on the farm size. The weighting coefficient for intermediate disease diagnosis results in each batch is allocated according to the number of laying hens in the batch: batches with more than 1500 hens have a weighting coefficient of 0.8-0.9, and batches with 1000-1500 hens have a weighting coefficient of 0.6-0.8. The weighting coefficient for growth status data is set according to the data collection frequency: 0.7-0.8 for collection frequency of once per day, and 0.5-0.6 for collection frequency of once every two days. The number of clusters in the K-means clustering algorithm is set based on historical disease occurrences: 5 clusters (extremely high, high, medium, low, and extremely low risk) are set during periods of high disease incidence, and 3 clusters (high, medium, and low risk) are set during periods of stable disease. The initial cluster centers are determined through random sampling. After each iteration, the variance of the data within each cluster is calculated, and iteration stops when the variance is less than 0.1. The distance between clusters is calculated using the Euclidean distance formula, with a distance threshold set at 5-8. Clusters exceeding the threshold are considered significantly different. The final generated health risk level cluster data includes information such as the number of batches in each cluster, risk level, and average values ​​of key indicators. This algorithm achieves accurate classification of the health risks of different batches of laying hens through population data cluster analysis, providing a basis for targeted prevention and control and improving the efficiency of population health management.

[0038] Preferably, the multi-omics data association analysis model expression used by the multi-omics map parsing module is as follows: , This represents a multi-omics association map of laying hens. This represents a function for correlation analysis of multi-omics data. , These represent data from the genomics, transcriptome, proteome, metabolome, microbiome, and epigenetics of laying hens, respectively. This indicates the number of levels of regulation in multi-omics data. Indicates the first The weighting coefficients of each regulatory level This represents the function for analyzing the regulatory relationship. They represent the first , Multi-omics data feature matrix at each regulatory level.

[0039] Specifically, when implementing the multi-omics data association analysis model, the following steps are taken: First, integrate multi-omics data from laying hens. For genomic data, the sequencing depth is set to 30×-50× based on analysis requirements, with a coverage of over 98% and an uncovered area not exceeding 2%. For transcriptome data, the sequencing volume is set to 10-20 Gb / sample, and the sequencing quality value Q30 must exceed 90%. For proteome data, liquid chromatography-mass spectrometry (LC-MS) is used, with a detection sensitivity of 0.1 ng / mL and a peptide matching rate exceeding 80%. For metabolome data, the detection resolution reaches 0.001 Da, and at least 500 metabolites are identified. For microbiome data, the sequencing depth reaches 10,000 records / sample, with a species annotation rate exceeding 95%. For epigenetics data, the detection accuracy reaches the single-base level, with methylation site coverage exceeding 90%. Multi-omics data correlation analysis uses a combination of Pearson and Spearman correlation coefficients. The Pearson correlation coefficient is used for linear relationship analysis, and the Spearman correlation coefficient is used for non-linear relationship analysis. A correlation coefficient with an absolute value exceeding 0.7 is considered a strong association. The number of regulatory levels in the multi-omics data was set to 4-6 (genome → transcriptome → proteome → metabolome → microbiome → epigenome). The weight coefficient of each level was set according to the intensity of the regulatory effect, with the upstream regulatory level (genome, transcriptome) having a weight coefficient of 0.8-0.9 and the downstream regulatory level (metabolome, microbiome) having a weight coefficient of 0.5-0.7. Path analysis was used to analyze regulatory relationships, identifying no fewer than 20 regulatory pathways. The final multi-omics association map included information such as omics data types, association relationships, and regulatory intensity. This model, through deep integration of multi-omics data, reveals the molecular mechanisms of disease occurrence and provides in-depth data support for knowledge graph construction.

[0040] Preferably, the decision scheme generation model expression adopted by the intelligent decision support module is: , RuleInfer represents an intelligent decision-making solution for the prevention and control of genetic diseases in laying hens and for their breeding management. ) represents a rule-based reasoning function based on a knowledge graph. This represents knowledge graph data on genetic diseases in laying hens. This represents a pre-defined base of decision rules. This represents the ranking function for optimizing decision-making schemes. Represents the set of candidate decision-making options. This represents the weight matrix of evaluation indicators for decision-making schemes. This represents the weighting coefficient for the fusion of historical schemes and real-time data, Update( This represents a scheme update function based on historical schemes and real-time data. This represents a database of historical decision-making schemes. This represents real-time monitoring data on the growth status of laying hens and the breeding environment.

[0041] Specifically, when implementing the decision-making scheme generation model, a rule-based reasoning mechanism based on a knowledge graph is first constructed. A pre-set decision rule base contains 500-800 rules, with rule priorities set according to the applicable scenarios. For emergency scenarios (such as disease outbreaks), the rule priority is set to levels 1-3, while for routine scenarios, it is set to levels 4-6. During rule reasoning, the number of rules matched each time is controlled to 10-20, the reasoning response time does not exceed 5 seconds, and the accuracy of the reasoning result is required to exceed 90%. The weights of the evaluation indicators for optimizing and ranking decision-making schemes are set according to the breeding objectives. When disease prevention and control is the core objective, the weights are: mortality rate 40%, prevention and control effectiveness 30%, and cost 30%. When economic benefits are the core objective, the weights are: egg production rate 40%, cost 30%, and mortality rate 30%. The number of candidate decision-making schemes is set to 5-8. The comprehensive score of each scheme is calculated through a weighted sum, with the score retained to two decimal places. The schemes are sorted from highest to lowest score, and the top three optimal schemes are output. The weighting coefficients for fusing historical plans and real-time data are set based on the accuracy of historical plans. Historical plans with an accuracy exceeding 90% have a weighting coefficient of 0.7-0.8, while those with an accuracy of 80%-90% have a weighting coefficient of 0.5-0.7. The final intelligent decision-making plan includes implementation steps, expected results, and cost budgets. Through knowledge reasoning and plan optimization, this model enhances the scientific rigor and practicality of decision-making, providing precise guidance for aquaculture management.

[0042] Preferably, the group health risk clustering module includes: a data preprocessing unit, which receives intermediate disease diagnosis results, performs preliminary processing of health data of different batches of laying hens by filtering outliers and filling in missing data to form a standardized group health dataset, removes non-compliant data through data verification, and supplements missing data by interpolation; a clustering parameter setting unit, which determines the number of clusters, distance calculation method, and number of iterations based on the scale of laying hen farming, breed characteristics, and differences in the farming area environment, compares the clustering effects of different parameter combinations, and selects the optimal parameter configuration; a clustering operation unit, which takes the standardized group health dataset as input, calculates sample similarity according to preset parameters, determines cluster centers, iteratively updates the results, divides the laying hen group into health risk level clusters, and outputs the results; and a result verification unit, which receives the cluster results, calculates the silhouette coefficient, adjusts the Rand index to verify the rationality, and if it does not meet the preset threshold, returns to the clustering parameter setting unit for readjustment until it meets the standard.

[0043] Specifically, the data preprocessing unit first receives intermediate disease diagnosis results and uses the 3σ principle to filter out outliers, i.e., removing data that deviates from the mean by more than 3 standard deviations. Missing data is filled using linear interpolation at an interval of 1 hour to ensure data continuity. The resulting standardized population health dataset must meet the requirements of data integrity ≥98% and outlier percentage ≤2%. Simultaneously, a format validation mechanism removes data that does not conform to XML format specifications to avoid errors in subsequent calculations. The clustering parameter setting unit determines parameters based on the farm size: 5 clusters for farms with over 50,000 animals, 4 clusters for 30,000-50,000 animals, and 3 clusters for farms with less than 30,000 animals. Euclidean distance is preferred for distance calculation; Manhattan distance is switched when the data dimension exceeds 10 dimensions. The initial iteration count is set to 20; if the cluster center change rate is still ≥5% after 15 iterations, the iteration count is increased to 30. The clustering operation unit takes a standardized dataset as input and iteratively operates according to the process of "calculating sample similarity → determining initial cluster centers → assigning samples to the nearest cluster → updating cluster centers." Sample similarity calculations are accurate to four decimal places. Cluster center updates use a weighted average method, with weights positively correlated with sample collection time. The weight coefficient for recent data is set to 0.8-0.9, and for older data, it is set to 0.4-0.6. The result verification unit calculates the silhouette coefficient and adjusts the Rand index. A silhouette coefficient ≥ 0.6 indicates reasonable clustering results, and an Rand index ≥ 0.7 indicates satisfactory cluster consistency. If these standards are not met, the system returns to the parameter setting unit for readjustment until the output of health risk level cluster data meets the requirements. This module ensures accurate group risk classification through multi-unit collaboration, providing clear targets for subsequent prevention and control.

[0044] Preferably, the multi-omics mapping module includes: a multi-omics data acquisition unit, which collects multi-omics data from laying hens using gene sequencing equipment, protein detection instruments, and metabolic analysis devices, converts it into a standardized digital format for storage, ensures consistent acquisition time through data synchronization, and reduces storage space through data compression; a data association analysis unit, which receives standardized omics data, analyzes omics data associations using Pearson correlation coefficient, partial correlation analysis, and mutual information calculation, constructs an association matrix, and screens strong association pairs to determine the labeling relationship; a regulatory relationship analysis unit, which, combined with prior knowledge of laying hen physiological metabolic pathways and gene regulatory networks, analyzes the regulatory direction, intensity, and path of omics data through path analysis and network construction to form a regulatory relationship network; and a mapping generation unit, which uses graphical modeling to present omics entities, associations, and regulatory paths as nodes, edges, and attribute labels to generate a multi-omics association mapping, and optimizes the node layout and edge connection methods through mapping.

[0045] Specifically, the multi-omics data acquisition unit acquired data using dedicated acquisition equipment. Genomic data was obtained using an Illumina sequencer with a sequencing depth of 30×-50×, a single-end read length of 150bp, and a sequencing quality Q30 ≥ 90%. Transcriptome data was obtained using the NovaSeq platform with a sequencing throughput of 10-20Gb / sample and gene coverage ≥ 95%. Proteome data was obtained using a ThermoQ Exactive mass spectrometer with a scanning range of 300-1800m / z, a resolution of 70000FWHM, and a peptide matching rate ≥ 80%. Metabolome data was obtained using an Agilent 7890A gas chromatograph with a column temperature program of holding at 60℃ for 2 min, increasing to 280℃ at 10℃ / min and holding for 5 min, and identifying ≥ 500 metabolites. Microbiome data was obtained using 16S rRNA V4 region sequencing with a sequencing depth of 10000 lines / sample and a species annotation rate ≥ 95%. Epigenetics data was obtained using Bisulfite sequencing with methylation site coverage ≥ 90% and detection accuracy down to the single-base level. The data association analysis unit uses Pearson correlation coefficient to analyze linear relationships, with a correlation coefficient absolute value ≥0.7 indicating a strong linear association. Spearman correlation coefficient is used to analyze nonlinear relationships, with a correlation coefficient absolute value ≥0.6 indicating a strong nonlinear association. The constructed association matrix must contain ≥2000 valid association pairs. The regulatory relationship analysis unit combines the KEGG chicken metabolic pathway database and identifies regulatory pathways through path enrichment analysis. Pathways with an enrichment fold ≥2 and a p-value <0.05 are considered significant regulatory pathways. ≥20 such pathways must be identified, and the regulatory direction (activation / inhibition) and regulatory intensity (quantified as 0-1, ≥0.8 indicating strong inhibition) of each pathway must be clearly defined. The map generation unit uses Cytoscape software to construct the map. Node size is positively correlated with data importance, and edge thickness is positively correlated with association strength. The map must contain ≥500 nodes and ≥1000 edges. A layout optimization algorithm (such as ForceAtlas2) is used to adjust the node distribution to ensure map readability. The final output multi-omics association map provides molecular-level data support for knowledge graph construction.

[0046] Preferably, the genetic disease knowledge graph construction module includes: a knowledge entity extraction unit, which receives multi-omics association graph data, extracts entities related to laying hen genetic diseases using named entity recognition, classifies and standardizes names to form an entity set, distinguishes entities with the same name through entity disambiguation, and links entities to an existing knowledge base; a relationship mining unit, which, based on the entity set, mines entity relationships from multi-omics association data, literature, and experimental data using relationship extraction, defines types and quantifies strengths to form a relationship set, eliminates false relationships through relationship verification, and integrates multi-source data through relationship fusion; an attribute definition unit, which, combined with multi-omics association graph attribute data, detection data, and clinical data, defines basic attributes, characteristic attributes, and associated attributes of entities, standardizes attribute values ​​and defines ranges, and forms entity-attribute-value triples; and a graph construction unit, which receives the entity set, relationship set, and triples, constructs a knowledge graph topology structure with entities as nodes, relationships as edges, and attributes as labels, integrates multi-source knowledge through knowledge fusion, detects contradictions and errors through knowledge verification, and forms a complete and accurate knowledge graph.

[0047] Specifically, the knowledge entity extraction unit uses the BERT-BiLSTM-CRF model, inputting text data from multi-omics association graphs (such as gene names and disease descriptions). The model training iterations are set to 100 rounds with a learning rate of 0.001. The entity extraction accuracy is ≥95% and the recall is ≥92%. The extracted entities are divided into five categories: disease type (such as leukemia, Marek's disease), gene name (such as ALV gene, MDV gene), symptom features (such as feather loss, sudden drop in egg production), detection indicators (such as viral antibody titer, gene mutation frequency), and prevention and control measures (such as vaccination, environmental disinfection). Each category has ≥100 entities. Entity disambiguation technology is used to distinguish entities with the same name (such as the meaning of "leukemia" in different contexts). The entity linking accuracy is ≥90%. The relation mining unit employs remote supervision combined with a CNN model to mine entity relationships from multi-omics association data and relevant PubMed literature (≥500 articles from the past 10 years). Relationship types include causal relationships (e.g., "ALV gene - causes - leukemia"), association relationships (e.g., "Marek's disease - associated with - feather loss"), and subordinate relationships (e.g., "chicken leukemia - belongs to - retroviral disease"). ≥50 of each type of relationship are mined. Relationships with a confidence score ≥0.8 are considered valid. False relationships (e.g., relationships with a confidence score <0.6) are removed through relation verification. Relationship fusion uses a weighted voting method, with weights positively correlated with the reliability of the data source. The attribute definition unit defines attributes for each entity. Disease entity attributes include disease number (10-digit code, such as JD001000001), age of onset (such as 30-60 days), mortality rate (statistical period of 30 days, accuracy ±2%), and susceptible breed (such as Lohmann Brown, Hy-Line White). Gene entity attributes include gene ID (such as NCBIGeneID), functional description (such as "encoding viral structural protein"), and mutation site (such as base A→G at position 120). Attribute integrity is ≥92%, and attribute values ​​are standardized (such as expressing mortality rate as a percentage, retaining one decimal place). The graph construction unit uses the Neo4j graph database, with node labels set to entity type and edge labels set to relation type. Attributes are stored as key-value pairs, and the database storage capacity supports ≥100,000 data entries. Knowledge fusion employs entity alignment (based on attribute similarity, with similarity ≥0.8 indicating the same entity) and relation merging (retaining the relation with the highest confidence for the same entity pair). Knowledge verification uses rule-based verification (e.g., the "disease-cause-symptom" relationship must conform to medical common sense) and consistency verification (e.g., no contradictory relationships, such as "A causes B" and "A inhibits B" cannot coexist). The verification accuracy is ≥98%. The final constructed knowledge graph contains ≥500 entities and ≥800 relations, providing systematic knowledge support for intelligent decision-making.

[0048] The weakly supervised graph extraction algorithm for laying hen gene diseases is a technique used to extract disease association graph structures from laying hen gene sequence data. Its implementation process first determines the number of laying hen gene sequence fragments (50-80 depending on the breed), and assigns fragment weight coefficients according to the impact of gene site mutations on disease (0.8-0.9 for core sites and 0.3-0.5 for minor sites). A sliding window comparison method of 20-30 base pairs is used to match the extracted gene sequence fragment features with a preset gene disease feature library (containing 100-150 disease feature vectors). A matching degree exceeding 0.85 is considered a valid association. Then, a graph convolution operation is performed between an adjacency matrix with dimensions adapted to the number of gene sequence fragments and a node feature matrix containing 8-10 feature dimensions. An attention mechanism with weight coefficients of 0.6-0.7 (weights are assigned based on cosine similarity, with higher weights given to similarities exceeding 0.75) is used, ultimately generating a gene disease association graph structure data containing 50-80 nodes and 100-150 edges. The algorithm aims to accurately mine the association between genes and diseases, providing reliable gene-level data support for subsequent disease diagnosis. Its significance lies in breaking through the limitations of traditional single gene locus analysis, improving the accuracy of gene-disease association analysis through multi-parameter collaborative optimization, and laying the foundation for early identification of genetic diseases in laying hens.

[0049] The multifactor diagnostic algorithm for laying hen diseases is a technical method that integrates genetic, physiological, and environmental data to determine the risk of diseases in laying hens. The implementation process first identifies 15-20 diagnostic factors (including three categories: genetic, physiological, and environmental), and assigns weights based on their diagnostic contribution (gene-related factors 0.3-0.4, physiological indicator factors 0.25-0.35, environmental parameter factors 0.2-0.3). For physiological indicator data (body temperature 38.5-40.5℃, weight 1.5-2.5kg, etc.), confidence coefficients are assigned based on the accuracy of the testing equipment (0.9-0.95 for high-precision equipment, 0.7-0.8 for ordinary-precision equipment). For environmental data (temperature 18-28℃, humidity 50%-70%, etc.), Min-Max standardization is used to compress the data to the 0-1 range. Gene association data influence coefficients are set according to the strength of the association between genes and diseases (strong association 0.8-0.9, moderate association 0.5-0.7). Then, the three types of data are weighted and fused, and the Sigmoid activation function is used to calculate the output value. A value exceeding 0.6 is considered indicative of disease risk, generating an intermediate diagnostic result containing risk level and the contribution of each factor. The algorithm aims to achieve accurate disease diagnosis by integrating multi-dimensional data. Its significance lies in avoiding the bias of diagnosis based on single data, improving the comprehensiveness and accuracy of disease diagnosis in laying hens, and providing accurate individual diagnostic basis for subsequent group health risk management.

[0050] The health risk clustering algorithm for laying hens is a technique that uses cluster analysis of health data from different batches of laying hens to classify risk levels. The process begins by identifying 10-20 breeding batches (depending on farm size). Disease diagnosis intermediate results are weighted according to the number of laying hens in each batch (0.8-0.9 for over 1500 hens, 0.6-0.8 for 1000-1500 hens), and growth status data is weighted according to data collection frequency (0.7-0.8 once / day, 0.5-0.6 once / 2 days). Then, a K-means clustering algorithm is used, with 3-5 clusters based on disease occurrence (5 during peak periods, 3 during stable periods). Distances are calculated using Euclidean distance (data dimension ≤ 10) or Manhattan distance (dimensional > 10). Initial iterations are performed 20 times. If the cluster center change rate is ≥ 5% after 15 iterations, the iterations are increased to 30. Iterations stop when the variance is < 0.1. Finally, the silhouette coefficient (≥ 0.6 is considered reasonable) and the adjusted Rand index (≥ 0.7 is considered acceptable) are calculated to verify the results. The output includes cluster data containing the number of batches and risk level. The algorithm aims to accurately classify the health risks of different batches of laying hens. Its significance lies in breaking through the limitations of traditional individual health management, providing a basis for developing differentiated prevention and control strategies for different risk clusters, and improving the efficiency of health management for laying hen populations.

[0051] The multi-omics mapping platform is a technical system that integrates multi-omics data from laying hens and analyzes their regulatory relationships. It is implemented by first collecting multi-omics data using specialized equipment: the genome is sequenced using an Illumina sequencer (depth 30×-50×, coverage ≥98%), the transcriptome using the NovaSeq platform (10-20Gb / sample, coverage ≥95%), the proteome using a ThermoQ Exactive mass spectrometer (sensitivity 0.1ng / mL, peptide matching rate ≥80%), the metabolome using an Agilent 7890A gas chromatograph (resolution 0.001Da, identification count ≥500 species), the microbiome using 16S rRNA sequencing (depth 10,000 lines / sample, annotation rate ≥95%), and the epigenetics using Bisulfite sequencing (single-base precision). The data coverage was ≥90%. Then, Pearson coefficient (linear relationship, absolute value ≥0.7 indicates strong association) and Spearman coefficient (non-linear relationship, absolute value ≥0.6 indicates strong association) were used to analyze data correlation, constructing an association matrix containing ≥2000 effective association pairs. Combined with the KEGG chicken metabolic pathway database, pathway enrichment analysis (enrichment fold ≥2, P<0.05) was used to identify ≥20 significant regulatory pathways, clarifying the direction and intensity of regulation (0-1 quantification, ≥0.8 indicates strong control). Finally, Cytoscape software was used to construct a multi-omics association map containing ≥500 nodes and ≥1000 edges (node ​​size is positively correlated with data importance, and edge thickness is positively correlated with association strength). The platform aims to reveal the molecular mechanisms of genetic diseases in laying hens. Its significance lies in integrating multi-dimensional omics data, providing in-depth data support for the construction of genetic disease knowledge graphs, and promoting the development of research on genetic diseases in laying hens from single-omics to multi-omics collaborative analysis.

[0052] like Figure 2As shown, the knowledge graph construction and intelligent decision support system for genetic diseases in laying hens includes the following steps: First, laying hen gene sequence data is collected using gene sequencing equipment and transmitted to a gene disease weakly supervised graph extraction module. This module parses gene sequence features according to preset rules, compares them multiple times with a preset gene disease feature database, generates a gene disease association graph structure data, and transmits it to a multi-factor disease diagnosis module. Second, the multi-factor disease diagnosis module receives the association graph structure data and simultaneously collects laying hen physiological indicator detection data and breeding environment data. It then weights and fuses these data according to multi-factor weight allocation rules to generate intermediate disease diagnosis results, which are transmitted to a population health risk clustering module. Third, the population health risk clustering module receives the intermediate diagnosis results, collects growth status data of laying hens from different batches, calculates data similarity using a hierarchical clustering algorithm, and divides the data into clusters. The process involves six steps: First, the health risk level cluster data is generated and transmitted to the multi-omics graph analysis module. Second, the multi-omics graph analysis module receives the cluster data, collects multi-omics data of laying hens, analyzes the regulatory relationships through multi-omics data association analysis, generates a multi-omics association graph, and transmits it to the genetic disease knowledge graph construction module. Third, the genetic disease knowledge graph construction module receives the association graph, extracts entities, mines relationships, and defines attributes to construct a genetic disease knowledge graph for laying hens, and transmits it to the intelligent decision support module. Fourth, the intelligent decision support module receives the knowledge graph, collects real-time data on the growth status and breeding environment of laying hens, generates candidate solutions through rule reasoning in combination with a preset decision rule base, evaluates and ranks the solutions using a solution optimization and ranking algorithm, and outputs an intelligent decision-making solution for the prevention and control of genetic diseases in laying hens and the management of breeding.

[0053] The knowledge graph construction and intelligent decision support system for laying hen genetic diseases achieves full-process coverage from data collection and analysis to decision output. It can efficiently integrate scattered data such as laying hen genes, physiological indicators, breeding environment, and multi-omics data, avoiding the problem of data isolation. The data analysis depth is sufficient. With the help of professional algorithms, it can perform correlation analysis and risk clustering on multi-source data, accurately uncover potential correlations between data, and provide comprehensive data support for disease diagnosis. The integration of knowledge and decision is high. The constructed genetic disease knowledge graph can systematically sort out disease-related knowledge and generate appropriate decision solutions by combining real-time data and algorithms. This greatly improves the accuracy and intelligence of breeding management, effectively reduces the interference of diseases on the growth and production performance of laying hens, and ensures the economic benefits of breeding.

[0054] This system addresses the issue of insufficient data integration and analysis capabilities. It integrates various omics data and analyzes regulatory relationships through a multi-omics graph analysis module, and combines this with a multi-factor disease diagnosis module to perform weighted fusion calculations on multi-source data, breaking through the limitations of single data processing and comprehensively analyzing the intrinsic connections between disease mechanisms and influencing factors. To address the lack of knowledge integration and intelligent decision support, the system relies on a genetic disease knowledge graph construction module to systematize and graph scattered disease knowledge, providing a complete knowledge system for disease diagnosis and management. Simultaneously, the intelligent decision support module combines multiple algorithms with real-time monitoring data to generate scientifically optimized decision-making schemes, completely eliminating the limitations of relying on human experience or single algorithms, and meeting the practical needs of precise prevention and intelligent management of genetic diseases in laying hens.

[0055] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0056] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A knowledge graph construction and intelligent decision support system for genetic diseases in laying hens, characterized in that, include: The gene disease weakly supervised graph extraction module adopts a weakly supervised learning mechanism based on the feature matching of laying hen gene sequence fragments. It extracts the image representation data of gene site mutation associations, compares them with the preset gene disease feature library in multiple dimensions, and outputs gene disease association graph structure data. The multi-factor disease diagnosis module receives gene-disease association graph structure data, combines it with laying hen physiological index detection data and environmental parameter data, and performs data fusion through a multi-factor weight allocation mechanism to generate intermediate disease diagnosis results. The population health risk clustering module obtains intermediate results of disease diagnosis and uses a hierarchical clustering algorithm to perform cluster analysis on the health data of different batches of laying hens, dividing them into health risk level clusters. The multi-omics mapping module receives cluster data on population health risk levels, integrates data from the laying hen genome, transcriptome, proteome, metabolome, microbiome, and epigenetics, analyzes the regulatory relationships between various omics data through a multi-omics data association analysis mechanism, and outputs a multi-omics association map. The genetic disease knowledge graph construction module receives multi-omics association graphs and uses a technical process of knowledge entity extraction, relationship mining, and attribute definition to construct a genetic disease knowledge graph for laying hens. The intelligent decision support module receives a knowledge graph of genetic diseases in laying hens and combines it with real-time monitoring data on the growth status of laying hens and the breeding environment. Through decision rule reasoning and scheme optimization and ranking techniques, it outputs intelligent decision-making schemes for the prevention and control of genetic diseases in laying hens and breeding management. Different modules transmit and interact with each other through data interfaces.

2. The knowledge graph construction and intelligent decision support system for laying hen genetic diseases according to claim 1, characterized in that, The weak supervised graph extraction algorithm for laying hen gene diseases used in the gene disease weak supervised graph extraction module is expressed as follows: This represents the structure data of the genetic disease association graph in laying hens. Indicates the number of gene sequence fragments in laying hens. Indicates the first The weighting coefficients of each gene sequence fragment, Match( This represents a matching function between gene sequence fragment features and a predefined gene disease feature database. Indicates the first Feature vectors of gene sequence fragments This indicates the first [item] in the preset gene disease feature library. Vectors of class features This represents the graph convolution operation function. Indicates the first The adjacency matrix corresponding to each gene sequence fragment. Indicates the first The node feature matrix of a gene sequence fragment Attention (represents the weight coefficients of the attention mechanism) ) represents the attention calculation function. They represent the first The feature matrix of a gene sequence fragment after graph convolution operation.

3. The knowledge graph construction and intelligent decision support system for laying hen genetic diseases according to claim 1, characterized in that, The multi-factor disease diagnosis module uses the following multi-factor diagnosis algorithm expression for laying hen diseases: This indicates intermediate results in the diagnosis of diseases in laying hens. Indicates the number of disease diagnostic factors. Indicates the first The weight values ​​of each diagnostic factor, Indicates the first Confidence coefficient of each physiological indicator test data This represents a function for standardizing environmental parameter data. Indicates the first Data on aquaculture environmental parameters, Indicates the first The influence coefficient of gene association data, Indicates the first Gene disease association graph structural data, Sigmod ( () represents the Sigmoid activation function. Indicates the first The results of a linear combination of diagnostic factors.

4. The knowledge graph construction and intelligent decision support system for laying hen genetic diseases according to claim 1, characterized in that, The expression for the layer hen health risk clustering algorithm used in the herd health risk clustering module is as follows: KMeans represents cluster data indicating the health risk level of laying hen populations. () represents the K-means clustering algorithm function. Indicates the number of breeding batches. Indicates the first Weighting coefficients of intermediate results of disease diagnosis in each batch. Indicates the first Interim results of disease diagnosis for each batch of laying hens. Indicates the first Weighting coefficients for each batch of growth status data Indicates the first Growth status characteristics data of each batch of laying hens Indicates the number of clusters. This represents the function for calculating distances between clusters. They represent the first A cluster of health risk levels.

5. The knowledge graph construction and intelligent decision support system for laying hen genetic diseases according to claim 1, characterized in that, The expression for the multi-omics data association analysis model used in the multi-omics map parsing module is as follows: , This represents a multi-omics association map of laying hens. This represents a function for correlation analysis of multi-omics data. , These represent data from the genomics, transcriptome, proteome, metabolome, microbiome, and epigenetics of laying hens, respectively. This indicates the number of levels of regulation in multi-omics data. Indicates the first The weighting coefficients of each regulatory level This represents the function for analyzing the regulatory relationship. They represent the first , Multi-omics data feature matrix at each regulatory level.

6. The knowledge graph construction and intelligent decision support system for laying hen genetic diseases according to claim 1, characterized in that, The decision scheme generation model expression adopted by the intelligent decision support module is as follows: , RuleInfer represents an intelligent decision-making solution for the prevention and control of genetic diseases in laying hens and for their breeding management. ) represents a rule-based reasoning function based on a knowledge graph. This represents knowledge graph data on genetic diseases in laying hens. This represents a pre-defined base of decision rules. This represents the ranking function for optimizing decision-making schemes. Represents the set of candidate decision-making options. This represents the weight matrix of evaluation indicators for decision-making schemes. This represents the weighting coefficient for the fusion of historical schemes and real-time data, Update( This represents a scheme update function based on historical schemes and real-time data. This represents a database of historical decision-making schemes. This represents real-time monitoring data on the growth status of laying hens and the breeding environment.

7. The knowledge graph construction and intelligent decision support system for laying hen genetic diseases according to claim 1, characterized in that, The population health risk clustering module includes: The data preprocessing unit receives intermediate disease diagnosis results, performs preliminary processing of health data of laying hens from different batches by filtering outliers and filling in missing data, forms a standardized group health dataset, removes non-compliant data through data verification, and uses interpolation to fill in missing data. The clustering parameter setting unit determines the number of clusters, distance calculation method, and number of iterations based on the scale of egg-laying hen farming, breed characteristics, and differences in the farming area environment. It compares the clustering effects of different parameter combinations and selects the optimal parameter configuration. The clustering operation unit takes a standardized population health dataset as input, calculates sample similarity, determines cluster centers, iteratively updates the results according to preset parameters, divides the laying hen population into health risk level clusters and outputs the results. The result verification unit receives the cluster results, calculates the profile coefficient, adjusts the Rand index to verify the rationality, and if it does not meet the preset threshold, it returns to the clustering parameter setting unit to readjust until it meets the standard.

8. The knowledge graph construction and intelligent decision support system for laying hen genetic diseases according to claim 1, characterized in that, The multi-omics mapping module includes: a multi-omics data acquisition unit, which collects multi-omics data from laying hens using gene sequencing equipment, protein detection instruments, and metabolic analysis devices, converts it into a standardized digital format for storage, ensures consistent acquisition time through data synchronization, and reduces storage space through data compression; a data association analysis unit, which receives standardized omics data, analyzes omics data associations using Pearson correlation coefficient, partial correlation analysis, and mutual information calculation, constructs an association matrix, and screens strong association pairs to determine the labeling relationship; a regulatory relationship analysis unit, which combines prior knowledge of laying hen physiological metabolic pathways and gene regulatory networks, analyzes the direction, intensity, and path of omics data regulation through path analysis and network construction, forming a regulatory relationship network; and a mapping generation unit, which uses graphical modeling to present omics entities, associations, and regulatory paths as nodes, edges, and attribute labels, generating a multi-omics association mapping, and optimizing the node layout and edge connection methods through mapping.

9. The knowledge graph construction and intelligent decision support system for laying hen genetic diseases according to claim 1, characterized in that, The genetic disease knowledge graph construction module includes: a knowledge entity extraction unit, which receives multi-omics association graph data, extracts entities related to laying hen genetic diseases using named entity recognition, classifies and standardizes names to form entity sets, distinguishes entities with the same name through entity disambiguation, and links entities to an existing knowledge base; a relationship mining unit, which, based on the entity sets, mines entity relationships from multi-omics association data, literature, and experimental data using relationship extraction, defines types and quantifies strengths to form relationship sets, eliminates false relationships through relationship verification, and integrates multi-source data through relationship fusion; an attribute definition unit, which, combining multi-omics association graph attribute data, detection data, and clinical data, defines basic attributes, characteristic attributes, and associated attributes of entities, standardizes attribute values ​​and defines ranges, forming entity-attribute-value triples; and a graph construction unit, which receives entity sets, relationship sets, and triples, constructs a knowledge graph topology structure with entities as nodes, relationships as edges, and attributes as labels, integrates multi-source knowledge through knowledge fusion, detects contradictions and errors through knowledge verification, and forms a complete and accurate knowledge graph.

10. The knowledge graph construction and intelligent decision support system for laying hen genetic diseases according to any one of claims 1-9, characterized in that, The system operation includes: The first step involves collecting the gene sequence data of laying hens using gene sequencing equipment and transmitting it to the gene disease weak supervision graph extraction module. This module parses the gene sequence features according to preset rules, compares them with the preset gene disease feature library in multiple rounds, generates gene disease association graph structure data, and transmits it to the disease multifactor diagnosis module. The second step is for the disease multifactor diagnosis module to receive the association graph structure data, and at the same time collect the physiological index detection data of laying hens and the breeding environment data. The data is then weighted and fused according to the multifactor weight allocation rules to generate intermediate disease diagnosis results, which are then transmitted to the population health risk clustering module. The third step is for the group health risk clustering module to receive intermediate diagnostic results, collect growth status data of laying hens from different batches, calculate data similarity and divide clusters according to the hierarchical clustering algorithm, and generate health risk level cluster data to the multi-omics graph analysis module. The fourth step involves the multi-omics graph analysis module receiving cluster data, collecting multi-omics data from laying hens, analyzing regulatory relationships through multi-omics data correlation analysis, generating a multi-omics correlation graph, and transmitting it to the genetic disease knowledge graph construction module. The fifth step involves the genetic disease knowledge graph construction module receiving the associated graph, extracting entities, mining relationships, and defining attributes to construct a knowledge graph of laying hen genetic diseases, which is then transmitted to the intelligent decision support module. The sixth step involves the intelligent decision support module receiving the knowledge graph, collecting real-time data on the growth status of laying hens and the breeding environment, generating candidate solutions through rule reasoning in conjunction with a pre-set decision rule base, evaluating and ranking the solutions using a solution optimization and ranking algorithm, and outputting an intelligent decision-making solution for the prevention and control of genetic diseases in laying hens and the management of breeding.

Citation Information

Cited By

  • River and lake health intelligent diagnosis and evaluation system and method fusing remote sensing and monitoring data

    CN121638687A