Laying hen epidemic disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN

Through the hybrid model of DistilBERT-GATv2 and BiLSTM-TCN, the complexity of information extraction in laying hen disease detection was solved, efficient and accurate entity recognition and relationship extraction were achieved, a visual knowledge graph was constructed, and the intelligence and systematization level of laying hen disease detection was improved.

CN120767002APending Publication Date: 2025-10-10SHANDONG AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510892029.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently and accurately achieve intelligent detection and diagnosis of laying hen diseases, especially when performing named entity recognition and relationship extraction in unstructured texts, which are prone to information dispersion, strong subjectivity, and low response efficiency.

Method used

A hybrid model of DistilBERT-GATv2 and BiLSTM-TCN is adopted. By constructing a dependency grammar structure graph and introducing a graph attention mechanism, combined with a cross-sentence reference resolution mechanism, high-quality named entity recognition and relationship extraction of laying hen disease information are achieved, and a knowledge graph is constructed and visualized.

Benefits of technology

It significantly improves the accuracy and intelligence of laying hen disease detection, enhances the robustness of the model in complex sentences and long-distance entity recognition, ensures the consistency and accuracy of entities in the knowledge graph, and supports applications such as intelligent question answering and auxiliary diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120767002A_ABST
    Figure CN120767002A_ABST
Patent Text Reader

Abstract

The invention discloses a laying hen epidemic disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN, and the method comprises the following steps: extracting related data from unstructured text data, and constructing a high-quality structured text database; a DistilBERT model is adopted to carry out fine tuning training on the labeled corpus, accurate recognition of a target entity is achieved, and standardized entity categories and texts are output; realizing extraction of a semantic relationship between entities, and constructing disease triple data; a cross-sentence anaphora resolution mechanism and an entity normalization algorithm are introduced, semantic references are unified, and consistency and uniqueness of node semantics in the knowledge graph are ensured; storing the constructed triple data in a graph database to complete the construction of the domain exclusive knowledge graph; visual presentation of entity nodes, relation edges and query paths is achieved by configuring a visual component. According to the method, the epidemic disease type of the laying hen can be efficiently detected and diagnosed, the accuracy of an intelligent monitoring system is improved, manual intervention is reduced, and the automation level of laying hen disease prevention and control is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and more specifically, relates to a laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN. Background Art

[0002] With the continued expansion of laying hen farming and the development of intelligent farming technologies, the industry has placed higher demands on the real-time and precise prevention and control of disease outbreaks. The spread of laying hen diseases not only severely impacts the production performance and health of flocks, but can also cause widespread economic losses and, in severe cases, raise food safety concerns. Currently, the diagnosis of laying hen diseases still primarily relies on veterinary experience and manual review of paper literature. These methods are subject to high subjectivity, fragmented information, and low response efficiency, making them unable to meet the practical needs of modern farms for intelligent management and efficient decision-making.

[0003] The rapid development of natural language processing (NLP) technology has provided new insights for information mining and knowledge acquisition in the agricultural sector. By identifying and extracting domain-specific entities and relationships from large amounts of unstructured text, it is possible to structure information about laying hen diseases, laying the foundation for downstream applications such as intelligent question answering and knowledge reasoning. Building disease knowledge graphs for laying hen farming scenarios has become a key path to intelligent diagnosis and treatment, as well as precise prevention and control. Knowledge graphs explicitly represent entities and their semantic relationships using a graph structure. They can systematically integrate domain knowledge such as disease names, symptoms, and prevention and control recommendations, providing data support for building efficient disease early warning and auxiliary diagnosis systems. However, high-quality named entity recognition (NER) and relation extraction (RE) in text related to laying hen diseases still face multiple challenges. Relevant corpora are generally characterized by dense terminology, complex syntactic structures, and strong semantic dependencies. Traditional methods based on rule-based template matching or shallow machine learning struggle to achieve accurate extraction and cover knowledge content with diverse expressions and implicit relationships.

[0004] To improve the quality of information extraction, various deep learning technologies have been actively applied in the field of agricultural information extraction in recent years. The pre-trained model BERT based on the Transformer architecture has significant advantages in contextual semantic modeling and is widely used in semantic representation tasks. Its lightweight version DistilBERT significantly reduces the model running load while maintaining semantic representation capabilities, making it more suitable for resource-constrained agricultural application scenarios. The graph neural network GATv2 has a strong ability in modeling non-Euclidean semantic relationships and effectively captures the complex relationship between disease names and disease symptoms; the hybrid time series structure that integrates BiLSTM and TCN can take into account both local dependency perception and global time series modeling, effectively improving the information expression and comprehension capabilities of epidemic-related texts.

[0005] While some current research explores the use of pre-trained language models or graph neural networks for agricultural information extraction, there is still a lack of integrated solutions for entity recognition, relationship extraction, and knowledge graph construction, making it difficult to build a complete closed loop of intelligent diagnosis and knowledge services. Therefore, there is an urgent need to develop a high-precision, layer-laying chicken disease detection method that integrates multiple structural advantages to achieve automated recognition, structured expression, and visual presentation of layer-laying chicken disease information, thereby improving the intelligent and systematic level of poultry disease prevention and control. Summary of the Invention

[0006] In response to the above problems, this paper takes laying hens as the research object. In order to achieve efficient and accurate laying hen disease detection, a laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN is designed. The method can use the organized text database to give the diagnosed disease name and prevention and control recommendations.

[0007] The present invention is achieved by adopting the following technical solutions: A laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN includes the following steps: Step S1, dataset construction: extract relevant data from unstructured text data, build a high-quality structured text database, and perform named entity annotation on the corpus of the text database; Step S2, named entity recognition: Use the DistilBERT model to fine-tune the annotated corpus to achieve accurate recognition of the target entity and output standardized entity categories and text; Step S3, relation extraction: extracting semantic relations between entities and constructing disease triple data; Step S4, entity reference resolution and alignment: By introducing a cross-sentence reference resolution mechanism and an entity normalization algorithm, the semantic reference is unified to ensure the consistency and uniqueness of the node semantics in the knowledge graph; Step S5, knowledge graph construction and visualization application: store the constructed triple data in the graph database to complete the construction of the domain-specific knowledge graph, and realize the visualization presentation of entity nodes, relationship edges and query paths by configuring visualization components.

[0008] Furthermore, step S1 includes the following steps: Step S11, collecting and organizing unstructured text data related to laying hen diseases, the data sources include authoritative animal husbandry books, scientific research papers, clinical diagnosis and treatment records and online knowledge platforms, and constructing a preliminary data corpus; Step S12: cleaning and preprocessing the original data to generate a high-quality structured text dataset; In step S13, a word segmentation tool is used to segment the text into sentences and words, and named entities are annotated in BIO format. A two-person cross-annotation method is used, combined with physical book comparison and expert review to ensure that entity boundaries are accurate and categories are consistent, improve accuracy and consistency, and form high-quality training corpus.

[0009] Furthermore, step S2 includes the following steps: Step S21: Input the annotated text sample into the lightweight pre-trained language model DistilBERT at the sentence granularity, perform feature encoding, and generate the context representation vector of each token. x i , as the basic features for subsequent entity recognition: , in, x i For the i The contextual semantic vector of each word; Step S22: Identify important entity information related to laying hen disease in the text, extract entity information using context modeling capabilities, and use a pre-training and fine-tuning strategy to accurately identify the target entity. Step S23, using precision, recall, and F1-score to evaluate the recognition performance and verify the entity extraction capability of the model; Step S24: Output standardized entity categories and text to provide basic data for relationship extraction and knowledge graph construction.

[0010] Furthermore, step S3 includes the following steps: Step S31: Based on the standardized entity categories and texts in step S2, a lightweight pre-trained language model DistilBERT is used to generate entity and context representations and construct a dependency syntactic graph structure between entities; Step S32: Combined with the dependency syntax graph, the GATv2 graph attention mechanism is introduced to aggregate entity adjacency information and construct a multi-dimensional entity pair semantic representation matrix; Step S33: Introduce the BiLSTM-TCN temporal modeling mechanism and input the entity pair context representation obtained by GATv2 into a temporal modeling structure consisting of two BiLSTM layers and one TCN layer in series. Step S34: The multi-dimensional time series representation output by BiLSTM-TCN is subjected to relationship classification through a maximum pooling layer and a Softmax classifier, and a standard triple structure is output to realize a closed loop of semantic relationship extraction; Furthermore, step S31 includes the following steps: Step S311: Based on the entity categories and texts standardized in step S2 and in combination with actual diagnostic requirements, the relationship types required to be identified in the core field of laying hen diseases are clarified, and a semantic relationship category set with guiding significance is constructed; Step S312: Based on the semantic relationship category set, a preliminary relationship annotation corpus is constructed. By combining the entity pairs identified in the text with the contextual semantics, the relationship labels are automatically matched using semantic templates to form training samples for supervised learning. In step S313, the annotated corpus is input into the lightweight encoder DistilBERT, which encodes the input text, generates entity and context representations, and constructs a dependency syntactic graph structure between entities.

[0011] Furthermore, step S32 includes the following steps: The graph attention mechanism GATv2 is introduced. Based on the graph embedding output by the GATv2 layer, the structural semantic representation of all candidate entity pairs is obtained. The semantic representation, position encoding, and entity type features are integrated to construct a multi-dimensional entity pair semantic representation matrix, providing semantically rich and structurally complete input feature representation for subsequent relationship modeling.

[0012] Furthermore, step S33 includes the following steps: Step S331: Input the entity pair context representation obtained by the GATv2 graph attention network into the BiLSTM network to extract the semantic changes and order information that have an impact on the entity relationship; In step S332, a temporal convolutional network (TCN) with a dilation factor is introduced to extract local semantic patterns of different scales in parallel while maintaining the temporal order.

[0013] Furthermore, step S34 includes the following steps: The multi-dimensional time series representation output by BiLSTM-TCN is input into the maximum pooling layer for feature compression to extract the most discriminative semantic features. It is then input into the relation classifier composed of the fully connected layer and the Softmax activation function to complete the classification task of the semantic relationship between entity pairs and output the standard "head entity-relationship-tail entity" triplet in the form of: , in, h is the head entity, t is the tail entity, r For the relationship between the two.

[0014] Furthermore, step S4 includes the following steps: Step S41: DistilBERT is used to obtain the full sentence-level semantic representation, and the original entity pointed to by the pronoun is automatically identified by combining syntactic position clues and entity semantic similarity; In step S42, an entity alignment model is constructed by combining context embedding, entity word vectors, and synonym expansion rules, semantic normalization is performed on entities with different expressions but synonyms, and a unique identification code is generated for each entity.

[0015] Furthermore, step S5 includes the following steps: Step S51: Import the triple data extracted and processed from the text into the graph database Neo4j to clarify the connection between entity nodes and relationship edges, ensuring that the data storage is efficient and scalable to support dynamic updates and queries of large-scale knowledge graphs; Step S52: Configuring a visualization component of the graph database to display entity nodes, relationship edges, and query paths through a graphical interface. This allows users to interactively query and visually browse the graph data, helping them to more intuitively understand the relationship between disease symptoms and prevention and control recommendations. In step S53, by building a graph database and visualization platform, support is provided for subsequent scenarios such as intelligent question-answering systems and diagnostic assistance systems, thereby improving diagnostic efficiency and accuracy.

[0016] Compared with the prior art, the present invention has the following beneficial effects: (1) This paper introduces the DistilBERT model and the GATv2 module to construct a dependency grammar structure graph, strengthens the aggregation ability of key clues in the semantic path between entities, improves the modeling robustness and generalization performance in situations such as complex sentence structures and long entity distances, and significantly improves the accuracy of semantic relationship recognition; (2) The present invention introduces a hybrid time series modeling module composed of BiLSTM and TCN, which not only retains the context-dependent structure but also extracts a multi-scale semantic model through dilated convolution, thereby enhancing the model's ability to identify fuzzy relationship boundary areas and word order coupling phenomena, and improving the overall modeling depth and expression stability; (3) The present invention points out the mechanism of reference resolution and entity alignment. By jointly modeling syntactic position and semantic similarity, it effectively restores fuzzy references such as "this disease" and "this disease", and uniformly identifies synonymous entities such as "pasteurellosis" and "fowl cholera", ensuring the consistency and accuracy of entity nodes in the knowledge graph and reducing the risk of mislabeling due to inconsistent expressions; (4) The present invention ultimately stores the extracted and normalized triples in a unified manner in a graph database, completing the construction of a knowledge graph of laying hen diseases, and presenting the graph structure and semantic path through a visualization component. It has good interactivity and application expansion capabilities, and can be widely used in intelligent question-answering, auxiliary diagnosis, and other breeding scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Fig. 1 It is an overall flow chart of the technical solution of the present invention.

[0018] Fig. 2 This is a model architecture diagram of the present invention.

[0019] Fig. 3 The overall flow chart constructed for the knowledge graph of invention. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the examples of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0021] This paper takes laying hens as the research object. To achieve efficient and accurate laying hen disease detection, it provides a laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN. It realizes the use of a collated text database to provide the diagnosis of disease names and prevention and control recommendations. This method can efficiently detect and diagnose the types of laying hen diseases, improve the accuracy of the intelligent monitoring system, reduce manual intervention, and improve the automation level of laying hen disease prevention and control.

[0022] The present invention first constructs a high-quality text database, organizes the knowledge base of literature, books, disease information and other knowledge in the field of laying hen breeding through crawling and manual sorting methods, and generates the standard input required for structured modeling after cleaning and preprocessing. The DistilBERT model is used as the semantic encoder to semantically encode text data, and a six-layer Transformer structure is used to extract the corresponding word vector representation of the context to capture the deep semantic features of the entity; an inter-word dependency graph with words as nodes is constructed, and a three-layer GATv2 graph attention mechanism is introduced to model the relationship between words, further enhancing the representation of semantic dependencies between entities; the structural semantic representation output by GATv2 is input into a temporal modeling structure consisting of two layers of BiLSTM and one layer of TCN in series. BiLSTM is used to capture the bidirectional context information of the text, and TCN models local and long-range dependencies through dilated convolution. The two jointly improve the expression effect of temporal features and semantic patterns; the fusion representation outputs named entity recognition and relation extraction branches respectively. In the named entity recognition branch, a sequence labeling classifier is constructed through a fully connected layer and a Softmax layer to achieve label prediction. In the relation extraction branch, the corresponding semantic vector is extracted through a maximum pooling operation and input into a relation classifier composed of a fully connected layer and a Softmax activation function to determine the relationship type of the candidate entity pair, output standard triples, and construct a structured laying hen disease knowledge graph.

[0023] like Figs. 1-3As shown, a laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN described in one embodiment of the present invention includes the following steps: Step S1, dataset construction: extract relevant data from unstructured text data, build a high-quality structured text database, and perform named entity annotation on the corpus of the text database; Specifically, step S1 includes the following steps: Step S11, collecting and organizing unstructured text data related to laying hen diseases, the data sources include authoritative animal husbandry books, scientific research papers, clinical diagnosis and treatment records and online knowledge platforms, and constructing a preliminary data corpus; Step S12: Clean and preprocess the raw data, including removing garbled characters, filtering duplicates, and removing irrelevant content, to generate a high-quality structured text dataset. Step S13: Use a word segmentation tool to segment the text and perform named entity annotation in BIO format. Then, through a two-person cross-annotation method, combined with physical book comparison and expert review, ensure that the entity boundaries are accurate and the categories are consistent, improve accuracy and consistency, and form high-quality training corpus. Step S2, named entity recognition: Use the DistilBERT model to fine-tune the annotated corpus to achieve accurate recognition of the target entity and output standardized entity categories and text; Specifically, step S2 includes the following steps: Step 21: Input the annotated text sample into the lightweight pre-trained language model DistilBERT at the sentence granularity, perform feature encoding, and generate the context representation vector of each token. x i , as the basic features for subsequent entity recognition: , in, x i For the i The contextual semantic vector of each word; Step 22: Identify important entity information related to laying hen diseases in the text. This includes disease names (e.g., Newcastle disease, infectious bronchitis), clinical symptoms (e.g., decreased feed intake, ruffled feathers), prevention and control recommendations, pathogen types, drug names, and transmission methods. Contextual modeling is used to extract entity information. The model uses a pre-training and fine-tuning strategy to accurately identify target entities. Step 23: Use precision, recall, and F1-score to evaluate the recognition performance and verify the entity extraction capability of the model. Step 24, output the standardized entity category and text, provide basic data for relationship extraction and knowledge graph construction; Step S3, relationship extraction: realize the extraction of semantic relationship between entities, and construct disease triple data; Specifically, in step S31, based on the standardized entity category and text of step S2, a lightweight pre-training language model DistilBERT is used to generate entity and context representation, and a dependency syntax graph structure between entities is constructed; Further, step S31 includes the following steps: Step S311, based on the standardized entity category and text of step S2, in combination with the actual diagnosis requirements, the relationship types required to be recognized in the layer of egg chicken epidemic disease are determined, and a set of semantic relationship categories with guiding significance is constructed; The relationship covers "disease name-disease symptom" (causing relationship), "disease name-prevention and treatment suggestion" (suggestion relationship), "disease symptom-prevention and treatment suggestion" (treatment relationship) and the like; Step S312, according to the set of semantic relationship categories, a preliminary relationship annotation corpus is constructed, the entity pairs and context semantics recognized in the text are combined, and the relationship labels are automatically matched and annotated through semantic templates to form training samples for supervised learning; Step S313, input the annotated corpus into the lightweight encoder DistilBERT, encode the input text, generate entity and context representation, and construct the dependency syntax graph structure between entities; Step S32, in combination with the dependency syntax graph, the GATv2 graph attention mechanism is introduced to aggregate the entity adjacency information; Further, step S32 includes the following steps: The graph attention mechanism GATv2 is introduced to aggregate the semantic information of entities and their adjacent nodes, to strengthen the weight of the syntax key path in the modeling of the relationship between entities, and to improve the robustness and generalization ability of the model in the scene of complex sentence structure and large entity semantic span. On the basis of the graph embedding output by the GATv2 layer, the structural semantic representation of all candidate entity pairs is obtained, the semantic representation, position encoding and entity type features are fused, and a multi-dimensional entity pair semantic representation matrix is constructed, to provide semantic-rich and structurally-complete input feature representation for subsequent relationship modeling; Step S33, introduce BiLSTM-TCN time series modeling mechanism, input the context representation of entity pairs obtained by GATv2 into the time series modeling structure of two layers of BiLSTM and one layer of TCN in series; Further, step S33 includes the following steps: Step 331: Input the entity pair context representation obtained by the GATv2 graph attention network into the BiLSTM network to extract the semantic changes and order information that affect the entity relationship, thereby improving the model's ability to model the semantics of span relationships and long-range dependencies. Step 332: Introduce a temporal convolutional network (TCN) with a dilation factor to extract local semantic patterns at different scales in parallel while maintaining temporal order. This enhances the model's sensitivity to special word orders and phrase combinations in the context, thereby further improving the model's ability to understand relationship boundaries and contextual semantics. Step S34: The multi-dimensional time series representation output by BiLSTM-TCN is subjected to relationship classification through a maximum pooling layer and a Softmax classifier, and a standard triple structure is output to realize a closed loop of semantic relationship extraction; Furthermore, step S34 includes the following steps: The multi-dimensional time series representation output by BiLSTM-TCN is input into the maximum pooling layer for feature compression to extract the most discriminative semantic features. It is then input into the relation classifier composed of the fully connected layer and the Softmax activation function to complete the classification task of the semantic relationship between entity pairs and output the standard "head entity-relationship-tail entity" triplet to realize the closed loop of semantic relationship extraction. Its form is: , in, h is the head entity, t is the tail entity, r For the relationship between the two; S4, Entity Coding Resolution and Alignment: To address entity coding and synonymous variants that appear in actual texts, we introduce a cross-sentence coding resolution mechanism and entity normalization algorithm to unify semantic references and ensure the consistency and uniqueness of node semantics in the knowledge graph. Specifically, step S4 includes the following steps: Step S41: DistilBERT is used to obtain a full-sentence semantic representation. Combining syntactic positional clues with entity semantic similarity, the original entity pointed to by the pronoun (e.g., "this disease" or "this condition") is automatically identified. This restores entity references within and across sentences, effectively reducing the risk of mislabeling ambiguous entities. Step S42: Build an entity alignment model by combining context embedding, entity word vectors, and synonym expansion rules. Perform semantic normalization on entities with different expressions but synonyms (e.g., "fowl cholera" and "pasteurellosis"), and generate a unique identification code for each entity to ensure node uniqueness and relationship accuracy in subsequent knowledge graph construction. Step S5, knowledge graph construction and visualization application: store the constructed triple data in the graph database to complete the construction of the domain-specific knowledge graph. The overall flow chart of knowledge graph construction is as follows: Fig. 3 As shown in the figure, by configuring visualization components, the entity nodes, relationship edges and query paths can be visualized to support subsequent scenarios such as intelligent question answering and diagnosis assistance. Specifically, step S5 includes the following steps: Step S51: Import the triple data extracted and processed from the text into the graph database Neo4j to clarify the connection between entity nodes (disease name, disease symptoms, prevention and treatment recommendations, etc.) and relationship edges (cause, treatment, etc.), ensuring that the data storage is efficient and scalable to support dynamic updates and queries of large-scale knowledge graphs; Step S52: Configuring a visualization component of the graph database to display entity nodes, relationship edges, and query paths through a graphical interface. This allows users to interactively query and visually browse the graph data, helping them to more intuitively understand the relationship between disease symptoms and prevention and control recommendations. In step S53, by building a graph database and visualization platform, support is provided for subsequent scenarios such as intelligent question-answering systems and diagnostic assistance systems, thereby improving diagnostic efficiency and accuracy.

[0024] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN, characterized in that: The following steps are involved: Step S1, dataset construction: extract relevant data from unstructured text data, build a high-quality structured text database, and perform named entity annotation on the corpus of the text database; Step S2, named entity recognition: Use the DistilBERT model to fine-tune the annotated corpus to achieve accurate recognition of the target entity and output standardized entity categories and text; Step S3, relation extraction: extracting semantic relations between entities and constructing disease triple data; Step S4, entity reference resolution and alignment: By introducing a cross-sentence reference resolution mechanism and an entity normalization algorithm, the semantic reference is unified to ensure the consistency and uniqueness of the node semantics in the knowledge graph; Step S5, knowledge graph construction and visualization application: store the constructed triple data in the graph database to complete the construction of the domain-specific knowledge graph, and realize the visualization presentation of entity nodes, relationship edges and query paths by configuring visualization components.

2. The laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN according to claim 1 is characterized in that: Step S1 includes the following steps: Step S11, collecting and organizing unstructured text data related to laying hen diseases, the data sources include authoritative animal husbandry books, scientific research papers, clinical diagnosis and treatment records and online knowledge platforms, and constructing a preliminary data corpus; Step S12: cleaning and preprocessing the original data to generate a high-quality structured text dataset; In step S13, a word segmentation tool is used to segment the text into sentences and words, and named entities are annotated in BIO format. A two-person cross-annotation method is used, combined with physical book comparison and expert review to ensure that entity boundaries are accurate and categories are consistent, improve accuracy and consistency, and form high-quality training corpus.

3. The laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN according to claim 1 is characterized in that: Step S2 includes the following steps: Step S21: Input the annotated text sample into the lightweight pre-trained language model DistilBERT at the sentence granularity, perform feature encoding, and generate the context representation vector of each token. x i , as the basic features for subsequent entity recognition: , in, x i For the i The contextual semantic vector of each word; Step S22: Identify important entity information related to laying hen disease in the text, extract entity information using context modeling capabilities, and use a pre-training and fine-tuning strategy to accurately identify the target entity. Step S23, using precision, recall, and F1-score to evaluate the recognition performance and verify the entity extraction capability of the model; Step S24: Output standardized entity categories and text to provide basic data for relationship extraction and knowledge graph construction.

4. The laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: Based on the standardized entity categories and texts in step S2, a lightweight pre-trained language model DistilBERT is used to generate entity and context representations and construct a dependency syntactic graph structure between entities; Step S32: Combined with the dependency syntax graph, the GATv2 graph attention mechanism is introduced to aggregate entity adjacency information and construct a multi-dimensional entity pair semantic representation matrix; Step S33: Introduce the BiLSTM-TCN temporal modeling mechanism and input the entity pair context representation obtained by GATv2 into a temporal modeling structure consisting of two BiLSTM layers and one TCN layer in series. In step S34, the multi-dimensional time series representation output by BiLSTM-TCN is subjected to relationship classification through the maximum pooling layer and the Softmax classifier, and a standard triple structure is output to realize the closed loop of semantic relationship extraction.

5. The laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN according to claim 4 is characterized in that: Step S31 includes the following steps: Step S311: Based on the entity categories and texts standardized in step S2 and in combination with actual diagnostic requirements, the relationship types required to be identified in the core field of laying hen diseases are clarified, and a semantic relationship category set with guiding significance is constructed; Step S312: Based on the semantic relationship category set, a preliminary relationship annotation corpus is constructed. By combining the entity pairs identified in the text with the contextual semantics, the relationship labels are automatically matched using semantic templates to form training samples for supervised learning. In step S313, the annotated corpus is input into the lightweight encoder DistilBERT, which encodes the input text, generates entity and context representations, and constructs a dependency syntactic graph structure between entities.

6. The laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN according to claim 4 is characterized in that: Step S32 includes the following steps: The graph attention mechanism GATv2 is introduced. Based on the graph embedding output by the GATv2 layer, the structural semantic representation of all candidate entity pairs is obtained. The semantic representation, position encoding, and entity type features are integrated to construct a multi-dimensional entity pair semantic representation matrix, providing semantically rich and structurally complete input feature representation for subsequent relationship modeling.

7. According to claim 4, a laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN is characterized in that: Step S33 includes the following steps: Step S331: Input the entity pair context representation obtained by the GATv2 graph attention network into the BiLSTM network to extract the semantic changes and order information that have an impact on the entity relationship; In step S332, a temporal convolutional network (TCN) with a dilation factor is introduced to extract local semantic patterns of different scales in parallel while maintaining the temporal order.

8. According to claim 4, a laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN is characterized in that: Step S34 includes the following steps: The multi-dimensional time series representation output by BiLSTM-TCN is input into the maximum pooling layer for feature compression to extract the most discriminative semantic features. It is then input into the relation classifier composed of the fully connected layer and the Softmax activation function to complete the classification task of the semantic relationship between entity pairs and output the standard "head entity-relationship-tail entity" triplet in the form of: , in, h is the head entity, t is the tail entity, r For the relationship between the two.

9. According to claim 1, a laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN is characterized in that: Step S4 includes the following steps: Step S41: DistilBERT is used to obtain the full sentence-level semantic representation, and the original entity pointed to by the pronoun is automatically identified by combining syntactic position clues and entity semantic similarity; In step S42, an entity alignment model is constructed by combining context embedding, entity word vectors, and synonym expansion rules, semantic normalization is performed on entities with different expressions but synonyms, and a unique identification code is generated for each entity.

10. A laying hen disease detection method based on DistilBERT-GATv2 and BiLSTM-TCN according to claim 1, characterized in that: Step S5 includes the following steps: Step S51: Import the triple data extracted and processed from the text into the graph database Neo4j to clarify the connection between entity nodes and relationship edges, ensuring that the data storage is efficient and scalable to support dynamic updates and queries of large-scale knowledge graphs; Step S52: Configuring a visualization component of the graph database to display entity nodes, relationship edges, and query paths through a graphical interface. This allows users to interactively query and visually browse the graph data, helping them to more intuitively understand the relationship between disease symptoms and prevention and control recommendations. In step S53, by building a graph database and visualization platform, support is provided for subsequent scenarios such as intelligent question-answering systems and diagnostic assistance systems, thereby improving diagnostic efficiency and accuracy.