Cow breeding information management system
By designing a dairy cattle breeding information management system, the keywords and values in the breeding records are automatically identified and stored, and the problem of low storage efficiency in the prior art is solved, efficient storage and query are achieved, and accurate milk production prediction and cow recommendation are provided.
Patent Information
- Application Number
- CN202510574915.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-15
AI Technical Summary
The existing dairy cattle breeding information is inefficient in storing information, and it is necessary to manually split the breeding record text for storage, which is inefficient.
Design a dairy cattle breeding information management system, including an information storage unit and an information retrieval unit, automatically identify keywords and their values in the breeding record text, match and store them, and use the Biaffine syntax analyzer and deep learning model for data processing and prediction.
It realizes efficient storage and query of dairy cattle breeding information without manual splitting, improves storage efficiency, and improves query efficiency through keyword matching and milk production prediction. The recommended dairy cattle are more in line with the needs of breeding technicians.
Smart Images

Figure CN120493916A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic information management, in particular to a dairy cow breeding information management system. Background Art
[0002] Breeding information is fundamental to dairy farm management. It provides detailed records of each cow's growth, reproduction, production, and health. This data allows ranchers to accurately assess the overall condition of their herd and promptly identify and address any issues. Breeding records are also crucial for genetic improvement and the selection of high-quality offspring. Analysis of these records allows identification of high-performing individuals, providing strong support for future breeding plans.
[0003] In their daily work, breeding technicians typically record cow information in a text file called a breeding record. However, when they need to store this information, they must manually separate the breeding information and its corresponding values from the text file before storing it. This storage method is extremely inefficient. Summary of the Invention
[0004] The purpose of the present invention is to provide a dairy cow breeding information management system to solve the problem of low storage efficiency of existing dairy cow breeding information.
[0005] The technical solution adopted by the present invention to solve the above technical problems is:
[0006] A dairy cow breeding information management system, the system comprising an information storage unit and an information retrieval unit;
[0007] The information storage unit specifically performs the following steps:
[0008] Step 11: receiving the cow breeding record text input by the user;
[0009] Step 12: Perform keyword recognition on the dairy cow breeding record text to obtain the keywords and their corresponding values;
[0010] Step 13: Determine the cow number based on the cow breeding record text;
[0011] Step 14: According to the cow number, retrieve the cow breeding records stored in the database and obtain the corresponding keywords and values;
[0012] Step 15: The keyword in step 14 is matched with the keyword in step 12. If the match is successful, the numerical values corresponding to the successfully matched keywords are averaged, the numerical value of the keyword is updated, and stored in the database. If the match is unsuccessful, the keyword in step 12 is used as a new keyword, and the keyword and its corresponding numerical value are stored in the database.
[0013] The information retrieval unit specifically performs the following steps:
[0014] Step 21: Receive the keyword queried by the user and the numerical range corresponding to the keyword;
[0015] Step 22: Match the keyword input by the user in step 21 with the keyword corresponding to each cow number in the database, and sort the cow numbers from largest to smallest according to the number of successful keyword matches;
[0016] Step 23: For the cow number ranked first, obtain the value of each keyword corresponding to the cow number, determine whether the value of each keyword is within the value range input by the user, and count the number of keywords corresponding to each cow number that are within the value range input by the user. Then, sort the cow numbers from largest to smallest based on the statistical results.
[0017] Step 24: Based on the ranking in step 23, the number of the cow ranked first is used as the query result, that is, the cow that best meets the user's needs.
[0018] Furthermore, the system further includes a cow recommendation module, which specifically performs the following steps:
[0019] Step 31: Based on the cow number obtained by the information retrieval unit, obtain the historical cow breeding record text of the cow number, and use the Biaffine syntax analyzer to obtain a dependency syntax structure diagram of the historical cow breeding record text;
[0020] Step 32: Calculate the relative dependency distances between different keywords using the dependency syntactic structure graph, and obtain the importance weights of different keywords in the historical dairy cow breeding record text relative to the dairy cow based on the relative dependency distances;
[0021] Step 33: Assign a value to each keyword of the cow according to the importance weight of the cow;
[0022] Step 34: Based on the assignment result of step 33, for each cow number in the database, obtain all the corresponding assigned keywords, add all the values, and then determine whether the addition result is greater than the threshold. If it is greater than the threshold, the corresponding cow number is recommended to the user, otherwise it is not recommended.
[0023] Furthermore, the system also includes a milk production prediction unit, which retrieves a keyword corresponding to the cow number according to the cow number obtained by the information retrieval unit, and predicts the milk production of the cow corresponding to the cow number according to the keyword.
[0024] Furthermore, the milk production prediction unit specifically performs the following steps:
[0025] Step 41: Retrieve the keyword and its corresponding value corresponding to the cow number, i.e., the cow's milk production characteristic data, according to the cow number obtained by the information retrieval unit;
[0026] Step 42: After normalizing the cow milk production characteristic data, the data is input into the trained cow milk production prediction model to obtain the predicted cow milk production.
[0027] Furthermore, the system further includes a structural variation detection module, which specifically performs the following steps:
[0028] Step 1: Obtain the BAM file corresponding to the cow that best meets the user's needs. Each BAM file contains M SNP sites.
[0029] Step 2: Obtain all clustering results of all variant type signals in the BAM file of individual dairy cows;
[0030] The signals contained in each cluster are integrated into a structural variation signal for output;
[0031] Structural variation signals include insertions, deletions, duplications, translocations, and inversions;
[0032] Step 3: Perform haplotype typing on each output structural variation signal to obtain the typed haplotype 1 file and haplotype 2 file;
[0033] Step 4: Use the BAM file of the individual dairy cow obtained in step 1 as the input of the deep learning model, and use the haplotype 1 file and haplotype 2 file after typing as the output of the deep learning model;
[0034] The BAM file of the individual cow obtained in step 1 is used as the input of the large language model, and the haplotype 1 file and haplotype 2 file after typing are used as the output of the large language model;
[0035] A comprehensive loss function is used to optimize the parameters of the deep learning model and the large language model, and the Adam optimizer is used to perform gradient updates until the comprehensive loss function converges to obtain the trained deep learning model and the large language model.
[0036] Step 5: Input the BAM file of the individual dairy cow to be tested into the trained deep learning model, and the trained deep learning model outputs the haplotype 1 file and haplotype 2 file after typing the BAM file of the individual dairy cow to be tested;
[0037] In step 2, all cluster results of all variant type signals in the BAM file of the individual dairy cow are obtained;
[0038] The signals contained in each cluster are integrated into a structural variation signal for output;
[0039] Structural variation signals include insertions, deletions, duplications, translocations, and inversions;
[0040] The specific process is:
[0041] Step 21: Obtain the clustering results of all signals in the mutation signal set of the "insert" mutation type. The specific process is as follows:
[0042] Step 211: Initialize an empty cluster and add the first signal in the mutation signal set of the "insert" mutation type as the starting signal;
[0043] Step 212: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0044] Step 213: If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster;
[0045] If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster;
[0046] Step 214: Repeat steps 212 and 213 until all signals in the mutation signal set of the “insertion” mutation type are determined, and clustering results of all signals in the mutation signal set of the “insertion” mutation type are obtained;
[0047] Step 22: Obtain the clustering results of all signals in the mutation signal set of the "deletion" mutation type. The specific process is as follows:
[0048] Step 221: Initialize an empty cluster and add the first signal in the mutation signal set of the “deletion” mutation type as the starting signal;
[0049] Step 222: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0050] Step 223: If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster;
[0051] If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster;
[0052] Step 224: Repeat steps 222 and 223 until all signals in the mutation signal set of the "delete" mutation type are determined;
[0053] Step 23: Obtain the clustering results of all signals in the mutation signal set of the "repeated" mutation type. The specific process is as follows:
[0054] Step 231: Initialize an empty cluster and add the first signal in the set of mutation signals of the "repeat" mutation type as the starting signal;
[0055] Step 232: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0056] Step 233: If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster;
[0057] If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster;
[0058] Step 234: Move to the next signal and repeat steps 232 to 233 until all signals in the set of variant signals of the "repeat" variant type are determined;
[0059] Step 24: Obtain the clustering results of all signals in the mutation signal set of the "translocation" mutation type. The specific process is as follows:
[0060] Step 241: Initialize an empty cluster and add the first signal in the set of mutation signals of the “translocation” mutation type as the starting signal;
[0061] Step 242: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0062] Step 243: If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster;
[0063] If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster;
[0064] Step 244: Repeat steps 242 to 243 until all signals in the mutation signal set of the “translocation” mutation type are determined, and a clustering result of all signals in the mutation signal set of the “translocation” mutation type is obtained;
[0065] Step 25: Obtain the clustering results of all signals in the mutation signal set of the "inversion" mutation type. The specific process is as follows:
[0066] Step 251: Initialize an empty cluster and add the first signal in the set of mutation signals of the “inversion” mutation type as the starting signal;
[0067] Step 252: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0068] Step 253: If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster;
[0069] If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster;
[0070] Step 254: Repeat steps 252 to 253 until all signals in the mutation signal set of the “inversion” mutation type are determined;
[0071] Get the clustering results of all signals in the mutation signal set of "inversion" mutation type;
[0072] The specific process of calculating the similarity between the current signal and the last signal in each existing cluster is as follows:
[0073] For insertion, deletion, inversion, and duplication variation signals, a comprehensive similarity score S is obtained, which is expressed as:
[0074] S=w1×S1+w2×S2+w3×S3
[0075] Among them, w1, w2, and w3 are the weight coefficients of position similarity score, variant size similarity score / chromosome similarity score, respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, w1 + w2 + w3 = 1;
[0076] For translocation mutation signals, a comprehensive similarity score S is obtained, which is expressed as:
[0077] S=w4×S4+w5×S5
[0078] Wherein, w4 is the weight coefficient of the starting chromosome similarity score S4, w4=0.5, w5 is the weight coefficient of the ending chromosome similarity score S5, w5=0.5, w4+w5=1;
[0079] For the insertion, deletion, inversion and duplication variation signals, a comprehensive similarity score S is obtained, which is expressed as:
[0080] S=w1×S1+w2×S2+w3×S3
[0081] Among them, w1, w2, and w3 are the weight coefficients of position similarity score, variant size similarity score / chromosome similarity score, respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, w1 + w2 + w3 = 1;
[0082] The specific process is:
[0083] 1) Calculate the location similarity S1 score. The specific process is as follows:
[0084] The calculation formula for the position similarity S1 score is as follows:
[0085] S1=|Start1-Start2|
[0086] Among them, Start1 and Start2 are the starting positions of the two variant signals respectively;
[0087] 2) Calculate the variation signal size similarity score S2. The specific process is as follows:
[0088] The mutation lengths of the two mutation signals are SVlen1 and SVlen2 respectively. The similarity score S2 of the mutation length of the mutation signal is calculated by the following formula:
[0089]
[0090] 3) Calculate the chromosome similarity score S3. The specific process is as follows:
[0091] Assume that the chromosome numbers of the two mutation signals are Chrom1 and Chrom2 respectively, and the chromosome similarity score S3 is calculated by the following discrete function:
[0092]
[0093] When the chromosome names of the two variant signals are the same, the chromosome similarity score is 1;
[0094] When the chromosome names of the two variant signals are different, the chromosome similarity score is 0;
[0095] 4) Perform weighted summation of the position similarity S1 score, the variation size similarity S2 score, and the chromosome similarity S3 score to obtain a comprehensive similarity score S, which is expressed as:
[0096] S=w1×S1+w2×S2+w3×S3
[0097] Among them, w1, w2, and w3 are the weight coefficients of position similarity score, variant size similarity score / chromosome similarity score, respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, w1 + w2 + w3 = 1;
[0098] For the translocation variation signal, a comprehensive similarity score S is obtained, which is expressed as:
[0099] S=w4×S4+w5×S5
[0100] Wherein, w4 is the weight coefficient of the starting chromosome similarity score S4, w4=0.5, w5 is the weight coefficient of the ending chromosome similarity score S5, w5=0.5, w4+w5=1;
[0101] The specific process is:
[0102] Calculate the starting chromosome similarity score:
[0103]
[0104] Among them, S4 represents the starting chromosome similarity score;
[0105] Indicates the name of the starting chromosome of the first signal;
[0106] Indicates the name of the starting chromosome of the second signal;
[0107] Calculate the terminating chromosome similarity score:
[0108]
[0109] Among them, S5 represents the termination chromosome similarity score;
[0110] The name of the chromosome that terminates the first signal;
[0111] The name of the chromosome that terminates the second signal;
[0112] The weighted sum of the starting chromosome similarity score S4 and the ending chromosome similarity score S5 is used to obtain a comprehensive similarity score S, which is expressed as:
[0113] S=w4×S4+w5×S5
[0114] Wherein, w4 is the weight coefficient of the starting chromosome similarity score S4, w4=0.5, w5 is the weight coefficient of the ending chromosome similarity score S5, w5=0.5, w4+w5=1;
[0115] In S3, each output structural variation signal is haplotyped to obtain the hap1 file and hap2 file after typing. The specific process is as follows:
[0116] Step 31: Input the VCF file of each structural variation signal into the variation detection tool;
[0117] The variant detection tool outputs a typing VCF file, which contains M SNP sites;
[0118] The VCF file of the classification contains 8 columns of information, namely:
[0119] Chromosome name, SNP position, SNP ID, reference gene, alternative allele, quality value, filter flag, annotation information column;
[0120] Step 32: Use Bcftools software to filter the typing VCF file to obtain a filtered typing VCF file;
[0121] Step 33: Input the filtered typing VCF file obtained in step 22 into the WhatsHap typing tool, which outputs two typing files, namely, haplotype 1 file and haplotype 2 file;
[0122] In step 32, the typing VCF file is filtered using Bcftools software to obtain a filtered typing VCF file. The specific process is as follows:
[0123] Keep the rows in the VCF file with the "Filter Flag" value equal to PASS;
[0124] Delete the rows in the typing VCF file where the "Filter Flag" value is not equal to PASS;
[0125] Finally, the typing VCF file is generated;
[0126] Each SNP site in the typing VCF file is divided into haplotype 1 and haplotype 2;
[0127] In step 33, the filtered typing VCF file obtained in step 22 is input into the WhatsHap typing tool, and the WhatsHap typing tool outputs two typing files, namely, haplotype 1 file and haplotype 2 file. The specific process is as follows:
[0128] Step 331: Set Modki software parameters: pileup, traditional;
[0129] Use Modki software to convert the filtered typing VCF file obtained in step 22 into a tsv format file;
[0130] Step 332: Use the WhatsHap typing tool to generate a BED.gz format file from the tsv format file;
[0131] Step 333: Input the typing VCF file obtained in step 22 and the BED.gz format file obtained in step 332 into the WhatsHap typing tool, and the WhatsHap typing tool outputs a haplotype 1 file and a haplotype 2 file;
[0132] In step 4, the BAM file of the individual dairy cow obtained in step 1 is used as the input of the deep learning model, and the haplotype 1 file and haplotype 2 file after typing are used as the output of the deep learning model;
[0133] The BAM file of the individual cow obtained in step 1 is used as the input of the large language model, and the haplotype 1 file and haplotype 2 file after typing are used as the output of the large language model;
[0134] A comprehensive loss function is used to optimize the parameters of the deep learning model and the large language model, and the Adam optimizer is used to perform gradient updates until the comprehensive loss function converges to obtain the trained deep learning model and the large language model.
[0135] The specific process is:
[0136] Step 41: Build a deep learning model. The deep learning model includes:
[0137] The first 1×1 convolution layer, BN layer, the first 7×7 depth convolution layer, the first LN layer, the second 1×1 convolution layer, the first GELU, the first GRN, the third 1×1 convolution layer, the second 7×7 depth convolution layer, the second LN layer, the fourth 1×1 convolution layer, the second GELU, the second GRN, the fifth 1×1 convolution layer, the third 7×7 depth convolution layer, the third LN layer, the sixth 1×1 convolution layer, the third GELU, the third GRN, the seventh 1×1 convolution layer, the fully connected layer, and the softmax layer.
[0138] The working process of the deep learning model is as follows:
[0139] The BAM file of the individual cow obtained in step 1 is input into the first 1×1 convolutional layer and the BN layer in sequence. The BN layer outputs feature A.
[0140] The output feature A of the BN layer is sequentially input into the first 7×7 depth convolution layer, the first LN layer, the second 1×1 convolution layer, the first GELU, the first GRN, and the third 1×1 convolution layer. The third 1×1 convolution layer outputs feature A′;
[0141] The output feature A′ of the third 1×1 convolutional layer is added element-by-element to the output feature A of the BN layer to obtain feature A″;
[0142] Feature A″ is sequentially input into the second 7×7 depth convolution layer, the second LN layer, the fourth 1×1 convolution layer, the second GELU, the second GRN, and the fifth 1×1 convolution layer. The fifth 1×1 convolution layer outputs feature A″′;
[0143] The fifth 1×1 convolutional layer outputs feature A″′ and feature A″ and adds them element by element to obtain feature
[0144] feature Input the third 7×7 depth convolution layer, the third LN layer, the sixth 1×1 convolution layer, the third GELU, the third GRN, the seventh 1×1 convolution layer, and the seventh 1×1 convolution layer output features.
[0145] The seventh 1×1 convolutional layer outputs features and features Perform element-by-element summation to obtain feature B;
[0146] Feature B is input into the fully connected layer and the softmax layer in sequence, and the softmax layer outputs the classification result;
[0147] Step 42: Generate a haplotype 1 file and a haplotype 2 file corresponding to the BAM file of the individual dairy cow obtained in step 1 using the large language model;
[0148] Step 43: Optimize the parameters of the deep learning model and the large language model using a comprehensive loss function, and perform gradient updates in combination with the Adam optimizer until the comprehensive loss function converges, thereby obtaining the trained deep learning model and the large language model.
[0149] The comprehensive loss function is
[0150] Where N represents the total number of data in the BAM file of individual cows, i represents the i-th data, and k represents the k-th data;
[0151] F i 1 The i-th data in the BAM file representing an individual dairy cow is input into the deep learning model, and the features output by the deep learning model;
[0152] F i 2 The i-th data in the BAM file representing the individual cow is input into the large language model, and the large language model outputs the semantic features of the text description information;
[0153] The kth data in the BAM file representing the individual cow is input into the large language model, and the large language model outputs the semantic features of the text description information;
[0154] s(F i 1 ,F i 2 ) represents the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the i-th data;
[0155] Represents the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the k-th data;
[0156] τ represents a temperature hyperparameter.
[0157] Furthermore, the system also includes a DHI report generation module, which is used to receive user instructions, obtain the DHI data of the ranch after receiving the user's instructions, and then interpret the DHI report using the completed knowledge graph. The completed knowledge graph is obtained by the following steps:
[0158] First, we obtain the benchmark DHI domain knowledge graph and record the edge weights in the benchmark DHI domain knowledge graph as the basic edge weight ω m1 , and then complete it based on the benchmark DHI domain knowledge graph, including the following steps:
[0159] S201. Take all entities in the benchmark DHI domain knowledge graph as reference entities, extract entities and entity relationships in all texts in the incremental text library based on the language big model, and predict the tail entity through the language big model to obtain the predicted triple (h, r, t), where h and t are used to represent the head entity and the tail entity in the triple, and r is used to represent the head entity relationship in the triple. Based on the predicted triples, the benchmark DHI domain knowledge graph structure is completed to obtain a complete DHI domain knowledge graph structure, which is recorded as the full DHI domain knowledge graph. The difference between the full DHI domain knowledge graph and the benchmark DHI domain knowledge graph is recorded as the incremental DHI domain knowledge graph;
[0160] The neighbor nodes of a node in the benchmark DHI domain knowledge graph are recorded as basic connection nodes, and the corresponding edges are basic edges. The neighbor nodes in the incremental DHI domain knowledge graph are incremental connection nodes, and the corresponding edges are incremental edges. For a node, calculate the proportion k of the number of corresponding incremental edges to the total number of edges. m , as the incremental adjustment coefficient, 1-k m As the basic adjustment coefficient;
[0161] S202, based on the two entities of "influencing factors" and "performance indicators / symptoms" in the DHI knowledge graph, count the non-repeating triples (h, r, t) corresponding to all entities in the incremental text library, and among all the non-repeating triples (h, r, t), for a certain entity v m , will contain entity v m The triplet of statistics The number n m , and then based on n m Get entity v m The corresponding incremental edge weight ω m2 ;
[0162] S203, correct the basic edge weight in the process of building the benchmark DHI domain knowledge graph to (1-k m )ω m1, and correct the incremental edge weight to k m ω m2 , and then complete the knowledge graph of the benchmark DHI domain;
[0163] The construction process of the benchmark DHI domain knowledge graph includes:
[0164] (1) Constructing the DHI domain ontology, which includes three types of entities and entity relationships: "performance indicators / symptoms", "influencing factors", and "solutions". The performance indicators / symptoms refer to the indicators used to present the health status of dairy cows or the symptoms manifested by dairy cows;
[0165] (2) The electronic text obtained after the DHI measurement and application guidance is digitized is used as the annotation object, and the ontology is used as the annotation basis to perform semantic annotation on the electronic text data to form annotation data;
[0166] (3) Using the data in the labeled data as training data, according to the ontology structure of the DHI domain knowledge graph, extract entities and entity relationships from the text of the basic text library to obtain entity and entity relationship data, and construct triples of any two types of entities and entity relationships, as well as the DHI domain knowledge graph, which is recorded as the baseline DHI domain knowledge graph;
[0167] (4) For the triples containing two types of entities in the DHI knowledge graph, “influencing factors” and “performance indicators / symptoms”, the conditional probability between the two types of entities is calculated, denoted as P(fac|sym), and the weight of the edge between the entities is denoted as the basic edge weight ω m1 ;
[0168] The P(fac|sym) is obtained by crowdsourcing calculation;
[0169] Based on n m Get entity v m The corresponding incremental edge weight ω m2 The process includes:
[0170] For an entity v m , the statistical incremental text library contains Number of documents And calculate the entity v m The corresponding incremental edge weight Where N is the number of documents in the incremental text library;
[0171] The knowledge graph adopts the CompGCN network as the network framework;
[0172] The large language model selects the BERT model.
[0173] Furthermore, the keywords include body height, chest width, body depth, udder depth, central suspensory ligament length, front teat position, front teat length, rear udder attachment height, rear teat position, lactation days, parity and rear udder attachment width.
[0174] Furthermore, the keyword recognition is performed by the Chinese keyword extractor Jieba.
[0175] Furthermore, the dependency syntax structure graph is represented in the form of an adjacency matrix D, where each element in D is represented as:
[0176]
[0177] Among them, i represents the row index of the matrix, j represents the column index of the matrix, and w i and w j Represents any two keywords in the dairy cow breeding record text.
[0178] Furthermore, the specific steps of calculating the relative dependency distance using the dependency syntactic structure graph are:
[0179] Based on the adjacency matrix D, the Dijkstra algorithm is used to obtain the relative dependency distance between different keywords through the shortest distance of different keywords on the adjacency matrix.
[0180] The beneficial effects of the present invention are:
[0181] This application does not require manual splitting of breeding information. After the breeding technician enters the dairy cow breeding record text into this application system, the system will split the keywords and store the keywords and their corresponding values, greatly improving storage efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0182] Figure 1 This is the overall structure diagram of this application. DETAILED DESCRIPTION
[0183] It should be noted that, unless there is any conflict, the various embodiments disclosed in this application can be combined with each other.
[0184] Specific embodiment 1: A dairy cow breeding information management system according to this embodiment includes an information storage unit, an information retrieval unit, and a milk production prediction unit;
[0185] The information storage unit specifically performs the following steps:
[0186] Step 11: receiving the cow breeding record text input by the user;
[0187] Step 12: Perform keyword recognition on the dairy cow breeding record text to obtain the keywords and their corresponding values;
[0188] Step 13: Determine the cow number based on the cow breeding record text;
[0189] Step 14: According to the cow number, retrieve the cow breeding records stored in the database and obtain the corresponding keywords and values;
[0190] Step 15: The keyword in step 14 is matched with the keyword in step 12. If the match is successful, the numerical values corresponding to the successfully matched keywords are averaged, the numerical value of the keyword is updated, and stored in the database. If the match is unsuccessful, the keyword in step 12 is used as a new keyword, and the keyword and its corresponding numerical value are stored in the database.
[0191] The information retrieval unit specifically performs the following steps:
[0192] Step 21: Receive the keyword queried by the user and the numerical range corresponding to the keyword;
[0193] Step 22: Match the keyword input by the user in step 21 with the keyword corresponding to each cow number in the database, and sort the cow numbers from largest to smallest according to the number of successful keyword matches;
[0194] Step 23: For the cow number ranked first, obtain the value of each keyword corresponding to the cow number, determine whether the value of each keyword is within the value range input by the user, and count the number of keywords corresponding to each cow number that are within the value range input by the user. Then, sort the cow numbers from largest to smallest based on the statistical results.
[0195] Step 24: Based on the ranking in step 23, the number of the cow ranked first is used as the query result, that is, the cow that best meets the user's needs.
[0196] This application eliminates the need for manual decomposition of breeding information. After a breeding technician enters a dairy cow breeding record text into the application system, the system decomposes the keywords and stores the keywords and their corresponding values, greatly improving storage efficiency. Furthermore, when a technician needs to extract a corresponding keyword and value (for example, a technician needs to obtain a dairy cow with a height of 1.5 to 1.8 meters), this application performs two matches (the first for keywords and the second for values) to obtain the cow number in the database that meets the technician's requirements. This application offers high query efficiency, more than doubling the efficiency of traditional query methods.
[0197] In addition, when storing cow keywords, this application will add new keywords and their corresponding values, and will also update the existing cow keywords. This application will average the historical values and the new values as the new values, so as to reduce the error of the values.
[0198] The system further includes a cow recommendation module, which specifically performs the following steps:
[0199] Step 31: Based on the cow number obtained by the information retrieval unit, obtain the historical cow breeding record text of the cow number, and use the Biaffine syntax analyzer to obtain a dependency syntax structure diagram of the historical cow breeding record text;
[0200] Step 32: Calculate the relative dependency distances between different keywords using the dependency syntactic structure graph, and obtain the importance weights of different keywords in the historical dairy cow breeding record text relative to the dairy cow based on the relative dependency distances;
[0201] Step 33: Assign a value to each keyword of the cow according to the importance weight of the cow;
[0202] Step 34: Based on the assignment result of step 33, for each cow number in the database, obtain all the corresponding assigned keywords, add all the values, and then determine whether the addition result is greater than the threshold. If it is greater than the threshold, the corresponding cow number is recommended to the user, otherwise it is not recommended.
[0203] This implementation assigns different weights to each keyword in the user's query results based on relative dependency distance. Based on these weights, the remaining cows are assigned a value, and cows with a score greater than a threshold are recommended to the user. The cows recommended in this application meet multiple user requirements, and each one is more closely aligned with the breeding technician's emphasis on that keyword.
[0204] The system further comprises a milk production prediction unit, which retrieves a keyword corresponding to the cow number according to the cow number obtained by the information retrieval unit, and predicts the milk production of the cow corresponding to the cow number according to the keyword.
[0205] The milk production prediction unit specifically performs the following steps:
[0206] Step 41: Retrieve the keyword and its corresponding value corresponding to the cow number, i.e., the cow's milk production characteristic data, according to the cow number obtained by the information retrieval unit;
[0207] Step 42: After normalizing the cow milk production characteristic data, the data is input into the trained cow milk production prediction model to obtain the predicted cow milk production.
[0208] The keywords include body height, chest width, body depth, udder depth, central suspensory ligament length, front teat position, front teat length, udder attachment height, udder position, lactation days, parity and udder attachment width.
[0209] The keyword recognition is performed by the Chinese keyword extractor Jieba.
[0210] The dependency syntax structure graph is represented in the form of an adjacency matrix D, where each element in D is represented as:
[0211]
[0212] Among them, i represents the row index of the matrix, j represents the column index of the matrix, and w i and w j Represents any two keywords in the dairy cow breeding record text.
[0213] The specific steps of calculating the relative dependency distance using the dependency syntactic structure graph are as follows:
[0214] Based on the adjacency matrix D, the Dijkstra algorithm is used to obtain the relative dependency distance between different keywords through the shortest distance of different keywords on the adjacency matrix.
[0215] DHI technology is a complete production record and management system. By measuring the milk production performance data of lactating cows and analyzing the basic information of the herd, it can understand the milk production level, milk composition, somatic cells, etc. of the existing herd and individual cows. It has an early warning effect on udder health and reproduction-related problems, and comprehensively evaluates the production performance and genetic performance of individual cows and herds to identify problems in dairy cow production management and breeding.
[0216] Gao Meng and others from Northeast Agricultural University proposed a "DHI report interpretation method based on knowledge graph (application number CN202110969609.X)". This method combines the DHI domain knowledge graph to diagnose problems based on the results of dynamic analysis. Based on the DHI domain knowledge graph, the DHI domain knowledge graph contains three types of entities and entity relationships: "performance indicators / symptoms", "influencing factors", and "solutions". Therefore, it realizes the automatic interpretation of DHI reports, enabling any dairy herd management unit to obtain scientific and effective guidance; at the same time, it forms a unified standard for DHI report interpretation, enabling different pasture management to form a unified system, ensuring the scientificity and effectiveness of dairy herd breeding. However, its effect is limited.
[0217] The specific process of the structural variation detection module in this application is as follows:
[0218] Step 1: Obtain the BAM file of the individual dairy cow; each BAM file contains M SNP sites
[0219] Step 2: Obtain all clustering results of all variant type signals in the BAM file of individual cows
[0220] The signals contained in each cluster are integrated into a structural variation signal for output;
[0221] Structural variation signals include insertions, deletions, duplications, translocations, and inversions;
[0222] Step 3: Perform haplotype typing on each output structural variation signal to obtain the typed haplotype 1 file and haplotype 2 file;
[0223] Step 4: Use the BAM file of the individual dairy cow obtained in step 1 as the input of the deep learning model, and use the haplotype 1 file and haplotype 2 file after typing as the output of the deep learning model;
[0224] The BAM file of the individual cow obtained in step 1 is used as the input of the large language model, and the haplotype 1 file and haplotype 2 file after typing are used as the output of the large language model;
[0225] A comprehensive loss function is used to optimize the parameters of the deep learning model and the large language model, and the Adam optimizer is used to perform gradient updates until the comprehensive loss function converges to obtain the trained deep learning model and the large language model.
[0226] Step 5: Input the BAM file of the individual dairy cow to be tested into the trained deep learning model, and the trained deep learning model outputs the haplotype 1 file and haplotype 2 file after typing the BAM file of the individual dairy cow to be tested.
[0227] In step 2, all cluster results of all variant type signals in the BAM file of the individual dairy cow are obtained;
[0228] The signals contained in each cluster are integrated into a structural variation signal for output;
[0229] Structural variation signals include insertions, deletions, duplications, translocations, and inversions;
[0230] The specific process is:
[0231] Step 21: Obtain the clustering results of all signals in the mutation signal set of the "insert" mutation type; the specific process is:
[0232] Step 211: Initialize an empty cluster and add the first signal in the mutation signal set of the "insert" mutation type as the starting signal;
[0233] Step 212: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0234] Step 213
[0235] If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster;
[0236] If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster;
[0237] Step 214: Repeat steps 212 to 213 until all signals in the mutation signal set of the “insertion” mutation type are determined; and obtain clustering results for all signals in the mutation signal set of the “insertion” mutation type.
[0238] Step 22: Obtain the clustering results of all signals in the mutation signal set of the "deletion" mutation type; the specific process is:
[0239] Step 221: Initialize an empty cluster and add the first signal in the mutation signal set of the “deletion” mutation type as the starting signal;
[0240] Step 222: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0241] Step 223
[0242] If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster;
[0243] If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster;
[0244] Step 224: Repeat steps 222 and 223 until all signals in the mutation signal set of the "delete" mutation type are determined;
[0245] Step 23: Obtain the clustering results of all signals in the set of mutation signals of the "repeated" mutation type; the specific process is as follows:
[0246] Step 231: Initialize an empty cluster and add the first signal in the set of mutation signals of the "repeat" mutation type as the starting signal;
[0247] Step 232: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0248] Step 233
[0249] If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster;
[0250] If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster;
[0251] Step 234: Move to the next signal and repeat steps 232 to 233 until all signals in the set of variant signals of the "repeat" variant type are determined;
[0252] Step 24: Obtain the clustering results of all signals in the mutation signal set of the "translocation" mutation type; the specific process is as follows:
[0253] Step 241: Initialize an empty cluster and add the first signal in the set of mutation signals of the “translocation” mutation type as the starting signal;
[0254] Step 242: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0255] Step 243
[0256] If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster;
[0257] If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster;
[0258] Step 244: repeat steps 242 to 243 until all signals in the mutation signal set of the “translocation” mutation type are determined; and obtain clustering results of all signals in the mutation signal set of the “translocation” mutation type;
[0259] Step 25: Obtain the clustering results of all signals in the set of mutation signals of the “inversion” mutation type; the specific process is as follows:
[0260] Step 251: Initialize an empty cluster and add the first signal in the set of mutation signals of the “inversion” mutation type as the starting signal;
[0261] Step 252: Calculate the similarity between the current signal and the last signal in each existing cluster;
[0262] Step 253
[0263] If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster;
[0264] If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster;
[0265] Step 254: Repeat steps 252 to 253 until all signals in the mutation signal set of the “inversion” mutation type are determined;
[0266] Get the clustering results of all signals in the mutation signal set of the "inversion" mutation type.
[0267] Other steps and parameters are the same as those in the first embodiment.
[0268] The similarity between the current signal and the last signal in each existing cluster is calculated; the specific process is:
[0269] For insertion, deletion, inversion, and duplication variation signals, a comprehensive similarity score S is obtained; it is expressed as:
[0270] S=w1×S1+w2×S2+w3×S3
[0271] Among them, w1, w2, and w3 are the weight coefficients of position similarity score, variant size similarity score / chromosome similarity score, respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, w1 + w2 + w3 = 1;
[0272] For translocation mutation signals, a comprehensive similarity score S is obtained; it is expressed as:
[0273] S=w4×S4+w5×S5
[0274] in,
[0275] w4 is the weight coefficient of the initial chromosome similarity score S4, w4 = 0.5,
[0276] w5 is the weight coefficient of the termination chromosome similarity score S5, w5 = 0.5,
[0277] w4+w5=1.
[0278] For insertion, deletion, inversion, and duplication variation signals, a comprehensive similarity score S is obtained; it is expressed as:
[0279] S=w1×S1+w2×S2+w3×S3
[0280] Among them, w1, w2, and w3 are the weight coefficients of position similarity score, variant size similarity score / chromosome similarity score, respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, w1 + w2 + w3 = 1;
[0281] The specific process is:
[0282] 1) Calculate the location similarity S1 score; the specific process is:
[0283] The calculation formula for the position similarity S1 score is as follows:
[0284] S1=|Start1-Start2|
[0285] Among them, Start1 and Start2 are the starting positions of the two variant signals respectively;
[0286] 2) Calculate the variation signal size similarity score S2; the specific process is:
[0287] The mutation lengths of the two mutation signals are SVlen1 and SVlen2 respectively. The similarity score S2 of the mutation length of the mutation signal is calculated by the following formula:
[0288]
[0289] 3) Calculate the chromosome similarity score S3; the specific process is:
[0290] Assume that the chromosome numbers of the two mutation signals are Chrom1 and Chrom2 respectively, and the chromosome similarity score S3 is calculated by the following discrete function:
[0291]
[0292] When the chromosome names of the two variant signals are the same, the chromosome similarity score is 1;
[0293] When the chromosome names of the two variant signals are different, the chromosome similarity score is 0;
[0294] 4) Perform weighted summation of the position similarity S1 score, the variation size similarity S2 score, and the chromosome similarity S3 score to obtain a comprehensive similarity score S; expressed as:
[0295] S=w1×S1+w2×S2+w3×S3
[0296] Among them, w1, w2, and w3 are the weight coefficients of the position similarity score and the variation size similarity score / chromosome similarity score, respectively, w1=0.2, w2=0.2, w3=0.6, w1+w2+w3=1.
[0297] The other steps and parameters are the same as those in the first to third embodiments.
[0298] Specific embodiment 5: This embodiment differs from specific embodiments 1 to 4 in that: for the translocation variation signal, a comprehensive similarity score S is obtained; it is expressed as:
[0299] S=w4×S4+w5×S5
[0300] in,
[0301] w4 is the weight coefficient of the initial chromosome similarity score S4, w4 = 0.5,
[0302] w5 is the weight coefficient of the termination chromosome similarity score S5, w5 = 0.5,
[0303] w4+w5=1;
[0304] The specific process is:
[0305] Calculate the starting chromosome similarity score:
[0306]
[0307] Among them, S4 represents the starting chromosome similarity score;
[0308] Indicates the name of the starting chromosome of the first signal;
[0309] Indicates the name of the starting chromosome of the second signal;
[0310] Calculate the terminating chromosome similarity score:
[0311]
[0312] Among them, S5 represents the termination chromosome similarity score;
[0313] The name of the chromosome that terminates the first signal;
[0314] The name of the chromosome that terminates the second signal;
[0315] The weighted sum of the starting chromosome similarity score S4 and the ending chromosome similarity score S5 is used to obtain a comprehensive similarity score S; it is expressed as:
[0316] S=w4×S4+w5×S5
[0317] in,
[0318] w4 is the weight coefficient of the initial chromosome similarity score S4, w4 = 0.5,
[0319] w5 is the weight coefficient of the termination chromosome similarity score S5, w5 = 0.5,
[0320] w4+w5=1.
[0321] The other steps and parameters are the same as those in the first to fourth embodiments.
[0322] Specific embodiment 6: This embodiment differs from any one of specific embodiments 1 to 5 in that: in S3, each output structural variation signal is haplotyped to obtain a hap1 file and a hap2 file after haplotyped;
[0323] The specific process is:
[0324] Step 31: Input the VCF file of each structural variation signal into the variation detection tool;
[0325] The variant detection tool outputs a typing VCF file, which contains M SNP sites;
[0326] The VCF file of the classification contains 8 columns of information, namely:
[0327] Chromosome name, SNP position, SNP ID, reference gene, alternative allele, quality value, filter flag, annotation information column;
[0328] Step 32: Use Bcftools software to filter the typing VCF file to obtain a filtered typing VCF file;
[0329] Step 33: Input the filtered typing VCF file obtained in step 22 into the WhatsHap typing tool. The WhatsHap typing tool outputs two typing files, namely, haplotype 1 file and haplotype 2 file.
[0330] In step 32, the typing VCF file is filtered using Bcftools software to obtain a filtered typing VCF file; the specific process is as follows:
[0331] Keep the rows in the VCF file with the "Filter Flag" value equal to PASS;
[0332] Delete the rows in the typing VCF file where the "Filter Flag" value is not equal to PASS;
[0333] Finally, the typing VCF file is generated;
[0334] Each SNP site in the typing VCF file is divided into haplotype 1 and haplotype 2.
[0335] In step 33, the filtered typing VCF file obtained in step 22 is input into the WhatsHap typing tool, and the WhatsHap typing tool outputs two typing files, namely, a haplotype 1 file and a haplotype 2 file.
[0336] The specific process is:
[0337] Step 331: Set Modki software parameters: pileup, traditional;
[0338] Use Modki software to convert the filtered typing VCF file obtained in step 22 into a tsv format file;
[0339] Step 332: Use the WhatsHap typing tool to generate a BED.gz format file from the tsv format file;
[0340] Step 333: Input the typing VCF file obtained in step 22 and the BED.gz format file obtained in step 332 into the WhatsHap typing tool, and the WhatsHap typing tool outputs a haplotype 1 file and a haplotype 2 file.
[0341] In step 4, the BAM file of the individual dairy cow obtained in step 1 is used as the input of the deep learning model, and the haplotype 1 file and haplotype 2 file after typing are used as the output of the deep learning model;
[0342] The BAM file of the individual cow obtained in step 1 is used as the input of the large language model, and the haplotype 1 file and haplotype 2 file after typing are used as the output of the large language model;
[0343] A comprehensive loss function is used to optimize the parameters of the deep learning model and the large language model, and the Adam optimizer is used to perform gradient updates until the comprehensive loss function converges to obtain the trained deep learning model and the large language model.
[0344] The specific process is:
[0345] Step 41: Build a deep learning model. The deep learning model includes:
[0346] The first 1×1 convolution layer, BN layer, the first 7×7 depth convolution layer, the first LN layer, the second 1×1 convolution layer, the first GELU, the first GRN, the third 1×1 convolution layer, the second 7×7 depth convolution layer, the second LN layer, the fourth 1×1 convolution layer, the second GELU, the second GRN, the fifth 1×1 convolution layer, the third 7×7 depth convolution layer, the third LN layer, the sixth 1×1 convolution layer, the third GELU, the third GRN, the seventh 1×1 convolution layer, the fully connected layer, and the softmax layer.
[0347] The working process of the deep learning model is as follows:
[0348] The BAM file of the individual cow obtained in step 1 is input into the first 1×1 convolutional layer and the BN layer in sequence. The BN layer outputs feature A.
[0349] The output feature A of the BN layer is sequentially input into the first 7×7 depth convolution layer, the first LN layer, the second 1×1 convolution layer, the first GELU, the first GRN, and the third 1×1 convolution layer. The third 1×1 convolution layer outputs feature A′;
[0350] The output feature A′ of the third 1×1 convolutional layer is added element-by-element to the output feature A of the BN layer to obtain feature A″;
[0351] Feature A″ is sequentially input into the second 7×7 depth convolution layer, the second LN layer, the fourth 1×1 convolution layer, the second GELU, the second GRN, and the fifth 1×1 convolution layer. The fifth 1×1 convolution layer outputs feature A″′;
[0352] The fifth 1×1 convolutional layer outputs feature A″′ and feature A″ and adds them element by element to obtain feature
[0353] feature Input the third 7×7 depth convolution layer, the third LN layer, the sixth 1×1 convolution layer, the third GELU, the third GRN, the seventh 1×1 convolution layer, and the seventh 1×1 convolution layer output features.
[0354] The seventh 1×1 convolutional layer outputs features and features Perform element-by-element summation to obtain feature B;
[0355] Feature B is input into the fully connected layer and the softmax layer in sequence, and the softmax layer outputs the classification result;
[0356] Step 42: Generate a haplotype 1 file and a haplotype 2 file corresponding to the BAM file of the individual dairy cow obtained in step 1 using the large language model;
[0357] Step 43: Use the comprehensive loss function to optimize the parameters of the deep learning model and the large language model, and use the Adam optimizer to perform gradient updates until the comprehensive loss function converges to obtain the trained deep learning model and the large language model.
[0358] The other steps and parameters are the same as those in the specific implementation modes 1 to 8-1.
[0359] Specific embodiment 10: This embodiment differs from any one of specific embodiments 1 to 9 in that the comprehensive loss function is:
[0360] Where N represents the total number of data in the BAM file of individual cows, i represents the i-th data, and k represents the k-th data.
[0361] F i 1The i-th data in the BAM file representing an individual dairy cow is input into the deep learning model, and the features output by the deep learning model;
[0362] F i 2 The i-th data in the BAM file representing the individual cow is input into the large language model, and the large language model outputs the semantic features of the text description information;
[0363] The kth data in the BAM file representing the individual cow is input into the large language model, and the large language model outputs the semantic features of the text description information;
[0364] s(F i 1 ,F i 2 ) represents the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the i-th data;
[0365] Represents the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the k-th data;
[0366] τ represents a temperature hyperparameter.
[0367] This implementation utilizes feature information acquired by a deep learning model and semantic information from a large language model to classify individual cow BAM files, significantly improving classification accuracy. Compared to traditional classification methods, the present invention offers significant advantages in classification accuracy. This implementation reduces data requirements, eliminates the need for pedigree data, and is suitable for large-scale studies, improving the method's applicability. The method avoids complex variant typing steps and directly analyzes the data, improving computational efficiency and reducing computing resource consumption.
[0368] The specific steps for knowledge graph completion are:
[0369] S1. Construction of benchmark DHI domain knowledge graph:
[0370] (1) Constructing the DHI domain ontology, which includes three types of entities and entity relationships: "performance indicators / symptoms", "influencing factors", and "solutions". The performance indicators / symptoms refer to the indicators used to present the health status of dairy cows or the symptoms manifested by dairy cows.
[0371] (2) The electronic texts obtained by digitizing professional books and documents related to DHI reports and application guidance are used as annotation objects. Based on the ontology, the electronic text data is semantically annotated to form annotated data;
[0372] (3) Using the data in the labeled data as training data, according to the ontology structure of the DHI domain knowledge graph, supervised, semi-supervised and unsupervised methods are used to extract entities and entity relationships from the text of the basic text library, obtain entity and entity relationship data, construct triples of any two types of entities and entity relationships, and the DHI domain knowledge graph, which are recorded as the baseline DHI domain knowledge graph; the knowledge graph uses the CompGCN network as the network framework. The basic text library includes professional books, documents and DHI reports on dairy cows and pastures, as well as texts on the Internet such as Baidu Encyclopedia. It can be constructed according to actual conditions. It should be noted that the more texts in this basic text library, the better the effect of constructing the baseline DHI domain knowledge graph will be in theory, but the workload will increase greatly. This application can appropriately construct the size of the basic text library, but it needs to be as comprehensive as possible to cover all known entities and entity relationships.
[0373] (4) For the triples containing two types of entities in the DHI knowledge graph, “influencing factors” and “performance indicators / symptoms”, the conditional probability between the two types of entities is calculated, denoted as P(fac|sym), and the weight of the edge between the entities is denoted as the basic edge weight ω m1 .
[0374] P(fac|sym) is obtained through crowdsourcing calculation, that is, all influencing factors that may lead to a certain performance indicator / symptom are provided to participants (ranch production personnel, ranch managers, field experts, etc.) through crowdsourcing software, and each participant ranks and scores the degree of influence of these factors; based on the data, the weight of each influencing factor and a certain performance indicator / symptom is calculated according to the principle of sorting first and then scoring. That is, according to the principle of minority obeys majority, the influencing factor with the most occurrences in each position is selected as the influencing factor of that position, and the average of the participant scores of the influencing factor in that position is taken as the weight between the influencing factor and the performance indicator / symptom. The same method is used to obtain the correlation strength between each performance indicator / symptom and each influencing factor, forming a correlation coefficient matrix, which represents the weight of the edge between the two types of entities: performance indicator / symptom and influencing factor.
[0375] S2. Completion of the benchmark DHI domain knowledge graph:
[0376] For professional books, literature, and DHI reports on existing dairy cows and pastures, the annotation of data is based on the entities and relationships between entities determined by the knowledge that has been recognized. For example, only when it is known which "influencing factors" will affect a certain "performance indicator / symptom" can the entities and entity relationships be labeled and extracted. The extraction of entities and entity relationships based on supervised, semi-supervised, or even unsupervised methods is basically based on this "known" information. However, there will be situations where "there are actually relationships between entities objectively, but we are not aware of it yet." This situation will lead to the problem that the constructed knowledge graph is actually incomplete, which will limit the interpretation effect of the knowledge graph to a certain extent. Therefore, this implementation method focuses on completing the benchmark DHI field knowledge graph:
[0377] S201. Take all entities in the benchmark DHI domain knowledge graph as reference entities, extract entities and entity relations (h, r, ?) in all texts in the incremental text library based on the BERT model, and predict the tail entity through the BERT model to obtain the predicted triple (h, r, t);
[0378] Based on the predicted triples, the benchmark DHI domain knowledge graph structure is completed to obtain a complete DHI domain knowledge graph structure, which is recorded as the full DHI domain knowledge graph. The difference between the full DHI domain knowledge graph and the benchmark DHI domain knowledge graph is recorded as the incremental DHI domain knowledge graph.
[0379] The neighbor nodes of a node in the baseline DHI domain knowledge graph are recorded as basic connection nodes, and the corresponding edges are basic edges. The neighbor nodes in the incremental DHI domain knowledge graph are incremental connection nodes (actually completed nodes), and the corresponding edges are incremental edges. For a node, calculate the proportion k of the number of corresponding incremental edges to the total number of edges. m , as the incremental adjustment coefficient; 1-k m As a basic adjustment coefficient. For example, a node has four neighboring nodes in the benchmark DHI domain knowledge graph. After completion, there are a total of six nodes, of which two are incremental connection nodes. The edges corresponding to the nodes and the incremental connection nodes are incremental edges. For this node, the incremental adjustment coefficient The basic adjustment coefficient is
[0380] S202, based on the two entities of "influencing factors" and "performance indicators / symptoms" in the DHI knowledge graph, count the non-repeating triples (h, r, t) corresponding to all entities in the incremental text library; among all the non-repeating triples (h, r, t), for a certain entity v m , will contain entity v m The triplet of statistics The number n m , in this process entity v m The number of triplets corresponding to the head entity h and the tail entity t must be counted;
[0381] For an entity v m , the statistical incremental text library contains Number of documents And calculate the entity v m The corresponding incremental edge weight Where N is the number of documents in the incremental text library.
[0382] In the process of completing the knowledge graph of the benchmark DHI domain, it is essentially necessary to supplement the entities that actually exist objectively but are not yet known. It is precisely because the relationship between entities is unknown to us that the weight of the edge cannot be obtained through step (4). In order to determine that the weight of the edge cannot be obtained through (4), the present invention turns to mining the objectively existing conditions. Since the existing texts in the incremental text library objectively record the relationship between entities, although some relationships are unknown, they objectively exist. Therefore, based on the BERT model, these unknown but actually existing relationships in the incremental text library can be extracted, and a certain entity v in the incremental text library m The number of corresponding texts can, to a certain extent, characterize the objective coexistence strength of entities and their relationships. Therefore, as long as the number of texts in the incremental text library is large enough, the relationship can be characterized. Therefore, the present invention calculates incremental edge weights by adopting objectively existing phenomena. In the process of calculating edge weights, considering that the higher the probability of an objective potential relationship being established, the higher the probability of its occurrence should be, or the more frequently it is potentially "represented", the present invention uses the frequency of occurrence of documents corresponding to a certain entity relationship (the proportion of entity documents) as its objective potential probability. At the same time, it is taken into account that for the potential relationships extracted by the large language model, this potential relationship is essentially an "estimated result", that is, a certain entity is estimated to have an entity relationship with other entities. The more relationships an entity has with other entities, the weaker its characterization ability is, and therefore its relationship confidence is relatively lower. For example, for an entity v A With ten entities v a1 -v a10 There is an entity relationship, which indicates that there may be ten factors that can cause a certain symptom to appear or a certain indicator to be abnormal; if the entity v A Only with 3 entities v a1 -v a3There is an entity relationship, which means that only three influencing factors will cause a certain symptom or a certain indicator to be abnormal; it can be seen that for the estimated potential relationship, compared with "one symptom / one indicator is related to ten influencing factors", "one symptom / one indicator is related to three influencing factors" indicates that the stronger the representation ability of the symptom / indicator is, the higher its confidence is. Therefore, the present invention adopts The number n m The reciprocal represents the confidence and is involved in the edge weight calculation.
[0383] S203, correct the basic edge weight in the process of building the benchmark DHI domain knowledge graph to (1-k m )ω m1 , and correct the incremental edge weight to k m ω m2 ;
[0384] The completion of the benchmark DHI domain knowledge graph is now complete. The DHI report is then interpreted using the completed DHI domain knowledge graph. The report interpretation process can be performed according to the DHI report interpretation method based on the knowledge graph in the prior art. In some embodiments, the interpretation process includes:
[0385] S301, DHI indicator data acquisition:
[0386] Obtain the DHI data of the ranch. The DHI indicator data includes this month's data and historical data. This month's data is directly obtained by using the DHI report file produced by the DHI Testing Center based on the China Dairy Cattle Production Performance Measurement and Analysis System (CNDHI) and automatically extracted by the software. This method is simple and efficient, and can provide basic measurement data and related statistical indicators, such as average calving interval, lactation days, milk fat rate, protein rate, fat-to-egg ratio, peak milk, peak day, sustainability, urea nitrogen, etc.; historical data is obtained by traversing the relevant indicator data in the historical DHI report file through the software, and obtaining the relevant indicator data according to the established fields and storing it in the database.
[0387] S302. Statistical analysis of indicator data:
[0388] Analyzing indicator data includes static analysis and dynamic analysis;
[0389] Static analysis is to find abnormal indicators based on the monthly data of various indicators and the normal range values of each indicator, and then form a corresponding fact description;
[0390] Dynamic analysis combines the current month's data and historical data of each indicator to analyze the recent changing patterns of each indicator and form corresponding factual descriptions.
[0391] S303, Problem Diagnosis:
[0392] Combine the DHI domain knowledge graph to diagnose problems based on the results of dynamic analysis. Problem diagnosis includes problem location and guidance measures.
[0393] Problem location is based on the DHI domain knowledge graph. The fact description of dynamic analysis is used as the "performance indicator / symptom" entity. The probability that the fact description is affected by a certain influencing factor is calculated, which is recorded as P(fac):
[0394] P(fac)=P(fac|sym)·P prior (sym)
[0395] Among them, P(fac|sym) is the conditional probability between the performance indicator / symptom and the influencing factor, that is, the weight of the edge between the entities (corrected basic edge weight and incremental edge weight); P prior (sym) is the prior probability of the performance indicator / symptom, P prior The initial value of (sym) is calculated based on the historical DHI reports and farm record data. That is, the prior probability of the performance indicator is the proportion of the number of reports with abnormal indicators in the historical DHI reports to the total number of reports; the prior probability of the symptom is the proportion of the number of cows with the symptom in the farm record data to the total number of cows. In addition, P prior (sym) can be updated monthly;
[0396] The guidance measures are recommended based on the influencing factors obtained by positioning. According to the relationship between the two entities of "influencing factors" and "solutions" in the DHI domain knowledge graph, the corresponding solutions to the influencing factors are found and fed back to the users.
[0397] It should be noted that the specific embodiments are merely explanations and illustrations of the technical solutions of the present invention and cannot be used to limit the scope of protection. Any minor changes made based on the claims and description of the present invention shall still fall within the scope of protection of the present invention.
Claims
1. A dairy cattle breeding information management system, characterized in that The system includes an information storage unit and an information retrieval unit; The information storage unit specifically performs the following steps: Step 11: receiving the cow breeding record text input by the user; Step 12: Perform keyword recognition on the dairy cow breeding record text to obtain the keywords and their corresponding values; Step 13: Determine the cow number based on the cow breeding record text; Step 14: According to the cow number, retrieve the cow breeding records stored in the database and obtain the corresponding keywords and values; Step 15: The keyword in step 14 is matched with the keyword in step 12. If the match is successful, the numerical values corresponding to the successfully matched keywords are averaged, the numerical value of the keyword is updated, and stored in the database. If the match is unsuccessful, the keyword in step 12 is used as a new keyword, and the keyword and its corresponding numerical value are stored in the database. The information retrieval unit specifically performs the following steps: Step 21: Receive the keyword queried by the user and the numerical range corresponding to the keyword; Step 22: Match the keyword input by the user in step 21 with the keyword corresponding to each cow number in the database, and sort the cow numbers from largest to smallest according to the number of successful keyword matches; Step 23: For the cow number ranked first, obtain the value of each keyword corresponding to the cow number, determine whether the value of each keyword is within the value range input by the user, and count the number of keywords corresponding to each cow number that are within the value range input by the user. Then, sort the cow numbers from largest to smallest based on the statistical results. Step 24: Based on the ranking in step 23, the number of the cow ranked first is used as the query result, that is, the cow that best meets the user's needs.
2. A dairy cattle breeding information management system according to claim 1, characterized in that The system further includes a cow recommendation module, which specifically performs the following steps: Step 31: Based on the cow number obtained by the information retrieval unit, obtain the historical cow breeding record text of the cow number, and use the Biaffine syntax analyzer to obtain a dependency syntax structure diagram of the historical cow breeding record text; Step 32: Calculate the relative dependency distances between different keywords using the dependency syntactic structure graph, and obtain the importance weights of different keywords in the historical dairy cow breeding record text relative to the dairy cow based on the relative dependency distances; Step 33: Assign a value to each keyword of the cow according to the importance weight of the cow; Step 34: Based on the assignment result of step 33, for each cow number in the database, obtain all the corresponding assigned keywords, add all the values, and then determine whether the addition result is greater than the threshold. If it is greater than the threshold, the corresponding cow number is recommended to the user, otherwise it is not recommended.
3. A dairy cattle breeding information management system according to claim 1, characterized in that The system further comprises a milk production prediction unit, which retrieves a keyword corresponding to the cow number according to the cow number obtained by the information retrieval unit, and predicts the milk production of the cow corresponding to the cow number according to the keyword.
4. A dairy cattle breeding information management system according to claim 3, characterized in that The milk production prediction unit specifically performs the following steps: Step 41: Retrieve the keyword and its corresponding value corresponding to the cow number, i.e., the cow's milk production characteristic data, according to the cow number obtained by the information retrieval unit; Step 42: After normalizing the cow milk production characteristic data, the data is input into the trained cow milk production prediction model to obtain the predicted cow milk production.
5. A dairy cattle breeding information management system according to claim 1, characterized in that The system further includes a structural variation detection module, which specifically performs the following steps: Step 1: Obtain the BAM file corresponding to the cow that best meets the user's needs. Each BAM file contains M SNP sites. Step 2: Obtain all clustering results of all variant type signals in the BAM file of individual dairy cows; The signals contained in each cluster are integrated into a structural variation signal for output; Structural variation signals include insertions, deletions, duplications, translocations, and inversions; Step 3: Perform haplotype typing on each output structural variation signal to obtain the typed haplotype 1 file and haplotype 2 file; Step 4: Use the BAM file of the individual dairy cow obtained in step 1 as the input of the deep learning model, and use the haplotype 1 file and haplotype 2 file after typing as the output of the deep learning model; The BAM file of the individual cow obtained in step 1 is used as the input of the large language model, and the haplotype 1 file and haplotype 2 file after typing are used as the output of the large language model; A comprehensive loss function is used to optimize the parameters of the deep learning model and the large language model, and the Adam optimizer is used to perform gradient updates until the comprehensive loss function converges to obtain the trained deep learning model and the large language model. Step 5: Input the BAM file of the individual dairy cow to be tested into the trained deep learning model, and the trained deep learning model outputs the haplotype 1 file and haplotype 2 file after typing the BAM file of the individual dairy cow to be tested; In step 2, all cluster results of all variant type signals in the BAM file of the individual dairy cow are obtained; The signals contained in each cluster are integrated into a structural variation signal for output; Structural variation signals include insertions, deletions, duplications, translocations, and inversions; The specific process is: Step 21: Get the clustering results of all signals in the mutation signal set of "insert" mutation type. The specific process is as follows: Step 211: Initialize an empty cluster and add the first signal in the mutation signal set of the "insert" mutation type as the starting signal; Step 212: Calculate the similarity between the current signal and the last signal in each existing cluster; Step 213: If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster; If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster; Step 214: Repeat steps 212 to 213 until all signals in the mutation signal set of the "insertion" mutation type are determined, and clustering results of all signals in the mutation signal set of the "insertion" mutation type are obtained; Step 22: Get the clustering results of all signals in the mutation signal set of the "deletion" mutation type. The specific process is as follows: Step 221: Initialize an empty cluster and add the first signal in the mutation signal set of the "deletion" mutation type as the starting signal; Step 222: Calculate the similarity between the current signal and the last signal in each existing cluster; Step 223: If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster; If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster; Step 224: Repeat steps 222 to 223 until all signals in the variation signal set of the "delete" variation type are determined; Step 23: Obtain the clustering results of all signals in the set of mutation signals of the "repeated" mutation type. The specific process is as follows: Step 231: Initialize an empty cluster and add the first signal in the set of mutation signals of the "repeat" mutation type as the starting signal; Step 232: Calculate the similarity between the current signal and the last signal in each existing cluster; Step 233: If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster; If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster; Step 234: Move to the next signal and repeat steps 232 to 233 until all signals in the set of variant signals of the "repeat" variant type are determined; Step 24: Obtain the clustering results of all signals in the mutation signal set of the "translocation" mutation type. The specific process is as follows: Step 241: Initialize an empty cluster and add the first signal in the set of mutation signals of the "translocation" mutation type as the starting signal; Step 242: Calculate the similarity between the current signal and the last signal in each existing cluster; Step 243: If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster; If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster; Step 244: Repeat steps 242 to 243 until all signals in the mutation signal set of the “translocation” mutation type are determined, and a clustering result of all signals in the mutation signal set of the “translocation” mutation type is obtained; Step 25: Obtain the clustering results of all signals in the mutation signal set of the "inversion" mutation type. The specific process is as follows: Step 251: Initialize an empty cluster and add the first signal in the set of mutation signals of the "inversion" mutation type as the starting signal; Step 252: Calculate the similarity between the current signal and the last signal in each existing cluster; Step 253: If the similarity is less than a predetermined threshold, a new cluster is constructed and the current signal is added to the new cluster; If the similarity is greater than or equal to a predetermined threshold, the current signal is added to the current cluster; Step 254: Repeat steps 252 to 253 until all signals in the mutation signal set of the "inversion" mutation type are determined; Get the clustering results of all signals in the mutation signal set of "inversion" mutation type; The specific process of calculating the similarity between the current signal and the last signal in each existing cluster is as follows: For insertion, deletion, inversion, and duplication variation signals, a comprehensive similarity score S is obtained, which is expressed as: S=w1×S1+w2×S2+w3×S3 Among them, w1, w2, and w3 are the weight coefficients of position similarity score, variant size similarity score / chromosome similarity score, respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, w1 + w2 + w3 = 1; For translocation mutation signals, a comprehensive similarity score S is obtained, which is expressed as: S=w4×S4+w5×S5 Wherein, w4 is the weight coefficient of the starting chromosome similarity score S4, w4=0.5, w5 is the weight coefficient of the ending chromosome similarity score S5, w5=0.5, w4+w5=1; For the insertion, deletion, inversion and duplication variation signals, a comprehensive similarity score S is obtained, which is expressed as: S=w1×S1+w2×S2+w3×S3 Among them, w1, w2, and w3 are the weight coefficients of position similarity score, variant size similarity score / chromosome similarity score, respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, w1 + w2 + w3 = 1; The specific process is: 1) Calculate the location similarity S1 score. The specific process is as follows: The calculation formula for the position similarity S1 score is as follows: S1=|Start1-Start2| Among them, Start1 and Start2 are the starting positions of the two variant signals respectively; 2) Calculate the variation signal size similarity score S2. The specific process is as follows: The mutation lengths of the two mutation signals are SVlen1 and SVlen2 respectively. The similarity score S2 of the mutation length of the mutation signal is calculated by the following formula: 3) Calculate the chromosome similarity score S3. The specific process is as follows: Assume that the chromosome numbers of the two mutation signals are Chrom1 and Chrom2 respectively, and the chromosome similarity score S3 is calculated by the following discrete function: When the chromosome names of the two variant signals are the same, the chromosome similarity score is 1; When the chromosome names of the two variant signals are different, the chromosome similarity score is 0; 4) Perform weighted summation of the position similarity S1 score, the variation size similarity S2 score, and the chromosome similarity S3 score to obtain a comprehensive similarity score S, which is expressed as: S=w1×S1+w2×S2+w3×S3 Among them, w1, w2, and w3 are the weight coefficients of position similarity score, variant size similarity score / chromosome similarity score, respectively, w1 = 0.2, w2 = 0.2, w3 = 0.6, w1 + w2 + w3 = 1; For the translocation variation signal, a comprehensive similarity score S is obtained, which is expressed as: S=w4×S4+w5×S5 Wherein, w4 is the weight coefficient of the starting chromosome similarity score S4, w4=0.5, w5 is the weight coefficient of the ending chromosome similarity score S5, w5=0.5, w4+w5=1; The specific process is: Calculate the starting chromosome similarity score: Among them, S4 represents the starting chromosome similarity score; Indicates the name of the starting chromosome of the first signal; Indicates the name of the starting chromosome of the second signal; Calculate the terminating chromosome similarity score: Among them, S5 represents the termination chromosome similarity score; The name of the chromosome that terminates the first signal; The name of the chromosome that terminates the second signal; The weighted sum of the starting chromosome similarity score S4 and the ending chromosome similarity score S5 is used to obtain a comprehensive similarity score S, which is expressed as: S=w4×S4+w5×S5 Wherein, w4 is the weight coefficient of the starting chromosome similarity score S4, w4=0.5, w5 is the weight coefficient of the ending chromosome similarity score S5, w5=0.5, w4+w5=1; In S3, each output structural variation signal is haplotyped to obtain the hap1 file and hap2 file after typing. The specific process is as follows: Step 31: Input the VCF file of each structural variation signal into the variation detection tool; The variant detection tool outputs a typing VCF file, which contains M SNP sites; The VCF file of the classification contains 8 columns of information, namely: Chromosome name, SNP position, SNP ID, reference gene, alternative allele, quality value, filter flag, annotation information column; Step 32: Use Bcftools software to filter the typing VCF file to obtain a filtered typing VCF file; Step 33: Input the filtered typing VCF file obtained in step 22 into the WhatsHap typing tool, which outputs two typing files, namely, haplotype 1 file and haplotype 2 file; In step 32, the typing VCF file is filtered using Bcftools software to obtain a filtered typing VCF file. The specific process is as follows: Keep the rows in the VCF file with the "Filter Flag" value equal to PASS; Delete the rows in the typing VCF file where the "Filter Flag" value is not equal to PASS; Finally, the typing VCF file is generated; Each SNP site in the typing VCF file is divided into haplotype 1 and haplotype 2; In step 33, the filtered typing VCF file obtained in step 22 is input into the WhatsHap typing tool, and the WhatsHap typing tool outputs two typing files, namely, haplotype 1 file and haplotype 2 file. The specific process is as follows: Step 331: Set Modki software parameters: pileup, traditional; Use Modki software to convert the filtered typing VCF file obtained in step 22 into a tsv format file; Step 332: Use the WhatsHap typing tool to generate a BED.gz format file from the tsv format file; Step 333: Input the typing VCF file obtained in step 22 and the BED.gz format file obtained in step 332 into the WhatsHap typing tool, and the WhatsHap typing tool outputs a haplotype 1 file and a haplotype 2 file; In step 4, the BAM file of the individual dairy cow obtained in step 1 is used as the input of the deep learning model, and the haplotype 1 file and haplotype 2 file after typing are used as the output of the deep learning model; The BAM file of the individual cow obtained in step 1 is used as the input of the large language model, and the haplotype 1 file and haplotype 2 file after typing are used as the output of the large language model; A comprehensive loss function is used to optimize the parameters of the deep learning model and the large language model, and the Adam optimizer is used to perform gradient updates until the comprehensive loss function converges to obtain the trained deep learning model and the large language model. The specific process is: Step 41: Build a deep learning model. The deep learning model includes: The first 1×1 convolution layer, BN layer, the first 7×7 depth convolution layer, the first LN layer, the second 1×1 convolution layer, the first GELU, the first GRN, the third 1×1 convolution layer, the second 7×7 depth convolution layer, the second LN layer, the fourth 1×1 convolution layer, the second GELU, the second GRN, the fifth 1×1 convolution layer, the third 7×7 depth convolution layer, the third LN layer, the sixth 1×1 convolution layer, the third GELU, the third GRN, the seventh 1×1 convolution layer, the fully connected layer, and the softmax layer. The working process of the deep learning model is as follows: The BAM file of the individual cow obtained in step 1 is input into the first 1×1 convolutional layer and the BN layer in sequence. The BN layer outputs feature A. The output feature A of the BN layer is sequentially input into the first 7×7 depth convolution layer, the first LN layer, the second 1×1 convolution layer, the first GELU, the first GRN, and the third 1×1 convolution layer. The third 1×1 convolution layer outputs feature A′; The output feature A′ of the third 1×1 convolutional layer is added element-by-element to the output feature A of the BN layer to obtain feature A″; Feature A″ is sequentially input into the second 7×7 depth convolution layer, the second LN layer, the fourth 1×1 convolution layer, the second GELU, the second GRN, and the fifth 1×1 convolution layer. The fifth 1×1 convolution layer outputs feature A″′; The fifth 1×1 convolutional layer outputs feature A″′ and feature A″ and adds them element by element to obtain feature feature Input the third 7×7 depth convolution layer, the third LN layer, the sixth 1×1 convolution layer, the third GELU, the third GRN, the seventh 1×1 convolution layer, and the seventh 1×1 convolution layer output features. The seventh 1×1 convolutional layer outputs features and features Perform element-by-element summation to obtain feature B; Feature B is input into the fully connected layer and the softmax layer in sequence, and the softmax layer outputs the classification result; Step 42: Generate a haplotype 1 file and a haplotype 2 file corresponding to the BAM file of the individual dairy cow obtained in step 1 using the large language model; Step 43: Optimize the parameters of the deep learning model and the large language model using a comprehensive loss function, and perform gradient updates in combination with the Adam optimizer until the comprehensive loss function converges, thereby obtaining the trained deep learning model and the large language model. The comprehensive loss function is Where N represents the total number of data in the BAM file of individual cows, i represents the i-th data, and k represents the k-th data; F i 1 The i-th data in the BAM file representing an individual dairy cow is input into the deep learning model, and the features output by the deep learning model; F i 2 The i-th data in the BAM file representing the individual cow is input into the large language model, and the large language model outputs the semantic features of the text description information; The kth data in the BAM file representing the individual cow is input into the large language model, and the large language model outputs the semantic features of the text description information; s(F i 1 ,F i 2 ) represents the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the i-th data; Represents the similarity between the feature corresponding to the i-th data and the semantic feature corresponding to the k-th data; τ represents a temperature hyperparameter.
6. A dairy cattle breeding information management system according to claim 1, characterized in that The system also includes a DHI report generation module, which is used to receive user instructions and, after receiving the user instructions, obtain the DHI data of the ranch and then interpret the DHI report using the completed knowledge graph. The completed knowledge graph is obtained by the following steps: First, we obtain the benchmark DHI domain knowledge graph and record the edge weights in the benchmark DHI domain knowledge graph as the basic edge weight ω m1 , and then complete it based on the benchmark DHI domain knowledge graph, including the following steps: S201. Take all entities in the benchmark DHI domain knowledge graph as reference entities, extract entities and entity relationships in all texts in the incremental text library based on the language big model, and predict the tail entity through the language big model to obtain the predicted triple (h, r, t), where h and t are used to represent the head entity and tail entity in the triple, and r is used to represent the head entity relationship in the triple. Based on the predicted triples, the benchmark DHI domain knowledge graph structure is completed to obtain a complete DHI domain knowledge graph structure, which is recorded as the full DHI domain knowledge graph. The difference between the full DHI domain knowledge graph and the benchmark DHI domain knowledge graph is recorded as the incremental DHI domain knowledge graph; The neighbor nodes of a node in the benchmark DHI domain knowledge graph are recorded as basic connection nodes, and the corresponding edges are basic edges. The neighbor nodes in the incremental DHI domain knowledge graph are incremental connection nodes, and the corresponding edges are incremental edges. For a node, calculate the proportion k of the number of corresponding incremental edges to the total number of edges. m , as the incremental adjustment coefficient, 1-k m As the basic adjustment coefficient; S202, based on the two entities of "influencing factors" and "performance indicators / symptoms" in the DHI knowledge graph, count the non-repeating triples (h, r, t) corresponding to all entities in the incremental text library, and among all the non-repeating triples (h, r, t), for a certain entity v m , will contain entity v m The triplet of statistics The number n m , and then based on n m Get entity v m The corresponding incremental edge weight ω m2 ; S203, correct the basic edge weight in the process of building the benchmark DHI domain knowledge graph to (1-k m )ω m1 , and correct the incremental edge weight to k m ω m2 , and then complete the knowledge graph of the benchmark DHI domain; The construction process of the benchmark DHI domain knowledge graph includes: (1) Constructing the DHI domain ontology, which includes three types of entities and entity relationships: "performance indicators / symptoms", "influencing factors", and "solutions". The performance indicators / symptoms refer to the indicators used to present the health status of dairy cows or the symptoms manifested by dairy cows; (2) The electronic text obtained after the DHI measurement and application guidance is digitized is used as the annotation object, and the ontology is used as the annotation basis to perform semantic annotation on the electronic text data to form annotation data; (3) Using the data in the labeled data as training data, according to the ontology structure of the DHI domain knowledge graph, extract entities and entity relationships from the text of the basic text library to obtain entity and entity relationship data, and construct triples of any two types of entities and entity relationships, as well as the DHI domain knowledge graph, which is recorded as the baseline DHI domain knowledge graph; (4) For the triples containing two types of entities, “influencing factors” and “performance indicators / symptoms”, in the DHI knowledge graph, the conditional probability between the two types of entities is calculated, denoted as P(fac|sym), and the weight of the edge between the entities is denoted as the basic edge weight ω m1 ; The P(fac|sym) is obtained by crowdsourcing calculation; Based on n m Get entity v m The corresponding incremental edge weight ω m2 The process includes: For an entity v m , the statistical incremental text library contains Number of documents And calculate the entity v m The corresponding incremental edge weight Where N is the number of documents in the incremental text library; The knowledge graph adopts the CompGCN network as the network framework; The large language model selects the BERT model.
7. A dairy cattle breeding information management system according to claim 1, characterized in that The keywords include body height, chest width, body depth, udder depth, central suspensory ligament length, front teat position, front teat length, udder attachment height, udder position, lactation days, parity and udder attachment width.
8. A dairy cattle breeding information management system according to claim 1, characterized in that The keyword recognition is performed by the Chinese keyword extractor Jieba.
9. A dairy cattle breeding information management system according to claim 8, characterized in that The dependency syntax structure graph is represented in the form of an adjacency matrix D, where each element in D is represented as: Among them, i represents the row index of the matrix, j represents the column index of the matrix, and w i and w j Represents any two keywords in the dairy cow breeding record text.
10. A dairy cattle breeding information management system according to claim 9, characterized in that The specific steps of calculating the relative dependency distance using the dependency syntactic structure graph are as follows: Based on the adjacency matrix D, the Dijkstra algorithm is used to obtain the relative dependency distance between different keywords through the shortest distance of different keywords on the adjacency matrix.
Citation Information
Patent Citations
DHI report interpretation method and system based on knowledge graph and storage medium
CN113656600A
Cited By
A multi-dimensional dairy cow breeding database construction method and system
CN122412385A