A method and system for constructing a veterinary knowledge graph
By analyzing the scope and consistency of key paragraphs in the veterinary knowledge graph and identifying synonyms, the problem of indistinguishability of entity words is solved, and the accuracy and quality of graph completion are improved.
Patent Information
- Application Number
- CN202510615224.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-14
AI Technical Summary
In the veterinary knowledge graph, many entity words have similar surface forms or semantic meanings and cannot be effectively distinguished, resulting in low accuracy of map completion, which in turn affects the quality of the map.
By analyzing the scope of key paragraphs in the veterinary knowledge literature, a comparison literature group was constructed, the consistency of the literature description content was compared, synonyms were identified and the map was updated.
Improve the accuracy of map completion and improve the quality of veterinary knowledge graphs.
Smart Images

Figure CN120123436B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for constructing a veterinary knowledge graph. Background Art
[0002] As an innovative way of knowledge representation and application in the field of veterinary medicine, the veterinary knowledge graph is gradually becoming an effective means to achieve animal health management and disease prevention and control. It organizes various knowledge elements in the field of veterinary medicine in a structured form, forming a huge and orderly knowledge network. The veterinary knowledge graph not only helps veterinarians quickly and accurately obtain the required information, but also provides strong data support for animal health research. Therefore, constructing a high-quality and complete veterinary knowledge graph is of crucial significance for promoting the development of the field of veterinary medicine.
[0003] Currently, in the process of constructing a veterinary knowledge graph, the graph completion technology is an important means to improve the quality of the graph. It makes the graph more complete and accurate by mining and supplementing the missing or imperfect information in the graph.
[0004] However, in the veterinary knowledge graph, many entity words have similar surface forms or semantic meanings, and it is impossible to effectively distinguish similar entity words. As a result, the accuracy of graph completion is relatively low, and the quality of the veterinary knowledge graph is poor. Summary of the Invention
[0005] The embodiments of the present invention provide a method and system for constructing a veterinary knowledge graph, which can improve the accuracy of graph completion and thus improve the quality of the veterinary knowledge graph.
[0006] In the first aspect of the embodiments of the present invention, a method for constructing a veterinary knowledge graph is provided, including:
[0007] Based on a first database, construct an initial veterinary knowledge graph, where the first database includes a plurality of veterinary knowledge documents;
[0008] For a first entity word and a second entity word in the initial veterinary knowledge graph, determine the key paragraph ranges of the first knowledge information and the second knowledge information in each veterinary knowledge document according to the distribution positions of the first entity word and the second entity word in each veterinary knowledge document, where the first entity word and the second entity word are two similar entity words;
[0009] Construct a comparison document group according to the detailed description information of the first knowledge information and the second knowledge information within the corresponding key paragraph ranges, where the comparison document group includes a first veterinary knowledge document and a second veterinary knowledge document with similar detailed description degrees;
[0010] Determine the content consistency between the first veterinary knowledge document and the second veterinary knowledge document in the comparative document group according to the document description content of the first veterinary knowledge document and the second veterinary knowledge document in the comparative document group;
[0011] Based on the content consistency, determine the relationship recognition result between the first entity word and the second entity word, and the relationship recognition result is used to represent whether the first entity word and the second entity word are synonyms;
[0012] Update the initial veterinary knowledge graph according to the relationship recognition result to obtain the target veterinary knowledge graph.
[0013] In the second aspect of the embodiments of the present invention, there is provided a construction system for a veterinary knowledge graph, including:
[0014] A graph construction module for constructing an initial veterinary knowledge graph based on a first database, where the first database includes multiple veterinary knowledge documents;
[0015] A paragraph determination module for determining the key paragraph ranges of the first knowledge information and the second knowledge information in each veterinary knowledge document according to the distribution positions of the first entity word and the second entity word in each veterinary knowledge document for the first entity word and the second entity word of the initial veterinary knowledge graph, and the first entity word and the second entity word are two similar entity words;
[0016] A document group construction module for constructing a comparative document group according to the detailed description information of the first knowledge information and the second knowledge information within the corresponding key paragraph ranges, where the comparative document group includes a first veterinary knowledge document and a second veterinary knowledge document with similar degrees of detailed description;
[0017] A consistency evaluation module for determining the content consistency between the first veterinary knowledge document and the second veterinary knowledge document in the comparative document group according to the document description content of the first veterinary knowledge document and the second veterinary knowledge document in the comparative document group;
[0018] A relationship recognition module for determining the relationship recognition result between the first entity word and the second entity word based on the content consistency, and the relationship recognition result is used to represent whether the first entity word and the second entity word are synonyms;
[0019] A graph update module for updating the initial veterinary knowledge graph according to the relationship recognition result to obtain the target veterinary knowledge graph.
[0020] In the method for constructing a veterinary knowledge graph provided by an embodiment of the present invention, for the first entity word and the second entity word that are similar in the initial veterinary knowledge graph, first, according to their distribution positions in veterinary knowledge documents, the key paragraph ranges of each veterinary knowledge document are found. Then, according to the key paragraph ranges of each veterinary knowledge document, the first veterinary knowledge document corresponding to the first entity word and the second veterinary knowledge document corresponding to the second entity word with similar degrees of detailed description are found to form a comparison document group. Thus, according to the document description contents of the first veterinary knowledge document and the second veterinary knowledge document in the comparison document group, their content consistency is determined. Finally, according to the content consistency, it can be accurately identified whether the first entity word and the second entity word belong to synonyms. In this way, the present invention screens the key paragraph ranges of each veterinary knowledge document, and then compares the document description contents of the first veterinary knowledge document and the second veterinary knowledge document, so as to accurately judge whether the first entity word and the second entity word belong to synonyms. Furthermore, the accuracy of graph completion can be improved, and the quality of the veterinary knowledge graph can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 It is a schematic flowchart of the first method for constructing a veterinary knowledge graph provided by an embodiment of the present invention;
[0023] Figure 2 It is a schematic flowchart of the second method for constructing a veterinary knowledge graph provided by an embodiment of the present invention;
[0024] Figure 3 It is a schematic flowchart of the third method for constructing a veterinary knowledge graph provided by an embodiment of the present invention;
[0025] Figure 4 It is an update schematic diagram of an initial veterinary knowledge graph provided by an embodiment of the present invention;
[0026] Figure 5 It is a schematic structural diagram of a system for constructing a veterinary knowledge graph provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, detail the specific implementation manners, structures, features and effects of a method and system for constructing a veterinary knowledge graph proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0029] It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solution of the present invention all comply with the relevant provisions of laws and regulations.
[0030] It should be noted that in the embodiments of the present invention, certain existing industry solutions such as software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present invention, but it does not mean that the applicant has already or necessarily used this solution.
[0031] As an innovative way of knowledge representation and application in the field of veterinary medicine, the veterinary knowledge graph is gradually becoming an effective means to achieve animal health management and disease prevention and control. It organizes various knowledge elements in the field of veterinary medicine in a structured form, forming a huge and orderly knowledge network. The veterinary knowledge graph not only helps veterinarians quickly and accurately obtain the required information and improve the diagnosis efficiency, but also provides strong data support for animal health research. Therefore, constructing a high-quality and complete veterinary knowledge graph is of crucial significance for promoting the development of the field of veterinary medicine.
[0032] Currently, in the process of constructing a veterinary knowledge graph, the graph completion technology is an important means to improve the quality of the graph. It makes the graph more complete and accurate by mining and supplementing the missing or imperfect information in the graph. However, in the veterinary knowledge graph, many entity words have similar surface forms or semantic meanings, and it is impossible to effectively distinguish similar entity words. As a result, the accuracy of graph completion is relatively low, and further the quality of the veterinary knowledge graph is poor.
[0033] The object of the present invention is to provide a method and system for constructing a veterinary knowledge graph. In the method for constructing a veterinary knowledge graph provided by the embodiments of the present invention, for the first entity word and the second entity word that are similar in the initial veterinary knowledge graph, first, according to their distribution positions in the veterinary knowledge literature, the key paragraph ranges of each veterinary knowledge literature are found. Then, according to the key paragraph ranges of each veterinary knowledge literature, the first veterinary knowledge literature corresponding to the first entity word and the second veterinary knowledge literature corresponding to the second entity word with similar degrees of detail description are found to form a comparison literature group. Thus, according to the literature description contents of the first veterinary knowledge literature and the second veterinary knowledge literature in the comparison literature group, their content consistency is determined. Finally, according to the content consistency, it can be accurately identified whether the first entity word and the second entity word belong to synonyms. In this way, the present invention can accurately determine whether the first entity word and the second entity word belong to synonyms by screening the key paragraph ranges of each veterinary knowledge literature and then comparing the literature description contents of the first veterinary knowledge literature and the second veterinary knowledge literature. Furthermore, the accuracy of graph completion can be improved, and the quality of the veterinary knowledge graph can be improved.
[0034] The following introduces specific embodiments of a method and system for constructing a veterinary knowledge graph provided by the embodiments of the present invention.
[0035] Figure 1 A flowchart of a method for constructing a veterinary knowledge graph is provided. The method for constructing the veterinary knowledge graph can be applied to a server, and the method for constructing the veterinary knowledge graph may include the following S101 to S106.
[0036] S101, based on a first database, construct an initial veterinary knowledge graph, where the first database includes multiple veterinary knowledge literatures.
[0037] In this embodiment, the first database includes multiple veterinary knowledge literatures. By way of example, the veterinary knowledge literatures may include academic papers in the veterinary field, clinical cases in the veterinary field, and professional books in the veterinary field, etc.
[0038] As an example, the server uses natural language processing technology to extract entity words in the veterinary field from the veterinary knowledge literatures, such as disease names, drug names, symptom descriptions, etc. Then, the contents of the veterinary knowledge literatures are further analyzed to identify the entity relationships between the entity words, such as "disease - symptom", "drug - mechanism of action", etc.
[0039] Finally, the identified entity words and the relationships between them are organized in the form of a graph to form an initial veterinary knowledge graph.
[0040] S102. For the first entity term and the second entity term of the initial veterinary knowledge graph, determine the key paragraph ranges of the first knowledge information and the second knowledge information in each veterinary knowledge document according to the distribution positions of the first entity term and the second entity term in each veterinary knowledge document. The first entity term and the second entity term are two similar entity terms.
[0041] In this embodiment, the first entity term and the second entity term are two similar entity terms in the initial veterinary knowledge graph. For example, "dog" and "canine" can be the first entity term and the second entity term respectively.
[0042] The first knowledge information is a triple formed by the first entity term and the corresponding entity relationship, and the second knowledge information is a triple formed by the second entity term and the corresponding entity relationship. The first knowledge information and the second knowledge information are two similar knowledge information. Specifically, that is, the entity terms in the first knowledge information and the second knowledge information are similar, and the entity relationships are the same.
[0043] The key paragraph range is used to represent the key occurrence paragraphs of the first knowledge information or the second knowledge information in the veterinary knowledge document. By way of example, the first knowledge information appears 10 times in veterinary knowledge document A, and 8 of them are in the second paragraph. Then the second paragraph is the key paragraph range in veterinary knowledge document A.
[0044] As an example, the server first analyzes the initial veterinary knowledge graph to obtain the similar first entity term and second entity term in the initial veterinary knowledge graph.
[0045] Then, analyze the distribution positions of the first entity term in each veterinary knowledge document, and determine the key paragraph ranges of the first knowledge information in each veterinary knowledge document according to the concentrated distribution of the first entity term in each veterinary knowledge document; and analyze the distribution positions of the second entity term in each veterinary knowledge document, and determine the key paragraph ranges of the second knowledge information in each veterinary knowledge document according to the concentrated distribution of the second entity term in each veterinary knowledge document.
[0046] S103. Construct a group of comparative documents according to the detailed description information of the first knowledge information and the second knowledge information within the corresponding key paragraph ranges. The group of comparative documents includes the first veterinary knowledge document and the second veterinary knowledge document with similar degrees of detailed description.
[0047] In this embodiment, the first veterinary knowledge document is the veterinary knowledge document corresponding to the first knowledge information, the second veterinary knowledge document is the veterinary knowledge document corresponding to the second knowledge information, and the first veterinary knowledge document and the second veterinary knowledge document with similar degrees of detailed description form a group of comparative documents.
[0048] The detailed description information is used to characterize the degree of detailed description of the first knowledge information or the second knowledge information within the corresponding key paragraph range.
[0049] As an example, the server extracts the detailed description text of the first knowledge information from within the key paragraph range of the first veterinary knowledge document, and extracts the detailed description text of the second knowledge information from within the key paragraph range of the second veterinary knowledge document.
[0050] Then, the server uses a text similarity algorithm (such as cosine similarity, Jaccard similarity, etc.) to calculate the similarity between the detailed descriptions in the first veterinary knowledge document and the second veterinary knowledge document.
[0051] Finally, based on the similarity calculation result, the server pairs the first veterinary knowledge document and the second veterinary knowledge document with a similarity greater than the preset similarity threshold to form a group of comparative documents.
[0052] S104. According to the document description content of the first veterinary knowledge document and the second veterinary knowledge document in the group of comparative documents, determine the content consistency of the first veterinary knowledge document and the second veterinary knowledge document in the group of comparative documents.
[0053] In this embodiment, the content consistency is used to characterize the degree of consistency of the description content between the first veterinary knowledge document and the second veterinary knowledge document.
[0054] As an example, the server deeply analyzes the document description content of the first veterinary knowledge document and the second veterinary knowledge document in the group of comparative documents, including the theme, viewpoints, data, etc.
[0055] Then, a quantitative or qualitative method is used to evaluate the content consistency between the first veterinary knowledge document and the second veterinary knowledge document, such as calculating the content overlap degree, viewpoint consistency, etc.
[0056] S105. Based on the content consistency, determine the relationship recognition result of the first entity word and the second entity word. The relationship recognition result is used to characterize whether the first entity word and the second entity word are synonyms.
[0057] In this embodiment, the relationship recognition result is used to characterize the relationship between the first entity word and the second entity word. Exemplarily, the relationship recognition result may include that the first entity word and the second entity word are synonyms, and that the first entity word and the second entity word are not synonyms.
[0058] As an example, when the content consistency is greater than the preset consistency threshold, the recognition result of the relationship between the first entity word and the second entity word is determined as that the first entity word and the second entity word are synonyms; when the content consistency is not greater than the preset consistency threshold, the recognition result of the relationship between the first entity word and the second entity word is determined as that the first entity word and the second entity word are not synonyms.
[0059] S106. Update the initial veterinary knowledge graph according to the relationship recognition result to obtain a target veterinary knowledge graph.
[0060] In this embodiment, the target veterinary knowledge graph is the final veterinary knowledge graph obtained after updating the initial veterinary knowledge graph.
[0061] As an example, the server first determines a graph update strategy according to the relationship recognition result. For example, if synonyms are recognized, they can be merged into one node or a synonym link can be added.
[0062] Then, modify the initial veterinary knowledge graph according to the graph update strategy, such as merging nodes, modifying node relationships, etc., so as to obtain a target veterinary knowledge graph.
[0063] In the method for constructing a veterinary knowledge graph provided in this embodiment, for similar first entity word and second entity word in the initial veterinary knowledge graph, first find the key paragraph range of each veterinary knowledge document according to their distribution positions in the veterinary knowledge documents. Then, according to the key paragraph ranges of each veterinary knowledge document, find the first veterinary knowledge document corresponding to the first entity word and the second veterinary knowledge document corresponding to the second entity word with similar degrees of detail description to form a comparison document group. Thus, determine their content consistency according to the document description contents of the first veterinary knowledge document and the second veterinary knowledge document in the comparison document group. Finally, according to the content consistency, it can be accurately recognized whether the first entity word and the second entity word are synonyms. In this way, the present invention can accurately judge whether the first entity word and the second entity word are synonyms by screening the key paragraph ranges of each veterinary knowledge document and then comparing the document description contents of the first veterinary knowledge document and the second veterinary knowledge document. Furthermore, the accuracy of graph completion can be improved, and the quality of the veterinary knowledge graph can be improved.
[0064] As an optional embodiment, S101 may specifically include:
[0065] Extract each target entity word from the first database based on the veterinary term library, and the veterinary term library includes each entity word in the veterinary field;
[0066] Extract each target entity relationship from the first database based on the veterinary relationship library, and the veterinary relationship library includes each entity relationship in the veterinary field;
[0067] Construct each target entity relationship and the corresponding target entity word into triples to obtain multiple target knowledge information;
[0068] Utilize each target knowledge information to construct an initial veterinary knowledge graph.
[0069] In this embodiment, the veterinary terminology library contains all entity words in the veterinary field, such as disease names, drug names, surgical names, animal species, etc.
[0070] The veterinary relationship library contains all entity relationships between entity words in the veterinary field, such as the relationship between diseases and symptoms, the relationship between drugs and treatments, etc.
[0071] The triple is in the form of "entity - relationship - entity". For example, "cat - suffers from - feline distemper" is a triple, where "cat" and "feline distemper" are entity words, and "suffers from" is an entity relationship.
[0072] As an example, the server uses natural language processing techniques, such as named entity recognition technology, to extract target entity words that match the veterinary terminology library from the first database. Specifically, it can be achieved by training a specific named entity recognition model, which can recognize and extract entity words related to the veterinary field.
[0073] Then, use the relationship extraction method in named entity recognition technology to extract target entity relationships that match the veterinary relationship library from the first database. Specifically, it can be achieved by training a specific relationship extraction model, which can recognize and extract entity relationships related to the veterinary field.
[0074] Then, construct the extracted target entity words and target entity relationships into triples in the form of "entity - relationship - entity" to obtain multiple target knowledge information.
[0075] Finally, utilize each target knowledge information to construct an initial veterinary knowledge graph through a knowledge representation model (such as TranE).
[0076] Among them, the core idea of the TransE model is to regard knowledge information as a kind of "translation" operation. By adjusting the vector representations of entity words and entity relationships, the head entity word vector plus the entity relationship vector is equal to the tail entity word vector, so as to represent the relationship of triples in the vector space. Specifically, in the initial veterinary knowledge graph, the target entity words are used as the nodes of the graph, and the target entity relationships are used as the edges of the graph.
[0077] Through this embodiment, based on the veterinary terminology database and the veterinary relationship database, target entity words and target entity relationships are extracted from the first database, thereby constructing an initial veterinary knowledge graph. In this way, the initial veterinary knowledge graph can be quickly generated according to the first database, improving the generation efficiency of the veterinary knowledge graph.
[0078] As an alternative embodiment, as Figure 2 shown, S102 may specifically include the following S201 to S204:
[0079] S201, obtain the occurrence frequency of the target knowledge information in each document paragraph of the target veterinary knowledge document, where the target veterinary knowledge document is any veterinary knowledge document, and the target knowledge information is any one of the first knowledge information and the second knowledge information;
[0080] S202, determine the document paragraph with the highest occurrence frequency as the starting point, and construct multiple attention windows according to the preset traversal rule;
[0081] S203, for each attention window, execute respectively: determine the knowledge concentration factor of the attention window according to the occurrence frequency of the target knowledge information in the attention window;
[0082] S204, determine the document paragraph range corresponding to the attention window with the largest knowledge concentration factor as the key paragraph range of the target veterinary knowledge document.
[0083] In this embodiment, the knowledge concentration factor is used to characterize the concentration degree of the target knowledge information within the attention window.
[0084] As an example, the server traverses each document paragraph of the target veterinary knowledge document, counts the number of times the target knowledge information appears in each document paragraph, and calculates its occurrence frequency (that is, the number of occurrences in the document paragraph divided by the total number of occurrences in the target veterinary knowledge document).
[0085] Then, according to the occurrence frequencies of the document paragraphs, find the document paragraph with the highest occurrence frequency and use it as the starting point of the traversal. Then, add a document paragraph in turn in the order from top to bottom to form an attention window, thereby obtaining multiple attention windows. Specifically, assume that the starting point of the traversal is the fifth paragraph of the target veterinary knowledge document, then the first attention window is the fifth paragraph of the target veterinary knowledge document, the second attention window is the fifth and fourth paragraphs of the target veterinary knowledge document, the third attention window is the fifth, fourth, and sixth paragraphs of the target veterinary knowledge document, and so on.
[0086] Then, for each attention window, calculate the total occurrence frequency of the target knowledge information therein. Specifically, it can be obtained by adding up the occurrence frequencies of the target knowledge information in all paragraphs within the attention window. According to the total occurrence frequency of the target knowledge information within the window, determine the knowledge concentration factor of this window. Specifically, the average occurrence frequency can be obtained by dividing the total occurrence frequency by the number of paragraphs within the window, and used as the knowledge concentration factor.
[0087] Finally, compare the knowledge concentration factors of all attention windows and find the largest one among them. Determine all paragraphs within this attention window as the key paragraph range of the target veterinary knowledge literature.
[0088] Through this embodiment, according to the preset traversal rule, multiple attention windows of the target veterinary knowledge literature are constructed. Then calculate the knowledge concentration factors of each attention window, so as to determine the key paragraph range of the target veterinary knowledge literature. In this way, this embodiment can accurately determine the key paragraph range of each veterinary knowledge literature, which helps to construct a group of comparative literatures based on the key paragraph ranges of each veterinary knowledge literature subsequently, and can improve the accuracy of map completion.
[0089] As an alternative embodiment, S203 may specifically include:
[0090] Compare the first occurrence frequency of the target knowledge information in the k-th attention window with the second occurrence frequency of the target knowledge information in the reference attention window of the k-th attention window, to obtain the frequency difference degree between the k-th attention window and the reference attention window. The reference attention window is N attention windows after the k-th attention window, and k and N are positive integers;
[0091] Based on the first frequency difference between the first paragraph and the second frequency difference between the second paragraph in the k-th attention window, determine the frequency concentration degree of the k-th attention window. The first paragraph is the first paragraph in the k-th attention window, the first frequency difference is the difference in occurrence frequency between the first paragraph and the adjacent paragraph outside the k-th attention window, the second paragraph is the last paragraph in the k-th attention window, and the second frequency difference is the difference in occurrence frequency between the second paragraph and the adjacent paragraph outside the k-th attention window;
[0092] Use the frequency difference degree between the k-th attention window and the reference attention window, and the frequency concentration degree of the k-th attention window to determine the knowledge concentration factor of the k-th attention window.
[0093] In this embodiment, the knowledge concentration factor can be specifically determined by the following formula 1:
[0094] Formula 1
[0095] In formula 1, Used to characterize the knowledge concentration factor of knowledge information i in the k-th attention window of the p-th veterinary knowledge document. Used to characterize the difference in the occurrence frequency of knowledge information i between the first paragraph in the k-th attention window and its adjacent paragraph outside the k-th attention window. Used to characterize the difference in the occurrence frequency of knowledge information i between the last paragraph in the k-th attention window and its adjacent paragraph outside the k-th attention window. By way of example, assuming that the k-th attention window is from the 5th paragraph to the 8th paragraph, then is the difference in occurrence frequency between the 5th paragraph and the 4th paragraph. is the difference in occurrence frequency between the 8th paragraph and the 9th paragraph.
[0096] Wherein, Used to characterize the occurrence frequency of knowledge information i in the k-th attention window. Used to characterize the occurrence frequency of knowledge information i in the (k + b)-th attention window. N is a preset positive integer. By way of example, N can be 5.
[0097] Wherein, Used to characterize the frequency concentration degree of the k-th attention window. And The larger it is, the greater the frequency concentration degree of the k-th attention window, and the greater the knowledge concentration factor of the k-th attention window. Used to characterize the frequency difference degree between the k-th attention window and the reference attention window. The larger it is, the greater the frequency difference between the k-th attention window and the reference attention window, and the greater the knowledge concentration factor of the k-th attention window.
[0098] Through this embodiment, according to the occurrence frequency of the target knowledge information in the attention window, the knowledge concentration factor of the attention window is accurately calculated. In this way, it helps to accurately determine the range of key paragraphs of the veterinary knowledge document according to the knowledge concentration factor of the attention window, and improve the accuracy of determining the range of key paragraphs.
[0099] As an alternative embodiment, as Figure 3 shown, S103 may specifically include:
[0100] S301, determining the first semantic discrimination ratio of the first knowledge information according to the occurrence frequency of the first knowledge information within the corresponding range of key paragraphs, and determining the second semantic discrimination ratio of the second knowledge information according to the occurrence frequency of the second knowledge information within the corresponding range of key paragraphs.
[0101] S302. When the difference between the first semantic discrimination ratio and the second semantic discrimination ratio is less than a preset difference, the first veterinary knowledge document to which the key paragraph range of the first knowledge information belongs and the second veterinary knowledge document to which the key paragraph range of the second knowledge information belongs are constructed into a comparison document group.
[0102] In this embodiment, the first semantic discrimination ratio is used to measure the importance and significance of the first knowledge information within the key paragraph range of the first veterinary knowledge document, and the second semantic discrimination ratio is used to measure the importance and significance of the second knowledge information within the key paragraph range of the second veterinary knowledge document.
[0103] The preset difference is a threshold used to determine whether the importance of the first knowledge information and the second knowledge information within the key paragraph range of the document is close. Among them, the preset difference can be set according to specific application scenarios and requirements.
[0104] As an example, the server respectively counts the occurrence frequencies of the first knowledge information and the second knowledge information within the corresponding key paragraph ranges. Then, based on the occurrence frequencies, the first semantic discrimination ratio of the first knowledge information and the second semantic discrimination ratio of the second knowledge information are analyzed.
[0105] Then, the difference between the first semantic discrimination ratio and the second semantic discrimination ratio is calculated. The calculated difference is compared with the preset difference. If the difference is less than the preset difference, it is considered that the importance of the first knowledge information and the second knowledge information within the key paragraph ranges of their respective documents is similar. Thus, the first veterinary knowledge document to which the key paragraph range of the first knowledge information belongs and the second veterinary knowledge document to which the key paragraph range of the second knowledge information belongs are constructed into a comparison document group.
[0106] Through this embodiment, the occurrence frequency of knowledge information within the key paragraph range of the document is calculated to determine its semantic discrimination ratio, and the comparison document group is constructed based on the semantic discrimination ratio. In this way, the comparison document group can be accurately constructed, which helps to subsequently determine whether the first entity word and the second entity word are synonyms according to the comparison document group, improving the quality of the veterinary knowledge graph.
[0107] As an alternative embodiment, S301 may specifically include:
[0108] Obtain the detailed description degree of the first knowledge information according to the number of interval sentences between adjacent first knowledge information within the key paragraph range;
[0109] Utilize the detailed description degree and the occurrence frequency of the first knowledge information within the corresponding key paragraph range to determine the first semantic discrimination ratio of the first knowledge information.
[0110] In this embodiment, the semantic discrimination ratio can be specifically determined by the following formula 2:
[0111] Formula 2
[0112] In Formula 2, represents the semantic discrimination ratio of knowledge information i, represents the occurrence frequency of knowledge information i within the scope of the key paragraph, represents the number of intervening sentences between the s-th sentence where knowledge information i is located and the (s + 1)-th sentence where knowledge information i is located, represents the number of sentences corresponding to knowledge information i. It should be noted that to ensure the significance of the calculation results, in the case of encountering a denominator of 0 during the fractional operation in the embodiments of the present invention, a tuning factor greater than 0 needs to be added to the denominator to prevent the denominator from being 0. The value of the tuning factor is set by the implementer according to the actual situation, and this application does not make special restrictions.
[0113] Among them, the larger it is, the more detailed the description of knowledge information i within the scope of the key paragraph, the greater the importance of the first knowledge information within the scope of the key paragraph in the first veterinary knowledge document, and the corresponding semantic discrimination ratio is larger; represents the degree of detailed description of knowledge information i. The larger this value is, the greater the degree of detailed description of knowledge information i, and the corresponding semantic discrimination ratio is larger.
[0114] Through this embodiment, according to the number of intervening sentences between adjacent first knowledge information within the scope of the key paragraph and the occurrence frequency of the first knowledge information within the corresponding scope of the key paragraph, the first semantic discrimination ratio of the first knowledge information can be accurately calculated. In this way, it helps to accurately construct a group of comparative documents according to the first semantic discrimination ratio later, and improve the accuracy of the group of comparative documents.
[0115] As an alternative embodiment, S104 may specifically include:
[0116] Determine the first information quantity of each group of comparative documents according to the document description content of the first veterinary knowledge document and the second veterinary knowledge document in each group of comparative documents. The first information quantity is the number of matching knowledge information of the first veterinary knowledge document and the second veterinary knowledge document within the scope of the key paragraph;
[0117] Use the first information quantity of each group of comparative documents and the position distribution of the matching knowledge information in each group of comparative documents to determine the content consistency of the first veterinary knowledge document and the second veterinary knowledge document in each group of comparative documents.
[0118] In this embodiment, the first information quantity is used to characterize the number of matching knowledge information of the first veterinary knowledge document and the second veterinary knowledge document within the key paragraph range in the group of comparative documents. Specifically, the matching knowledge information is used to characterize knowledge information with similar or identical entity words and the same entity relationship. By way of example, two entity words can be identified as similar based on semantic similarity.
[0119] As an example, the content consistency can be specifically determined by the following formula 3:
[0120] Formula 3
[0121] In formula 3, is used to characterize the content consistency between the key paragraph range c of the first veterinary knowledge document and the key paragraph range of the second veterinary knowledge document. is used to characterize the first information quantity of the group of comparative documents composed of the key paragraph range c and the key paragraph range and is used to characterize the maximum value among the first information quantities of each group of comparative documents.
[0122] Among them, is used to characterize the difference between the shortest distance between other matching knowledge information and the j-th matching knowledge information in the key paragraph range c and the shortest distance between other matching knowledge information and the j-th matching knowledge information in the key paragraph range .
[0123] Among them, When the value of is larger, it indicates that the number of matching knowledge information of the first veterinary knowledge document and the second veterinary knowledge document in the key paragraph range in the group of comparative documents is larger, and the content consistency is larger; When the value of is smaller, it indicates that the distribution positions of the matching knowledge information of the first veterinary knowledge document and the second veterinary knowledge document in the key paragraph range in the group of comparative documents are more similar, and the content consistency is larger.
[0124] Through this embodiment, according to the document description content of the first veterinary knowledge document and the second veterinary knowledge document in the group of comparative documents, the content consistency of the first veterinary knowledge document and the second veterinary knowledge document in the group of comparative documents can be accurately determined. Thus, it helps to determine whether the first entity word and the second entity word are synonyms according to the content consistency of the first veterinary knowledge document and the second veterinary knowledge document, and improves the recognition accuracy of the first entity word and the second entity word.
[0125] As an alternative embodiment, after S104, the method for constructing the veterinary knowledge graph may further include:
[0126] Obtain the second information quantity between the first extended knowledge information formed by the first entity words and the second extended knowledge information formed by the second entity words in each second database, where the second database is a knowledge literature database for each field except the veterinary field, and the second information quantity is the number of matching knowledge information between the first extended knowledge information and the second extended knowledge information in the second database;
[0127] Use the maximum value among the content consistencies and each second information quantity to determine the thing representation consistency between the first entity word and the second entity word;
[0128] S105 may specifically include:
[0129] Based on the thing representation consistency, determine the relationship recognition result between the first entity word and the second entity word.
[0130] In this embodiment, the first extended knowledge information is used to represent the knowledge information formed by the first entity word in other fields except the veterinary field, and the second extended knowledge information is used to represent the knowledge information formed by the second entity word in other fields except the veterinary field.
[0131] As an example, the thing representation consistency can be specifically determined by the following formula 4:
[0132] Formula 4
[0133] In formula 4, is used to represent the thing representation consistency between the j-th entity word and the -th entity word, is used to represent the maximum value of the content consistencies of each group of comparison documents, is used to represent the second information quantity between the extended knowledge information formed by the j-th entity word and the -th entity word, is used to represent the maximum value of the second information quantity in each second database.
[0134] Among them, The larger it is, the more similar the roles of the j-th entity word and the -th entity word in other fields are, that is, the more likely the j-th entity word and the -th entity word are the same thing, and the greater the thing representation consistency between the j-th entity word and the -th entity word.
[0135] Then, when the consistency degree of the thing representatives is greater than the preset consistency degree threshold, determine the relationship recognition result between the first entity word and the second entity word as that the first entity word and the second entity word are synonyms; when the consistency degree of the thing representatives is not greater than the preset consistency degree threshold, determine the relationship recognition result between the first entity word and the second entity word as that the first entity word and the second entity word are not synonyms.
[0136] Through this embodiment, further consider the roles of the first entity word and the second entity word in other fields except the veterinary field, so as to obtain the consistency degree of the thing representatives of the first entity word and the second entity word. In this way, according to the consistency degree of the thing representatives of the first entity word and the second entity word, it is possible to more accurately judge whether the first entity word and the second entity word are synonyms, and improve the recognition accuracy of the first entity word and the second entity word.
[0137] As an optional embodiment, S106 may specifically include:
[0138] When the relationship recognition result indicates that the first entity word and the second entity word are synonyms, perform entity alignment on the first entity word and the second entity word in the initial veterinary knowledge graph to obtain an entity alignment result;
[0139] Based on the entity alignment result, update the initial veterinary knowledge graph to obtain a target veterinary knowledge graph.
[0140] In this embodiment, when the server determines that the first entity word and the second entity word are synonyms, it performs entity alignment on the first entity word and the second entity word in the initial veterinary knowledge graph. Specifically, various strategies can be adopted for entity alignment, such as rule-based alignment, similarity-based alignment, or machine learning-based alignment.
[0141] Then, update the initial veterinary knowledge graph according to the entity alignment result. Specifically, that is, perform a merging process on the aligned first entity word and the second entity word, so as to obtain the updated target veterinary knowledge graph.
[0142] Exemplarily, as Figure 4 shown, a schematic diagram for updating the initial veterinary knowledge graph is provided. Among them, the meaning represented by the first entity word 430 in the initial veterinary knowledge graph 410 is "dog", and the meaning represented by the second entity word 440 is "canine". According to the analysis, the first entity word 430 and the second entity word 440 are synonyms. At this time, perform entity alignment on the first entity word 430 and the second entity word 440, and then perform a merging process on the aligned first entity word 430 and the second entity word 440, so as to obtain the updated target veterinary knowledge graph 420.
[0143] In this embodiment, when it is determined that the first entity word and the second entity word are synonyms, the first entity word and the second entity are entity-aligned, and after alignment, the first entity word and the second entity word are merged. In this way, the initial veterinary knowledge graph is updated to obtain the target veterinary knowledge graph, which can improve the accuracy of the veterinary knowledge graph.
[0144] A construction method of a veterinary knowledge graph. Correspondingly, the present invention also provides a specific embodiment of a veterinary knowledge graph construction system.
[0145] Figure 5 A structural schematic diagram of a veterinary knowledge graph construction system is provided. The veterinary knowledge graph construction system 500 includes a graph construction module 510, a paragraph determination module 520, a literature group construction module 530, a consistency evaluation module 540, a relationship recognition module 550, and a graph update module 560.
[0146] The graph construction module 510 is used to construct an initial veterinary knowledge graph based on the first database, and the first database includes a plurality of veterinary knowledge documents;
[0147] The paragraph determination module 520 is used to determine the key paragraph ranges of the first knowledge information and the second knowledge information in each veterinary knowledge document according to the distribution positions of the first entity word and the second entity word of the initial veterinary knowledge graph in each veterinary knowledge document. The first entity word and the second entity word are two similar entity words;
[0148] The literature group construction module 530 is used to construct a comparison literature group according to the detailed description information of the first knowledge information and the second knowledge information within the corresponding key paragraph ranges. The comparison literature group includes a first veterinary knowledge document and a second veterinary knowledge document with similar degrees of detailed description;
[0149] The consistency evaluation module 540 is used to determine the content consistency of the first veterinary knowledge document and the second veterinary knowledge document in the comparison literature group according to the literature description content of the first veterinary knowledge document and the second veterinary knowledge document in the comparison literature group;
[0150] The relationship recognition module 550 is used to determine the relationship recognition result of the first entity word and the second entity word based on the content consistency. The relationship recognition result is used to represent whether the first entity word and the second entity word belong to synonyms;
[0151] The graph update module 560 is used to update the initial veterinary knowledge graph according to the relationship recognition result to obtain the target veterinary knowledge graph.
[0152] In the construction system of the veterinary knowledge graph provided in this embodiment, for the similar first entity word and second entity word in the initial veterinary knowledge graph, first, according to their distribution positions in veterinary knowledge documents, the key paragraph ranges of each veterinary knowledge document are found. Then, according to the key paragraph ranges of each veterinary knowledge document, the first veterinary knowledge document corresponding to the first entity word and the second veterinary knowledge document corresponding to the second entity word with similar degrees of detail description are found to form a comparison document group. Thus, according to the document description contents of the first veterinary knowledge document and the second veterinary knowledge document in the comparison document group, their content consistency is determined. Finally, according to the content consistency, it can be accurately identified whether the first entity word and the second entity word belong to synonyms. In this way, the present invention screens the key paragraph ranges of each veterinary knowledge document, and then compares the document description contents of the first veterinary knowledge document and the second veterinary knowledge document, so as to accurately judge whether the first entity word and the second entity word belong to synonyms. Furthermore, the accuracy of graph completion can be improved, and the quality of the veterinary knowledge graph can be improved.
[0153] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed descriptions of known methods are omitted here. In the above embodiment, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0154] It should also be noted that the exemplary embodiments mentioned in the present invention describe some methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.
[0155] As described above, only the specific embodiments of the present invention are provided. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A method for constructing a veterinary knowledge graph, characterized in that, The method includes: Based on a first database, constructing an initial veterinary knowledge graph, where the first database includes multiple veterinary knowledge documents; For a first entity word and a second entity word of the initial veterinary knowledge graph, according to the distribution positions of the first entity word and the second entity word in each of the veterinary knowledge documents, determining the key paragraph ranges of the first knowledge information and the second knowledge information in each of the veterinary knowledge documents, where the first entity word and the second entity word are two similar entity words; According to the detailed description information of the first knowledge information and the second knowledge information within the corresponding key paragraph ranges, constructing a comparison document group, where the comparison document group includes a first veterinary knowledge document and a second veterinary knowledge document with similar degrees of detailed description; According to the document description contents of the first veterinary knowledge document and the second veterinary knowledge document in the comparison document group, determining the content consistency degree of the first veterinary knowledge document and the second veterinary knowledge document in the comparison document group; Based on the content consistency degree, determining the relationship recognition result of the first entity word and the second entity word, where the relationship recognition result is used to represent whether the first entity word and the second entity word belong to synonyms; According to the relationship recognition result, updating the initial veterinary knowledge graph to obtain a target veterinary knowledge graph.
2. The construction method of the veterinary knowledge graph according to claim 1, characterized in that, The constructing an initial veterinary knowledge graph based on a first database includes: Based on a veterinary term library, extracting each target entity word from the first database, where the veterinary term library includes each entity word in the veterinary field; Based on a veterinary relationship library, extracting each target entity relationship from the first database, where the veterinary relationship library includes each entity relationship in the veterinary field; Constructing each target entity relationship and the corresponding target entity word into a triple to obtain multiple target knowledge information; Using each target knowledge information to construct the initial veterinary knowledge graph.
3. The construction method of the veterinary knowledge graph according to claim 1, wherein The determining the key paragraph ranges of the first knowledge information and the second knowledge information in each of the veterinary knowledge documents according to the distribution positions of the first entity word and the second entity word of the initial veterinary knowledge graph in each of the veterinary knowledge documents includes: Obtaining the occurrence frequency of the target knowledge information in each document paragraph of a target veterinary knowledge document, where the target veterinary knowledge document is any one of the veterinary knowledge documents, and the target knowledge information is any one of the first knowledge information and the second knowledge information; Determining the document paragraph with the maximum occurrence frequency as the starting point, and constructing multiple attention windows according to a preset traversal rule; For each of the attention windows, respectively performing: determining the knowledge concentration factor of the attention window according to the occurrence frequency of the target knowledge information in the attention window; Determining the document paragraph corresponding to the attention window with the maximum knowledge concentration factor as the key paragraph range of the target veterinary knowledge document.
4. The construction method of the veterinary knowledge graph according to claim 3, characterized in that, The determining the knowledge concentration factor of the attention window according to the occurrence frequency of the target knowledge information in the attention window includes: Compare the first occurrence frequency of the target knowledge information in the k-th attention window with the second occurrence frequency of the target knowledge information in the reference attention window of the k-th attention window, to obtain the frequency difference degree between the k-th attention window and the reference attention window, where the reference attention window is N attention windows after the k-th attention window, and k and N are positive integers; Based on the first frequency difference between the first paragraph and the second paragraph in the k-th attention window, determine the frequency concentration degree of the k-th attention window. The first paragraph is the first paragraph in the k-th attention window, the first frequency difference is the difference in occurrence frequency between the first paragraph and the adjacent paragraph outside the k-th attention window, the second paragraph is the last paragraph in the k-th attention window, and the second frequency difference is the difference in occurrence frequency between the second paragraph and the adjacent paragraph outside the k-th attention window; Use the frequency difference degree between the k-th attention window and the reference attention window, and the frequency concentration degree of the k-th attention window, to determine the knowledge concentration factor of the k-th attention window.
5. The construction method of the veterinary knowledge graph according to claim 1, characterized in that, Construct a comparison literature group according to the detailed description information of the first knowledge information and the second knowledge information within the corresponding key paragraph range, including: Determine the first semantic discrimination ratio of the first knowledge information according to the occurrence frequency of the first knowledge information within the corresponding key paragraph range, and determine the second semantic discrimination ratio of the second knowledge information according to the occurrence frequency of the second knowledge information within the corresponding key paragraph range; In the case where the difference between the first semantic discrimination ratio and the second semantic discrimination ratio is less than a preset difference, construct the first veterinary knowledge literature to which the key paragraph range of the first knowledge information belongs and the second veterinary knowledge literature to which the key paragraph range of the second knowledge information belongs into the comparison literature group.
6. The construction method of the veterinary knowledge graph according to claim 5, characterized in that The determination of the first semantic discrimination ratio of the first knowledge information according to the occurrence frequency of the first knowledge information within the corresponding key paragraph range includes: Obtain the detailed description degree of the first knowledge information according to the number of interval sentences between adjacent first knowledge information within the key paragraph range; Use the detailed description degree and the occurrence frequency of the first knowledge information within the corresponding key paragraph range to determine the first semantic discrimination ratio of the first knowledge information.
7. The construction method of the veterinary knowledge graph according to claim 1, characterized in that The determination of the content consistency degree between the first veterinary knowledge literature and the second veterinary knowledge literature in the comparison literature group according to the literature description content of the first veterinary knowledge literature and the second veterinary knowledge literature in the comparison literature group includes: According to the literature description content of the first veterinary knowledge literature and the second veterinary knowledge literature in each of the comparison literature groups, determine the first information quantity of each of the comparison literature groups, where the first information quantity is the quantity of matching knowledge information of the first veterinary knowledge literature and the second veterinary knowledge literature in the key paragraph range in the comparison literature group; Using the first information quantity of each of the comparison literature groups and the position distribution of the matching knowledge information in each of the comparison literature groups, determine the content consistency of the first veterinary knowledge literature and the second veterinary knowledge literature in each of the comparison literature groups.
8. The construction method of the veterinary knowledge graph according to claim 1, characterized in that After determining the content consistency of the first veterinary knowledge literature and the second veterinary knowledge literature in the comparison literature group according to the literature description content of the first veterinary knowledge literature and the second veterinary knowledge literature in the comparison literature group, the method further includes: Obtain the second information quantity between the first extended knowledge information formed by the first entity word and the second extended knowledge information formed by the second entity word in each second database, where the second database is a knowledge literature database for each field except the veterinary field, and the second information quantity is the quantity of matching knowledge information of the first extended knowledge information and the second extended knowledge information in the second database; Using the maximum value of each of the content consistencies and each of the second information quantities, determine the thing representation consistency of the first entity word and the second entity word. Based on the content consistency, determining the relationship recognition result of the first entity word and the second entity word includes: Based on the thing representation consistency, determine the relationship recognition result of the first entity word and the second entity word.
9. The construction method of the veterinary knowledge graph according to claim 1, wherein, Updating the initial veterinary knowledge graph according to the relationship recognition result to obtain a target veterinary knowledge graph includes: In the case where the relationship recognition result indicates that the first entity word and the second entity word are synonyms, perform entity alignment of the first entity word and the second entity word in the initial veterinary knowledge graph to obtain an entity alignment result; Based on the entity alignment result, update the initial veterinary knowledge graph to obtain the target veterinary knowledge graph.
10. A construction system for a veterinary knowledge graph, characterized in that, The system includes: A graph construction module for constructing an initial veterinary knowledge graph based on a first database, where the first database includes a plurality of veterinary knowledge literatures; A paragraph determination module for, for the first entity word and the second entity word of the initial veterinary knowledge graph, determine the key paragraph range of the first knowledge information and the second knowledge information in each of the veterinary knowledge literatures according to the distribution positions of the first entity word and the second entity word in each of the veterinary knowledge literatures, where the first entity word and the second entity word are two similar entity words; A literature group construction module for constructing a comparison literature group according to the detailed description information of the first knowledge information and the second knowledge information in the corresponding key paragraph range, where the comparison literature group includes a first veterinary knowledge literature and a second veterinary knowledge literature with similar detailed description degrees; A consistency evaluation module, configured to determine the content consistency between the first veterinary knowledge document and the second veterinary knowledge document in the comparison document group according to the document description content of the first veterinary knowledge document and the second veterinary knowledge document in the comparison document group; A relationship recognition module, configured to determine a relationship recognition result between the first entity word and the second entity word based on the content consistency, where the relationship recognition result is used to characterize whether the first entity word and the second entity word are synonyms; A graph update module, configured to update the initial veterinary knowledge graph according to the relationship recognition result to obtain a target veterinary knowledge graph.
Citation Information
Patent Citations
Knowledge graph updating method, system and device for game question-answering system
CN110532399A
Storage method and system of common sense knowledge graph
CN116910276A