Knowledge base establishment method and system of service platform

By building a dynamic knowledge graph through heterogeneous graph attention network and deep learning technology, the problems of low accuracy and static knowledge extraction in existing technologies are solved, and efficient and intelligent knowledge base construction is achieved, meeting the needs of high-precision information query and knowledge reasoning.

CN120671793APending Publication Date: 2025-09-19CHINA TELECOM DIGITAL INTELLIGENCE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510857635.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

When constructing a knowledge base, existing technologies have problems such as low accuracy of knowledge extraction, static nature, and inability to fully utilize deep learning technology for deep semantic understanding. This results in a low level of intelligence in the knowledge base, making it difficult to meet the needs of high-precision information query and knowledge reasoning.

Method used

By adopting heterogeneous graph attention network and deep learning technology, we can achieve high-precision knowledge extraction and relational reasoning by constructing a corpus, identifying entity-relationship pairs, forming a knowledge graph, and dynamically updating the graph database.

Benefits of technology

It improves the accuracy and completeness of knowledge extraction, enhances the efficiency and intelligence of knowledge base construction, solves the shortcomings of traditional methods in complex text processing, and ensures the timeliness and scalability of knowledge base content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671793A_ABST
    Figure CN120671793A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge base establishment method and system of a service platform, and belongs to the technical field of natural language processing. Sentences in the corpus are input into a constructed network model, entity relation pairs are recognized, and the network model comprises an input layer, a network feature extraction layer and a relation extraction layer; the input layer encodes words and relations in sentences at the same time to serve as input vectors, and the feature extraction layer fuses the input vectors in the heterogeneous graph attention network; the relation extraction layer identifies entity relation pairs by using word nodes and relation nodes output by the feature extraction layer; and forming a knowledge graph based on the identified entity relationship pair, storing the entity relationship pair as a triple, and updating the knowledge graph when the text data in the corpus is updated. According to the method, deep analysis is carried out on complex texts by utilizing natural language processing and graph neural network technologies, so that high-precision knowledge extraction and relation reasoning are realized, and the construction efficiency and the intelligent degree of a knowledge base are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a method and system for establishing a knowledge base of a service platform. Background Art

[0002] With the rapid development of big data, artificial intelligence, and natural language processing technologies, service platforms are increasingly demanding knowledge management and application. These platforms aim to establish comprehensive, dynamic knowledge base systems to achieve efficient information acquisition, intelligent analysis, and automated reasoning, thereby improving user experience and service quality. However, due to the large volume and complex structure of text data, building a knowledge base that meets the specific needs of these platforms remains challenging.

[0003] Existing technical solutions primarily rely on simple keyword matching and rule extraction to build knowledge bases. While these methods can meet basic knowledge management needs, they struggle to process large amounts of unstructured text data. Furthermore, when extracting entities and relationships, existing solutions often overlook the complex semantic context, resulting in inaccurate and incomplete relationship extraction, which in turn impacts the effectiveness of knowledge base construction.

[0004] The above-mentioned existing technical solutions have the following problems: First, the accuracy of knowledge extraction is low, especially when dealing with multi-level and complex semantic structures, the existing technology has difficulty in accurately identifying entities and their relationships; second, the constructed knowledge base is often static and difficult to cope with the continuous updating of dynamic data; third, it is impossible to fully utilize the latest deep learning technology to perform deep semantic understanding of complex texts, resulting in a low level of intelligence in the knowledge base, which is difficult to meet the needs of high-precision information query and knowledge reasoning. Summary of the Invention

[0005] In response to the deficiencies in the existing technology, the present invention provides a knowledge base establishment method and system for a service platform, which can fully utilize natural language processing and graph neural network technology to perform in-depth analysis of complex texts, achieve high-precision knowledge extraction and relational reasoning, and significantly improve the construction efficiency and intelligence level of the knowledge base.

[0006] The present invention provides the following technical solutions:

[0007] In a first aspect, a method for establishing a knowledge base of a service platform is provided, comprising:

[0008] Build a corpus;

[0009] Input the sentences in the corpus into the constructed network model to identify entity-relationship pairs. The network model includes an input layer, a network feature extraction layer, and a relationship extraction layer. The input layer encodes the words and relationships in the sentence as input vectors. The feature extraction layer fuses the input vectors in the heterogeneous graph attention network and introduces a gate mechanism and residual connection to update the word nodes and relationship nodes. The relationship extraction layer uses the word node m output by the feature extraction layer to extract the entity-relationship pairs. l and relationship node r l , identify entity-relationship pairs;

[0010] Represent the identified entity-relationship pairs as nodes and edges, and merge the duplicate entities in all entity-relationship pairs to form a knowledge graph;

[0011] Using graph database tools, the entity-relationship pairs in the knowledge graph are stored as triples, and the knowledge graph is updated when the text data in the corpus is updated.

[0012] Optionally, the constructing of the corpus is specifically as follows:

[0013] Using public databases and text crawler technology to collect text data consistent with the current knowledge base construction goals;

[0014] Perform data cleaning on the collected text data;

[0015] With the help of automated text annotation tools, the cleaned data is formatted and sequence labeled to build a corpus.

[0016] Optionally, the input layer uses pre-trained BERT to simultaneously encode words and relations in the sentence to obtain word embedding vectors and relation embedding vectors;

[0017] [m1,m2,...,m n ]=BERT([x1,x2,...,x n ])

[0018] [r1,r2,...,r m ]=σ([a1,a2,...,a m ])

[0019] Among them, m i represents the encoded word embedding vector (i=1,2,…,n), r j represents the encoded relation embedding vector (i=1,2,…,m), σ represents the linear mapping, T(x1,x2,…,x n ) is the input text set, n represents the length of the entire text, x i Represents the i-th word, a1, a2, ..., a mis a predefined relationship label.

[0020] Optionally, the feature extraction layer fuses the input vectors in the heterogeneous graph attention network, specifically:

[0021] Embed the word output of the input layer into the vector m i and the relation embedding vector r j Use them as word nodes and relationship nodes respectively to construct a heterogeneous graph;

[0022] According to the similarity scores between the word node and its adjacent relationship nodes, weights are assigned to all adjacent relationship nodes of the word node to fuse all relationship nodes and update the word node;

[0023] Update the relationship node based on all word nodes adjacent to the relationship node;

[0024] Based on updated word nodes and relationship nodes The gate mechanism and residual connection are introduced to output the final word nodes and relation nodes.

[0025] Optionally, the method assigns weights to all adjacent relationship nodes of a word node according to the similarity scores between the word node and its adjacent relationship nodes, so as to merge all relationship nodes and update the word node, specifically:

[0026] w ij =w1[m i ,r j ]+b1

[0027]

[0028] Among them, w ij is the similarity score between the i-th word node and the j-th relationship node, The updated word node after fusion of all relationship nodes; [m i ,r j ] is the word embedding vector and relation embedding vector after feature concatenation; w1 is [m i ,r j ] weight coefficient; b1 is the first bias coefficient, exp(w ij ) represents w ij The exponential function, α ij is the probability of the attention weight of the i-th word node to the j-th relationship node, w2 is r j The weight coefficient is , N is the length of the current sentence text, and h is the number of relationship nodes actually associated with the current word node in the text;

[0029] The updating of the relationship node based on all word nodes adjacent to the relationship node is specifically as follows:

[0030]

[0031] Among them, q is the number of word nodes actually associated with the current relationship node in the text, w2′ is m i The weight coefficient of .

[0032] Optionally, the updated word node and relationship nodes The gate mechanism and residual connection are introduced to output the final word nodes and relationship nodes, specifically:

[0033] Based on the input word embedding vector m i and the updated word node Calculate the gate value e i ;

[0034]

[0035] Among them, w3 is The weight coefficient, sigmoid is the function, and b2 is the second bias term;

[0036] Using the gate value e i , update the word node and relationship node again, and get the updated word node and relationship nodes

[0037]

[0038] Use residual connections to output the final word nodes and relationship nodes;

[0039]

[0040] Among them, m l and r l They are the word nodes and relationship nodes output by the feature extraction layer respectively.

[0041] Optionally, the relationship extraction layer uses the word node m output by the feature extraction layer l and relationship node r l , identify entity-relationship pairs, specifically:

[0042] The subject tagger uses binary classification to calculate the probability of a word node being identified as the start point and end point of the main entity, and continuously optimizes the likelihood function to identify the subject;

[0043]

[0044] in, is the probability that the i-th word node represents the starting point, represents the probability that the i-th word node represents the termination point, w s and w e are two optimizable weight coefficients for subject identification, b s and b e are two optimizable bias coefficients for subject identification;

[0045] After the word nodes, relationship nodes and candidate main entities are fused, they are input into the object tagger. The object tagger uses binary classification to calculate the probability of the fused object entity being identified as the starting point and the ending point of the object entity, and continuously optimizes the likelihood function to identify the object;

[0046] m o =w4[m l ,r l ,δ k ]+b4

[0047]

[0048] Among them, m o is the fused object entity, δ k is the kth candidate main entity, w4 represents [m l ,r l ,δ k ] can be optimized weight, b4, and are the three optimizable bias coefficients for object recognition, and are two optimizable weights for object recognition, represents the probability that the i-th word node represents the starting point of the object entity, represents the probability that the i-th word node represents the end point of the object entity;

[0049] All possible entity pairs in the form of triples are extracted, and the entity pairs that satisfy the maximum likelihood estimation are output.

[0050] In a second aspect, a knowledge base establishment system for a service platform is provided, comprising:

[0051] Construction module: build corpus;

[0052] Recognition module: Input the sentences in the corpus into the constructed network model to identify entity-relationship pairs. The network model includes an input layer, a network feature extraction layer, and a relationship extraction layer. The input layer encodes the words and relationships in the sentence as input vectors. The feature extraction layer fuses the input vectors in the heterogeneous graph attention network and introduces gate mechanisms and residual connections to update word nodes and relationship nodes. The relationship extraction layer uses the word nodes m output by the feature extraction layer to extract the entity-relationship pairs. l and relationship node rl , identify entity-relationship pairs;

[0053] Knowledge graph building module: represents the identified entity relationship pairs as nodes and edges, and merges duplicate entities in all entity relationship pairs to form a knowledge graph;

[0054] Storage and update module: Use graph database tools to store entity relationship pairs in the knowledge graph as triples, and update the knowledge graph when the corpus updates text data.

[0055] According to a third aspect, a computer device is provided, comprising a processor and a memory; wherein, when the processor executes a computer program stored in the memory, the steps of the method for establishing a knowledge base of the service platform described in any one of the first aspects are implemented.

[0056] In a fourth aspect, a computer-readable storage medium is provided for storing a computer program; when the computer program is executed by a processor, the steps of the method for establishing a knowledge base of a service platform described in any one of the first aspects are implemented.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] (1) Introduction of heterogeneous graph attention network model: By embedding word nodes and relationship nodes together, the heterogeneous graph attention network is used to accurately extract entity and relationship information in the text. Especially when dealing with complex contexts and multi-level semantic structures, it can improve the accuracy and completeness of knowledge extraction and effectively solve the shortcomings of existing technologies in complex text processing.

[0059] (2) Automated text annotation and relationship extraction process: By using automated tools for text annotation and relationship extraction, the need for manual intervention is greatly reduced, and the efficiency of knowledge base construction is improved. This automated process performs well in large-scale data processing scenarios, significantly improving construction speed and optimizing resource utilization.

[0060] (3) Multi-layer feature fusion and extraction mechanism of deep learning models: In the model, multi-layer features of word nodes and relationship nodes are fused through deep learning technology, enabling the system to capture more complex features, improving the accuracy of relationship extraction and the learning ability of the model. This mechanism effectively solves the technical problem that traditional methods cannot fully understand complex contexts.

[0061] (4) Construction and update of dynamic knowledge graph: A dynamic knowledge graph is constructed, which can automatically expand and update the knowledge base according to the input of new data. This mechanism solves the problem of static knowledge base in traditional knowledge base, ensures the timeliness and scalability of knowledge base content, keeps the system knowledge up to date, and meets the needs of continuous data growth. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 This is a flow chart of the steps of the method for establishing a knowledge base of the service platform of the present invention;

[0063] Figure 2 It is a schematic diagram of updating word nodes in the heterogeneous graph of the present invention. DETAILED DESCRIPTION

[0064] The present invention will be further described below with reference to the accompanying drawings. The following examples are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. It should be noted that the term "comprising" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0065] Before explaining the embodiments of this solution, some terms that appear later are explained first:

[0066] Text crawling is an automated process used to collect large amounts of text data from web pages or databases on the internet. By simulating user browsing behavior, crawlers access websites and crawl the text content within them, extracting useful information such as articles, reviews, and product descriptions. This technology is commonly used to build datasets, gather information, and perform search engines. However, it must adhere to website usage protocols (such as robots.txt) and laws and regulations to ensure data acquisition is legal and compliant.

[0067] Natural language processing (NLP) is a field of computer science and artificial intelligence that aims to enable computers to understand, generate, and process human language. It encompasses tasks such as text analysis, word segmentation, part-of-speech tagging, sentiment analysis, and machine translation, helping computers extract useful information from unstructured text data and perform semantic understanding and reasoning. NLP is widely used in search engines, voice assistants, chatbots, and automatic translation.

[0068] Doccano is an open-source automated text annotation tool designed specifically for natural language processing (NLP) tasks. It supports tasks such as text classification, named entity recognition (NER), and relation extraction. Users can manually annotate text data using a simple graphical interface or automatically annotate using pre-trained models. Doccano supports importing and exporting data in multiple formats, such as JSON and CSV, facilitating subsequent model training and data processing. It is widely used to build high-quality training datasets.

[0069] The Heterogeneous Graph Attention Network (HGAT) is a deep learning model designed to process graph structures containing heterogeneous node and edge types. Unlike traditional graph neural networks, HGAT can handle complex graphs composed of multiple types of nodes (such as words and entities) and multiple relationships (such as associations and roles). By introducing an attention mechanism, HGAT dynamically adjusts the importance of different nodes and relationships and learns the semantic connections between them, thereby improving its ability to model complex data. It is widely used in tasks such as knowledge graph construction and relationship extraction.

[0070] BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture, proposed by Google. BERT learns deep semantic information about language through bidirectional context, effectively considering all words before and after a word to better understand its meaning. Pre-trained on large-scale text data, BERT can be fine-tuned for various natural language processing tasks, such as question-answering, text classification, and named entity recognition, significantly improving model performance and effectiveness.

[0071] Example 1

[0072] like Figure 1 As shown, a method for establishing a knowledge base of a service platform is provided, comprising the following steps:

[0073] S1: Build the corpus.

[0074] Specifically: S11: Utilize public databases and adopt text crawler technology to collect text data that is consistent with the current knowledge base construction goal.

[0075] Use text crawling technology to collect text data related to the service platform's topic from the internet, public databases, academic papers, industry reports, news articles, and relevant websites. These data sources can be tailored to the platform's needs, such as industry-specific portals, professional knowledge bases, or social media platforms. Crawling technology should comply with relevant laws and policies to ensure that the data is sourced legally and legitimately. Specific keywords, topics, and fields can also be set during the collection process to ensure that the data collected is consistent with the service platform's knowledge base objectives.

[0076] As an option, when establishing a service platform for traditional Chinese medicine knowledge, it is necessary to use text crawlers to obtain text data from authoritative medical literature, clinical guidelines and specifications, drug instructions and regulatory documents, electronic medical records (EMR) and real-world data, patent documents, patient forums and social media.

[0077] S12: Clean the collected text data.

[0078] First, remove HTML tags, special characters, spaces, and redundant punctuation marks in the text to ensure uniform data format. Second, check and delete duplicate records to ensure the uniqueness of each piece of data. Next, use natural language processing techniques for word segmentation and词性标注 (POS tagging), remove stop words (such as "的", "了", "是", etc.) and noise words irrelevant to the theme. Automatically correct spelling mistakes to ensure the correctness of the text. Subsequently, use regular expressions to identify and process abnormal text formats, such as dates, currency symbols, units, etc., to standardize them. Finally, convert the cleaned text into a unified encoding format (such as UTF-8) to ensure data consistency and compatibility for subsequent processing. Conduct a preliminary check on the data after cleaning to ensure that no key information is lost or damaged.

[0079] S13:借助自动化文本标注工具,对清理数据进行格式转化和序列标注,构建语料库。 S13: With the help of an automated text annotation tool, convert the format of the cleaned data and perform sequence annotation to build a corpus.

[0080] First, install and configure the doccano tool, and import the cleaned data into the doccano platform in JSON, CSV, or TXT format. Then, create a new project in doccano and define the annotation task type, such as named entity recognition (NER), text classification, or relation extraction, etc. Next, set the annotation tags according to the requirements of the service platform, such as entity tags (person, location, drug) or relation tags (association, function, ingredient, etc.). In the doccano interface, the data can be manually annotated by highlighting the text, or the automatic annotation function of doccano can be used to automatically generate preliminary annotations in combination with a pre-trained model or rules. After that, users can manually verify and correct the automatic annotation results to ensure the accuracy and consistency of the annotations. After completing the annotation, export the annotated data into the format required for model input (such as JSON or CSV) for subsequent model training and knowledge base construction.

[0081] S2: Input the sentences in the corpus into the constructed network model to identify entity relation pairs. The network model includes an input layer, a network feature extraction layer, and a relation extraction layer. The input layer encodes both the words and relations in the sentence as input vectors. The feature extraction layer fuses the input vectors in a heterogeneous graph attention network and introduces a gate mechanism and residual connections to update the word nodes and relation nodes. The relation extraction layer uses the word nodes m l and relation nodes r l , to identify entity relation pairs.

[0082] The constructed network model needs to be trained before use. The training of the network model can refer to existing technologies.

[0083] S21: input layer.

[0084] The model input is first of all word embedding operation. This model embeds word nodes and relationship nodes at the same time to form the model input. The process is to use the pre-trained BERT to encode the text. Assume that the input text set is T(x1,x2,…,x n ), where n represents the length of the entire text, x i To represent the i-th word, the entire text character token is input into BERT using the following formula to obtain the word node representation, and the predefined relationship label linear mapping layer is embedded:

[0085] [m1,m2,...,m n ]=BERT([x1,x2,...,x n ])

[0086] [r1,r2,...,r m ]=σ([a1,a2,...,a m ])

[0087] Among them, m i represents the encoded word embedding vector (i=1,2,…,n), r j represents the encoded relation embedding vector (i=1,2,…,m), σ represents the linear mapping, T(x1,x2,…,x n ) is the input text set, n represents the length of the entire text, x i Represents the i-th word, a1, a2, ..., a m is a predefined relationship label.

[0088] S22: Network feature extraction layer.

[0089] The word embedding operation input by the model will get m i and r j Two different types of node representations are connected to obtain [m i ,r j ], then uses the heterogeneous layers in the heterogeneous graph attention network model to incorporate diverse semantic information. In this layer, each node is considered a neighbor of another node, and the graph attention network is used to update the relationship between them. This structure uses a stacked layer to assign different weight coefficients to different nodes in the neighborhood based on the characteristics of the difference in attention between nodes, thereby improving the accuracy of relationship extraction.

[0090] Specifically, S221: embed the word output by the input layer into the vector m i and the relation embedding vector r jThey are used as word nodes and relationship nodes respectively to construct a heterogeneous graph.

[0091] The construction method of the heterogeneous graph refers to the existing technology.

[0092] S222: According to the similarity scores between the word node and its adjacent relationship nodes, weights are assigned to all adjacent relationship nodes of the word node to merge all relationship nodes and update the word node.

[0093] like Figure 2 As shown, specifically:

[0094] w ij =w1[m i ,r j ]+b1

[0095]

[0096]

[0097] Among them, w ij is the similarity score between the i-th word node and the j-th relationship node, The updated word node after fusion of all relationship nodes; [m i ,r j ] is the word embedding vector and relation embedding vector after feature concatenation; w1 is [m i ,r j ] weight coefficient; b1 is the first bias coefficient, exp(w ij ) represents w ij The exponential function, α ij is the probability of the attention weight of the i-th word node to the j-th relationship node, w2 is r j The weight coefficient is , N is the length of the current sentence text, and h is the number of relationship nodes actually associated with the current word node in the text.

[0098] S223: Update the relationship node based on all word nodes adjacent to the relationship node.

[0099] Specifically:

[0100]

[0101] Among them, q is the number of word nodes actually associated with the current relationship node in the text, w2′ is m i The weight coefficient of .

[0102] S224: Based on updated word nodes and relationship nodes The gate mechanism and residual connection are introduced to output the final word nodes and relation nodes.

[0103] In order to improve the ability of neural networks to capture complex features and maintain nonlinearity, a gate mechanism is added after the activation function. Specifically:

[0104] Based on the input word embedding vector m i and the updated word node Calculate the gate value e i ;

[0105]

[0106] Among them, w3 is The weight coefficient, sigmoid is the function, and b2 is the second bias term;

[0107] Using the gate value e i , update the word node and relationship node again, and get the updated word node and relationship nodes

[0108]

[0109] Add residual connection calculation to avoid the occurrence of gradient disappearance during training. The specific calculation is as follows:

[0110]

[0111] Among them, m l and r l They are the word nodes and relationship nodes output by the feature extraction layer respectively.

[0112] S23: Relation extraction layer.

[0113] The relationship extraction layer uses a two-stage subject-object entity identification method, breaking with the traditional approach of first extracting entities and then classifying relationships to form relationship pairs. The identifier uses a binary classifier to obtain the start and end positions of entities, better addressing the issue of nested entities and, to a certain extent, eliminating meaningless relationship groups.

[0114] By obtaining representations of word nodes and relationship nodes, the entity tagger can identify various possible entities (classical prescriptions and methods of taking medicine) within the word nodes, thereby deriving implicit relationship pairs within the text. Binary classifiers are well-utilized. They work by determining the start and end points of an entity and labeling each token with [0, 1]. A "1" in the start layer indicates that the token is the entity's starting point, while a "1" in the end layer indicates that the token is the entity's end point. Ultimately, combining these two layers of information to determine the entity.

[0115] The subject tagger uses binary classification to calculate the probability of a word node being identified as the start point and end point of the main entity, and continuously optimizes the likelihood function to identify the subject;

[0116]

[0117] in, is the probability that the i-th word node represents the starting point, represents the probability that the i-th word node represents the termination point, w s and w e are two optimizable weight coefficients for subject identification, b s and b e They are two optimizable bias coefficients for subject identification.

[0118] After the word nodes, relationship nodes and candidate main entities are fused, they are input into the object tagger. The object tagger uses binary classification to calculate the probability of the fused object entity being identified as the starting point and the ending point of the object entity, and continuously optimizes the likelihood function to identify the object;

[0119] m o =w4[m l ,r l ,δ k ]+b4

[0120]

[0121] Among them, m o is the fused object entity, δ k is the kth candidate main entity, w4 represents [m l ,r l ,δ k ] can be optimized weight, b4, and are the three optimizable bias coefficients for object recognition, and are two optimizable weights for object recognition, represents the probability that the i-th word node represents the starting point of the object entity, represents the probability that the i-th word node represents the end point of the object entity.

[0122] To determine the span of the main entity x in sentence T, the entity identifier is optimized and trained using the likelihood function formula. The specific formula is as follows:

[0123]

[0124] Among them, p θrepresents the probability distribution determined by the parameter θ, θ represents the model parameter, λ represents the boundary type of the entity (s(start, starting position), e(end, ending position)), represents the probability of the i-th token being the "start (s)" or "end (e)" position, f() represents the exponential function, Indicates that the i-th token is the starting / ending position of the main entity. Similarly, the same operation is performed on the guest entity to obtain p θ (o|x,T,r), that is, the likelihood function formula of the guest entity is the same as the likelihood function formula of the host entity.

[0125] All possible entity pairs in the form of triples are extracted, and the entity pairs (entity A, relation, entity B) that meet the maximum likelihood estimation are output.

[0126] That is, the loss function is as follows:

[0127]

[0128] Among them, L represents the loss function, D i Represents the main entity set in the training set, S represents the text feature, o represents the guest entity, and D represents the triple set in the training set.

[0129] For example, the text in "Treatise on Febrile Diseases" states that Mahuang Decoction, composed of ephedra, cinnamon twig, and apricot kernel, has the effect of diaphoresis and relieving exterior symptoms. The output relation pairs are: (Mahuang Decoction, composition, ephedra) and (Mahuang Decoction, effect, diaphoresis and relieving exterior symptoms). Mahuang Decoction is the subject, the relation is composition or effect, and the corresponding object entity is ephedra or diaphoresis and relieving exterior symptoms.

[0130] S3: Represent the identified entity relationship pairs as nodes and edges, and merge the duplicate entities in all entity relationship pairs to form a knowledge graph.

[0131] The constructed corpus is used as input, and the words and relations within the sentences are encoded simultaneously as input node vectors. The resulting word node and relationship node vectors are fed into the heterogeneous graph neural network layer for fused representation, enabling feature learning and extraction. By feeding the feature representations learned by the feature extraction layer into the relationship extraction layer, the entity attribute tagger can be used to identify associations between entities and medication attributes within the text, thereby determining entity and attribute relationship pairs throughout the entire text.

[0132] After extracting entity and relationship pairs, the next step is to organize the extracted information into a structured knowledge graph. First, the knowledge graph is constructed based on the identified entities (such as people, places, drugs, etc.) and the relationships between them (such as associations, functions, attributes, etc.). These entity and relationship pairs are represented as nodes and edges, where entities serve as nodes and relationships serve as edges connecting nodes. Next, these nodes and edges are combined to form a graph structure. During the construction process, duplicate entities can be merged and redundant information can be removed through techniques such as semantic similarity.

[0133] S4: Use graph database tools to store entity-relationship pairs in the knowledge graph as triples, and update the knowledge graph when the corpus updates text data.

[0134] Using graph databases (such as Neo4j or other graph database tools), these entities and relationships are stored as triples (entity A, relationship, entity B), making them easier to query and visualize. Furthermore, the knowledge graph can be dynamically expanded and updated as new data is added, ultimately forming a knowledge base that can be queried, analyzed, and reasoned about by the service platform, enabling systematic management and application of knowledge.

[0135] Example 2

[0136] A knowledge base building system for a service platform, comprising:

[0137] Construction module: build corpus;

[0138] Recognition module: Input the sentences in the corpus into the constructed network model to identify entity-relationship pairs. The network model includes an input layer, a network feature extraction layer, and a relationship extraction layer. The input layer encodes the words and relationships in the sentence as input vectors. The feature extraction layer fuses the input vectors in the heterogeneous graph attention network and introduces gate mechanisms and residual connections to update word nodes and relationship nodes. The relationship extraction layer uses the word nodes m output by the feature extraction layer to extract the entity-relationship pairs. l and relationship node r l , identify entity-relationship pairs;

[0139] Knowledge graph building module: represents the identified entity relationship pairs as nodes and edges, and merges duplicate entities in all entity relationship pairs to form a knowledge graph;

[0140] Storage and update module: Use graph database tools to store entity relationship pairs in the knowledge graph as triples, and update the knowledge graph when the corpus updates text data.

[0141] For more specific details about the above method, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be described again here.

[0142] Example 3

[0143] The present invention provides a computer device comprising a processor and a memory; wherein, when the processor executes a computer program stored in the memory, the steps of the method for establishing a knowledge base of the above-mentioned service platform are implemented.

[0144] For more specific details about the above method, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be described again here.

[0145] Example 4

[0146] The present invention provides a computer-readable storage medium for storing a computer program; when the computer program is executed by a processor, the steps of the method for establishing a knowledge base of the above-mentioned service platform are implemented.

[0147] For more specific details about the above method, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be described again here.

[0148] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments will be sufficient. The systems, devices, and storage media disclosed in the embodiments are described briefly because they correspond to the methods disclosed in the embodiments. For relevant details, refer to the method description.

[0149] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention or certain portions of the embodiments.

[0150] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A method for establishing a knowledge base of a service platform, characterized in that: include: Build a corpus; Input the sentences in the corpus into the constructed network model to identify entity-relationship pairs. The network model includes an input layer, a network feature extraction layer, and a relationship extraction layer. The input layer encodes the words and relationships in the sentence as input vectors. The feature extraction layer fuses the input vectors in the heterogeneous graph attention network and introduces a gate mechanism and residual connection to update the word nodes and relationship nodes. The relationship extraction layer uses the word node m output by the feature extraction layer to extract the entity-relationship pairs. l and relationship node r l , identify entity-relationship pairs; Represent the identified entity-relationship pairs as nodes and edges, and merge the duplicate entities in all entity-relationship pairs to form a knowledge graph; Using graph database tools, the entity-relationship pairs in the knowledge graph are stored as triples, and the knowledge graph is updated when the text data in the corpus is updated.

2. The method for establishing a knowledge base of a service platform according to claim 1, characterized in that: The construction of the corpus is specifically as follows: Using public databases and text crawler technology to collect text data consistent with the current knowledge base construction goals; Perform data cleaning on the collected text data; With the help of automated text annotation tools, the cleaned data is formatted and sequence labeled to build a corpus.

3. The method for establishing a knowledge base of a service platform according to claim 1, characterized in that: The input layer uses pre-trained BERT to simultaneously encode the words and relations in the sentence to obtain word embedding vectors and relationship embedding vectors; [m1,m2,...,m n ]=BERT([x1,x2,...,x n ]) [r1,r2,...,r m ]=σ([a1,a2,...,a m ]) Among them, m i represents the encoded word embedding vector (i=1,2,…,n), r j represents the encoded relation embedding vector (i=1,2,…,m), σ represents the linear mapping, T(x1,x2,…,x n ) is the input text set, n represents the length of the entire text, x i Represents the i-th word, a1, a2, ..., a m is a predefined relationship label.

4. The method for establishing a knowledge base of a service platform according to claim 1, characterized in that: The feature extraction layer fuses the input vectors in the heterogeneous graph attention network, specifically: Embed the word output of the input layer into the vector m i and the relation embedding vector r j Use them as word nodes and relationship nodes respectively to construct a heterogeneous graph; According to the similarity scores between the word node and its adjacent relationship nodes, weights are assigned to all adjacent relationship nodes of the word node to fuse all relationship nodes and update the word node; Update the relationship node based on all word nodes adjacent to the relationship node; Based on updated word nodes and relationship nodes The gate mechanism and residual connection are introduced to output the final word nodes and relation nodes.

5. The method for establishing a knowledge base of a service platform according to claim 4, characterized in that: According to the similarity scores between the word node and its adjacent relationship nodes, weights are assigned to all adjacent relationship nodes of the word node to merge all relationship nodes and update the word node. Specifically, w ij =w1[m i ,r j ]+b1 Among them, w ij is the similarity score between the i-th word node and the j-th relationship node, The updated word node after fusion of all relationship nodes; [m i ,r j ] is the word embedding vector and relation embedding vector after feature concatenation; w1 is [m i ,r j ] weight coefficient; b1 is the first bias coefficient, exp(w ij ) represents w ij The exponential function, α ij is the probability of the attention weight of the i-th word node to the j-th relationship node, w2 is r j The weight coefficient is , N is the length of the current sentence text, and h is the number of relationship nodes actually associated with the current word node in the text; The updating of the relationship node based on all word nodes adjacent to the relationship node is specifically as follows: Among them, q is the number of word nodes actually associated with the current relationship node in the text, w2′ is m i The weight coefficient of .

6. The method for establishing a knowledge base of a service platform according to claim 4, characterized in that: The updated word node and relationship nodes The gate mechanism and residual connection are introduced to output the final word nodes and relationship nodes, specifically: Based on the input word embedding vector m i and the updated word node Calculate the gate value e i ; Among them, w3 is The weight coefficient, sigmoid is the function, and b2 is the second bias term; Using the gate value e i , update the word node and relationship node again, and get the updated word node and relationship nodes Use residual connections to output the final word nodes and relationship nodes; Among them, m l and r l They are the word nodes and relationship nodes output by the feature extraction layer respectively.

7. The method for establishing a knowledge base of a service platform according to claim 1, characterized in that: The relationship extraction layer uses the word node m output by the feature extraction layer l and relationship node r l , identify entity-relationship pairs, specifically: The subject tagger uses binary classification to calculate the probability of a word node being identified as the start point and end point of the main entity, and continuously optimizes the likelihood function to identify the subject; in, is the probability that the i-th word node represents the starting point, represents the probability that the i-th word node represents the termination point, w s and w e are two optimizable weight coefficients for subject identification, b s and b e are two optimizable bias coefficients for subject identification; After the word nodes, relationship nodes and candidate main entities are fused, they are input into the object tagger. The object tagger uses binary classification to calculate the probability of the fused object entity being identified as the starting point and the ending point of the object entity, and continuously optimizes the likelihood function to identify the object; m o =w4[m l ,r l ,d k ]+b4 Among them, m o is the fused object entity, δ k is the kth candidate main entity, w4 represents [m l ,r l ,δ k ] can be optimized weight, b4, and are the three optimizable bias coefficients for object recognition, and are two optimizable weights for object recognition, represents the probability that the i-th word node represents the starting point of the object entity, represents the probability that the i-th word node represents the end point of the object entity; All possible entity pairs in the form of triples are extracted, and the entity pairs that satisfy the maximum likelihood estimation are output.

8. A knowledge base building system for a service platform, characterized in that: include: Construction module: build corpus; Recognition module: Input the sentences in the corpus into the constructed network model to identify entity-relationship pairs. The network model includes an input layer, a network feature extraction layer, and a relationship extraction layer. The input layer encodes the words and relationships in the sentence as input vectors. The feature extraction layer fuses the input vectors in the heterogeneous graph attention network and introduces gate mechanisms and residual connections to update word nodes and relationship nodes. The relationship extraction layer uses the word nodes m output by the feature extraction layer to extract the entity-relationship pairs. l and relationship node r l , identify entity-relationship pairs; Knowledge graph building module: represents the identified entity relationship pairs as nodes and edges, and merges duplicate entities in all entity relationship pairs to form a knowledge graph; Storage and update module: Use graph database tools to store entity relationship pairs in the knowledge graph as triples, and update the knowledge graph when the corpus updates text data.

9. A computer device, characterized in that: It comprises a processor and a memory; wherein, when the processor executes the computer program stored in the memory, it implements the steps of the method for establishing a knowledge base of the service platform described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that Used to store computer programs; when the computer programs are executed by the processor, the steps of the method for establishing a knowledge base of the service platform described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Relation extraction and knowledge graph construction method based on deep learning model

    CN110598000A

  • Knowledge graph construction method based on Chinese electronic medical records

    CN113688255A

  • Knowledge graph construction method and device for electric power operation text, medium and chip

    CN118469006A

  • Knowledge graph construction and intelligent question and answer method and device based on deep learning

    CN119691135A

  • Heterogeneous graph neural network-based automobile part supply chain entity relationship extraction method

    CN119808785A