Knowledge extraction method based on bert algorithm and related equipment

By applying a method based on BERT algorithm in knowledge extraction, the problems of low accuracy and low efficiency in the prior art are solved, and efficient and accurate knowledge extraction is achieved, which is suitable for different fields and text data.

CN120069025APending Publication Date: 2025-05-30BEIJING CHINA POWER INFORMATION TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510108992.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing knowledge extraction methods have problems such as insufficient accuracy and low efficiency, making it difficult to deal with complex language structures and adapt to diversity in different fields.

Method used

The knowledge extraction method based on the BERT algorithm is adopted, and the text data is preprocessed and formatted, and the pre-trained BERT model is input for fine-tuning training, the knowledge extraction result data is output, and cluster-based knowledge fusion and storage are carried out.

Benefits of technology

Improve the accuracy and efficiency of knowledge extraction, and can quickly extract useful knowledge from large-scale text data, reduce manual intervention and reduce costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069025A_ABST
    Figure CN120069025A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge extraction method based on a bert algorithm and related equipment. The method comprises the steps of obtaining text data of knowledge to be extracted; comprising any one of structured, semi-structured or unstructured data; performing preprocessing, word segmentation processing, labeling processing and format conversion processing on the text data of the to-be-extracted knowledge to obtain vector data of the to-be-extracted knowledge; the vector data of the knowledge to be extracted comprises word vector data and word vector data; inputting the vector data of the knowledge to be extracted into a pre-trained bert model, and outputting knowledge extraction result data; the pre-trained bert model is a bert model obtained through fine tuning training; performing clustering-based knowledge fusion on the knowledge extraction result data; storing the data after knowledge fusion into a relational database; the relational database has a knowledge query interface. The knowledge extraction accuracy and efficiency can be improved, and useful knowledge can be rapidly extracted from large-scale text data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of knowledge extraction, and in particular, to a knowledge extraction method and related devices based on the BERT algorithm. Background Art

[0002] Knowledge extraction is the process of automatically extracting useful knowledge from source texts or structured data. Knowledge extraction techniques include entity recognition, relation extraction, event extraction, etc., which are used to extract information such as entities, relations, and attributes from large-scale data sources to assist in constructing knowledge graphs.

[0003] In related technologies, knowledge extraction methods include rule- and template-based methods, machine learning-based methods, etc. The knowledge extraction methods in related technologies have problems of low accuracy and low efficiency. Summary of the Invention

[0004] In view of this, the purpose of this application is to propose a knowledge extraction method and related devices based on the BERT algorithm.

[0005] Based on the above purpose, this application provides a knowledge extraction method based on the BERT algorithm, including:

[0006] Obtain the text data of the knowledge to be extracted; the text data includes any one of structured, semi-structured, or unstructured data;

[0007] Preprocess the text data of the knowledge to be extracted, perform word segmentation, annotation, and format conversion processing to obtain the vector data of the knowledge to be extracted; the vector data of the knowledge to be extracted includes word vector data and character vector data;

[0008] Input the vector data of the knowledge to be extracted into a pre-trained BERT model to output the knowledge extraction result data; the pre-trained BERT model is a BERT model obtained through fine-tuning training;

[0009] Perform knowledge fusion based on clustering on the knowledge extraction result data;

[0010] Store the data after knowledge fusion in a relational database; the relational database has a knowledge query interface.

[0011] In some embodiments, the knowledge extraction result data includes the knowledge extraction result data of labeled entities; the performing knowledge fusion based on clustering on the knowledge extraction result data includes:

[0012] For the knowledge extraction result data of multiple labeled entities for the same task, determine the entity category of the knowledge extraction result data of the labeled entity according to the similarity between the knowledge extraction result data of each labeled entity and the existing entity categories respectively;

[0013] Determine whether the knowledge extraction result data of multiple labeled entities corresponding to the same entity category correspond to the same object;

[0014] In response to determining that the multiple knowledge extraction result data of the labeled entities corresponding to the same entity category correspond to the same object, construct an alignment relationship for the knowledge extraction result data of the multiple labeled entities and perform knowledge fusion.

[0015] In some embodiments, the determining the entity category of the knowledge extraction result data according to the similarity between each knowledge extraction result data and the existing entity categories includes:

[0016] Based on the first database, construct a candidate set of entity categories for the knowledge extraction result data of each labeled entity; the candidate set of entity categories includes multiple entity categories;

[0017] Based on the second database, respectively obtain the original text data corresponding to multiple entity entries in the candidate set of entity categories;

[0018] Calculate the similarity between the knowledge extraction result data of the labeled entity and the original text data corresponding to each entity entry in the candidate set of entity categories respectively, to obtain multiple similarities; the multiple similarities correspond to the multiple entity categories one by one;

[0019] Determine multiple target similarities that are not less than the threshold among the multiple similarities, and determine the entity category corresponding to the target similarity with the largest value as the entity category of the knowledge extraction result data of the labeled entity.

[0020] In some embodiments, the method further includes training the pre-trained bert model through the following method:

[0021] Obtain a training data set of knowledge to be extracted;

[0022] Preprocess the text data in the training data set, perform word segmentation, annotation, and format conversion processing to obtain a training vector data set of knowledge to be extracted; the training vector data set of knowledge to be extracted includes the training vector data of the knowledge to be extracted of multiple labeled entities;

[0023] Fine-tune and train the pre-trained bert model based on the vector data set of knowledge to be extracted;

[0024] Among them, the fine-tuning training includes:

[0025] Obtain the encoded sequence of the training vector data of the knowledge to be extracted for multiple labeled entities;

[0026] For each encoded sequence:

[0027] Randomly sample labeled entities from the encoded sequences in the same batch;

[0028] According to the labeled entities obtained by the random sampling, extract the encoded vectors corresponding to the head and tail in the encoded sequence;

[0029] Perform normalization processing on the encoded sequence with the encoded vector as the condition.

[0030] In some embodiments, the fine-tuning training further includes:

[0031] Test the pre-trained bert model that has been trained using the validation data set at preset time intervals, and calculate the evaluation metrics of the pre-trained bert model that has been trained to obtain the evaluation metrics of multiple tests; the evaluation metrics include at least one of accuracy, recall rate, and F1 value;

[0032] In response to determining that the difference between the evaluation metrics of the multiple tests does not exceed the preset threshold, determine that the pre-trained bert model is obtained.

[0033] In some embodiments, the preprocessing of the text data of the knowledge to be extracted includes: removing HTML tags, special characters, and duplicate content in the text data of the knowledge to be extracted;

[0034] The annotation processing includes: annotating at least one of entities, relationships, and events;

[0035] The format conversion processing includes converting the words in the text data after annotation processing into word vectors, and performing at least one of padding processing and truncation processing to make the text data of the knowledge to be extracted have a preset length;

[0036] Before performing knowledge fusion based on clustering on the knowledge extraction result data, it further includes:

[0037] Perform deduplication processing and correction processing on the knowledge extraction result data.

[0038] In some embodiments, the determining the entity category of the knowledge extraction result data of the labeled entity according to the similarity between the knowledge extraction result data of each labeled entity and the existing entity categories further includes:

[0039] In response to determining that all the similarities are less than the threshold, determine that the entity category of the knowledge extraction result data of the labeled entity is an unknown entity category, and add a new entity category to the first database;

[0040] Based on the knowledge extraction result data of multiple labeled entities of the unknown entity category, calculate the pairwise similarity of the knowledge extraction result data of multiple labeled entities of each unknown entity category;

[0041] In response to determining that the pairwise similarity is greater than the similarity threshold, merge the knowledge extraction result data of the corresponding two labeled entities into the same newly added entity category.

[0042] An embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in any one of the foregoing is implemented.

[0043] An embodiment of the present application further provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method described in any one of the foregoing.

[0044] An embodiment of the present application further provides a computer program product, including computer program instructions. When the computer program instructions run on a computer, the computer is caused to execute the method described in any one of the foregoing.

[0045] As can be seen from the above, the knowledge extraction method based on the bert algorithm provided by the present application obtains the text data of the knowledge to be extracted; the text data includes any one of structured, semi-structured or unstructured data; preprocesses the text data of the knowledge to be extracted, performs word segmentation processing, annotation processing, and format conversion processing to obtain the vector data of the knowledge to be extracted; the vector data of the knowledge to be extracted includes word vector data and character vector data; inputs the vector data of the knowledge to be extracted into a pre-trained bert model, and outputs knowledge extraction result data; the pre-trained bert model is a bert model obtained through fine-tuning training; performs knowledge fusion based on clustering on the knowledge extraction result data; stores the data after knowledge fusion in a relational database; the relational database has a knowledge query interface; can improve the accuracy and efficiency of knowledge extraction, and can quickly extract useful knowledge from large-scale text data. Description of the Drawings

[0046] In order to more clearly illustrate the technical solutions in the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1aFlowchart of the knowledge extraction method based on the BERT algorithm according to the embodiments of the present application;

[0048] Figure 1b Flowchart of the training method of the BERT model according to the embodiments of the present application;

[0049] Figure 2 Schematic flowchart of the knowledge fusion method according to the embodiments of the present application;

[0050] Figure 3 Schematic diagram of multi-source knowledge fusion according to the embodiments of the present application;

[0051] Figure 4 Schematic diagram of the electronic device according to the embodiments of the present application. Detailed implementation manners

[0052] To make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0053] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be the ordinary meanings understood by those of ordinary skill in the art to which the present application belongs. The "first", "second" and similar terms used in the embodiments of the present application do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0054] The technical solution of the knowledge extraction method based on rules and templates is as follows: By manually defining a series of rules and patterns, the text is matched and analyzed to extract specific knowledge. Usually, some patterns of keywords or phrases are defined, and when these patterns appear in the text, the corresponding knowledge is extracted. For example, when extracting company names and location information, the rule can be defined as "the company name is usually at the beginning of the sentence, followed by keywords such as 'located in' or'situated in', and then the location information". Then, by analyzing the text sentence by sentence, the company name and location are extracted according to this rule.

[0055] In addition, rule- and template-based knowledge extraction methods can identify Patterns from seed data through BootStrap for extracting more data and extracting more Patterns. Ontology-based extraction conducts knowledge mining through reasoning, mainly including technologies such as PRA (graph-based extraction) and TransE series (Embedding-based extraction).

[0056] Machine learning-based methods can include traditional machine learning-based knowledge extraction methods and deep learning-based knowledge extraction methods. Among them, the technical solution of traditional machine learning-based knowledge extraction methods is to use feature engineering combined with machine learning algorithms for knowledge extraction. First, various features of the text are manually extracted, such as word frequency, part of speech, syntactic structure, etc., and then these features are input into machine learning algorithms, such as support vector machines (SVM) (Logistic Model), random forests (e.g., conditional random fields (CRF)), etc., for training and prediction to extract knowledge. For example, for entity recognition tasks, morphological features, context features, etc. of the text can be extracted, and then SVM is used for training to identify entities such as person names, place names, and organization names in the text.

[0057] Among them, deep learning-based knowledge extraction methods can be, for example, DeepKE. Among them, DeepKE is a deep learning-based knowledge graph construction framework that supports complex knowledge graph construction scenarios such as single-sentence / multi-sentence passages, multi-modal, and low-resource, including entity, relationship, and attribute extraction tasks. In addition, DeepKE also provides a multi-modal entity relationship extraction model IFAformer based on Transformer, which can concatenate the multi-head attention key values of context and visual features at each Transformer layer, thereby implicitly aligning the features of objects in the text and related images.

[0058] Rule- and template-based methods, machine learning-based methods, etc. have problems such as low accuracy, low efficiency, difficulty in handling complex language structures, and poor generality. Among them, the low accuracy is mainly reflected in that rule- and template-based knowledge extraction methods rely heavily on manually defined rules, and the rules often fail to cover all situations. For complex language structures and diverse text contents, missing extraction or incorrect extraction is likely to occur. For example, some new expressions or special language phenomena may not be recognized by existing rules. Although traditional machine learning-based knowledge extraction methods can automatically learn features to a certain extent, due to the need for manual participation in feature engineering, they are easily affected by human factors, and the extracted features are not comprehensive and accurate enough, resulting in limited accuracy of knowledge extraction.

[0059] Among them, the low efficiency is mainly reflected in that: when dealing with large-scale text data, the rule- and template-based knowledge extraction method needs to match rules one by one for each text, with a large amount of calculation and slow speed. Moreover, when the rules need to be modified or extended, the entire dataset needs to be processed again, with poor flexibility. The knowledge extraction method of traditional machine learning requires a large amount of time and manpower in the feature engineering stage, and the training and prediction processes are also time-consuming, making it difficult to meet the real-time processing requirements of large-scale text data.

[0060] Among them, the difficulty in dealing with complex language structures is mainly as follows: for some complex sentence structures and nested entity relationships, etc., existing technologies (such as rule- and template-based methods and traditional machine learning methods, etc.) often have difficulty accurately extracting knowledge. For example, in a sentence containing multiple clauses and modifiers, it may be very difficult to determine the correct relationship between entities. For some ambiguous or vague language expressions, existing technologies also have difficulty making accurate judgments. For example, a word may have multiple meanings and require different interpretations in different contexts, but existing technologies (such as rule-based methods and traditional machine learning methods, etc.) may not be able to accurately identify the context and perform correct knowledge extraction.

[0061] Among them, the poor generality is mainly reflected in that: the rule- and template-based knowledge extraction method and the knowledge extraction method of traditional machine learning are usually designed for specific tasks and domains. When applied to different tasks or domains, the rules need to be redesigned or a large number of feature engineering adjustments need to be made, with poor generality. Texts in different domains have different characteristics and language styles, and it is difficult for existing technologies to adapt to this diversity and achieve good knowledge extraction effects in different domains.

[0062] Natural language processing technology is used to process and understand natural language text. In a knowledge graph, NLP technology can be used for tasks such as text parsing, semantic analysis, named entity recognition, etc., so as to extract and understand knowledge. The latest technologies in current NLP are still constantly evolving, and some important progress includes: Pre-trained language models: These models can learn rich language knowledge and structures through pre-training on a large corpus, and thus perform well in various natural language processing tasks. Representative pre-trained language models include bert, GPT series, etc. Knowledge graph embedding technology: By converting the information of entities, relationships, attributes, etc. in the knowledge graph into vector forms and embedding them into a low-dimensional space, the visualization and computability of the knowledge graph can be realized. Semantic role labeling and dependency syntactic analysis: These technologies can be used to analyze the syntactic structure and semantic relationships in a sentence, so as to better understand the text meaning. Cross-lingual natural language processing: With the development of globalization, cross-lingual natural language processing has become increasingly important. This technology can be used to implement tasks such as machine translation, text classification, sentiment analysis, etc.

[0063] In terms of historical application technologies, the research and application of NLP have gone through multiple stages. Early NLP technologies were mainly based on rules and manual feature engineering, such as lexical analysis, syntactic analysis, etc. With the development of machine learning and deep learning technologies, NLP technologies based on statistical learning and neural networks have emerged, such as N-gram models, hidden Markov models, maximum entropy models, support vector machines, etc. The application scope of these technologies is extensive, including machine translation, automatic summarization, text classification, sentiment analysis, etc.

[0064] NLP is a key technology in knowledge graphs. It can help machines understand and process human language, thus realizing automated and intelligent knowledge management and reasoning. With the continuous development of technology, the application prospects of NLP will be even broader.

[0065] Knowledge fusion is an important research direction in the field of knowledge engineering. It involves collecting, processing, representing, integrating, and reasoning about knowledge from different sources and domains to achieve in-depth understanding and application of knowledge. This article will introduce several key aspects of knowledge fusion, including knowledge acquisition, knowledge representation, knowledge integration, knowledge reasoning, and knowledge application. Knowledge acquisition is the first step in knowledge fusion, and its task is to extract useful knowledge from various data sources. Data sources can include text, images, audio, video, etc., and knowledge acquisition methods include techniques such as text mining, image recognition, and natural language processing. Through these methods, various forms of data can be transformed into structured knowledge, providing a basis for subsequent fusion and reasoning. Knowledge representation is the process of describing and expressing knowledge in a certain form. Effective knowledge representation can reduce the difficulty of understanding and using knowledge and improve the availability and reusability of knowledge. Commonly used knowledge representation methods include semantic networks, ontology models, natural language processing, etc. They can integrate and associate scattered knowledge points to form a logical and hierarchical knowledge structure. Knowledge integration is the process of integrating and unifying knowledge from different sources and in different forms. It can transform and map various forms of knowledge so that knowledge from different sources can be interconnected and complementary. At the same time, through knowledge integration, contradictions and conflicts between knowledge can also be discovered, further optimizing and correcting the quality of knowledge. Knowledge reasoning is the thinking process of deriving new knowledge based on known knowledge. In knowledge fusion, through knowledge reasoning, the associations and laws between knowledge can be discovered, further expanding and deepening the understanding of knowledge. Commonly used knowledge reasoning methods include rule-based reasoning, model-based reasoning, case-based reasoning, etc., which can achieve automated reasoning of knowledge. Knowledge application is the process of applying the acquired, represented, integrated, and reasoned knowledge to practical problems. Through the analysis and modeling of practical problems, the acquired knowledge can be used for problem-solving and decision-making support. At the same time, through knowledge application, the accuracy and availability of knowledge can also be further tested, providing feedback for subsequent knowledge acquisition and optimization.

[0066] Knowledge fusion is one of the important means to achieve intelligent decision-making. Through research and application in aspects such as knowledge acquisition, representation, integration, reasoning, and application, in-depth understanding and application of knowledge can be realized. With the continuous development of technology and the continuous expansion of application scenarios, knowledge fusion will be widely applied and promoted in more fields.

[0067] Based on this, the embodiments of the present application provide a knowledge extraction method and related devices based on the BERT algorithm. In terms of efficiency, aiming at the problem of low efficiency of traditional knowledge extraction methods when dealing with large-scale text data, the advantages of the BERT algorithm are fully utilized. By optimizing the process and parameter settings, a large amount of text can be quickly processed in a short time. Whether it is news, papers, or social media content and other data sources, the required knowledge can be quickly extracted from them. In terms of accuracy, given that traditional methods are prone to inaccurate knowledge extraction, by leveraging the powerful language understanding ability of the BERT algorithm, knowledge elements such as entities, relationships, and events in the text can be accurately identified. After fine data preprocessing, model training, and post-processing, the extraction results are highly accurate and reliable. At the same time, automatic extraction can be achieved, reducing manual intervention, constructing an automated knowledge extraction system, without a large amount of manual annotation and screening, reducing costs and improving convenience. Therefore, it can solve to a certain extent the problems of low accuracy and low efficiency in the existing knowledge extraction methods.

[0068] An embodiment of the present invention provides a knowledge extraction method based on the BERT algorithm, as Figure 1a shown, the knowledge extraction method based on the BERT algorithm may include:

[0069] S100, obtaining text data of knowledge to be extracted; the text data includes any one of structured, semi-structured, or unstructured data;

[0070] S200, preprocessing the text data of the knowledge to be extracted, performing word segmentation, annotation, and format conversion processing to obtain vector data of the knowledge to be extracted; the vector data of the knowledge to be extracted includes word vector data and character vector data;

[0071] S300, inputting the vector data of the knowledge to be extracted into a pre-trained BERT model, and outputting knowledge extraction result data; the pre-trained BERT model is a BERT model obtained through fine-tuning training;

[0072] S400, performing knowledge fusion on the knowledge extraction result data based on clustering and entity connection;

[0073] S500, storing the data after knowledge fusion into a relational database; the relational database has a knowledge query interface.

[0074] The knowledge extraction method based on the BERT algorithm provided by the embodiments of this application is committed to achieving efficient and accurate automatic extraction of useful knowledge from large-scale text data, so as to provide solid support for various application scenarios. It aims to serve a wide range of application scenarios, such as quickly and accurately answering users' questions in an intelligent question-answering system, improving the retrieval efficiency and quality in the field of information retrieval, providing comprehensive and accurate knowledge and data analysis results for decision-making support to reduce risks and improve scientificity. It can also be applied to multiple fields such as knowledge graph construction, text analysis, and machine translation, and can meet the knowledge extraction needs of different fields and industries.

[0075] In some of the embodiments, in step S100, the text data of the knowledge to be extracted can be the text data of the knowledge to be extracted for the same task. The task can include tasks such as entity annotation tasks, relationship annotation tasks, and event annotation.

[0076] In some of the embodiments, in step S200, the preprocessing can be data cleaning. The preprocessing of the text data of the knowledge to be extracted can include: removing HTML tags, special characters, and duplicate content, etc. from the text data of the knowledge to be extracted. In this way, the noise data in the text data of the knowledge to be extracted can be removed.

[0077] In some of the embodiments, the word segmentation processing can include using a word segmentation tool to perform word segmentation on the text data of the knowledge to be extracted after cleaning, so as to split the text data of the knowledge to be extracted into individual words. For example, Jieba can be used for word segmentation.

[0078] In some of the embodiments, the annotation processing can include: annotating at least one of entities, relationships, and events. That is, at least one of the knowledge elements such as entities, relationships, and events is annotated for the text data of the knowledge to be extracted after word segmentation.

[0079] In some of the embodiments, the annotation processing can include a first annotation processing and a second annotation processing. Among them, the first annotation processing can be performed using an automatic annotation tool for preliminary annotation. The second annotation processing can be understood as the review and correction of the first annotation processing. The second annotation processing can improve the annotation efficiency and accuracy. Specifically, an annotation task can be created, and the data after preliminary annotation can be viewed manually, etc., and the results of automatic annotation can be modified and new annotations can be added according to the actual situation.

[0080] In some of these embodiments, the format conversion process may include converting the words in the text data after annotation processing into word vectors, and performing at least one of padding processing and truncation processing, so that the text data for which knowledge is to be extracted has a preset length. Specifically, the text data after annotation processing can be converted into a format suitable for input to the bert model. The Word2Vec tool for word vectors is used to convert the words in the text data after annotation processing into corresponding word vectors, and processing such as padding and truncation is performed. In this way, the lengths of the input data can be kept consistent. Generally, after the format conversion process, each sentence in the text data for which knowledge is to be extracted can be represented as a word vector and a character vector.

[0081] In some of these embodiments, in step S300, the training of the pre-trained bert model may include optimizing the model parameters using an optimization algorithm, setting appropriate parameters such as the learning rate, batch size, and number of training epochs. And save the model parameters regularly during the training process for subsequent evaluation and application. Regularly evaluate the model using the validation set data, and observe the changes in the model performance using metrics such as accuracy, recall, and F1 value. If the model performance no longer improves, the training can be stopped or the parameters can be adjusted and retrained. The pre-trained bert model can be a bert-crf model.

[0082] In some of these embodiments, as Figure 1b shown, the training method of the pre-trained bert model may include:

[0083] Step S321, obtain the training data set of the knowledge to be extracted.

[0084] Step S322, perform preprocessing, word segmentation processing, annotation processing, and format conversion processing on the text data in the training data set to obtain the training vector data set of the knowledge to be extracted; the training vector data set of the knowledge to be extracted includes multiple pieces of training vector data of the knowledge to be extracted with annotated entities.

[0085] Step S323, perform fine-tuning training on the pre-trained bert model based on the vector data set of the knowledge to be extracted.

[0086] In some of these embodiments, in step S322, the preprocessing, word segmentation processing, annotation processing, and format conversion processing may be the same as the preprocessing, word segmentation processing, annotation processing, and format conversion processing in the aforementioned step S200. Details are not described herein again.

[0087] In some of these embodiments, in step S323, the fine-tuning training may include:

[0088] Obtain the encoded sequence of the training vector data of multiple pieces of knowledge with labeled entities. Usually, the training vector data of the knowledge to be extracted can be passed into the encoder of BERT, and appropriate parameters such as the learning rate, batch size, and number of training epochs are set to obtain the encoded sequence.

[0089] For each encoded sequence, perform the following same operations respectively:

[0090] Randomly sample the labeled entities in the encoded sequences of the same batch. That is, randomly sample a labeled s (i.e., the main entity, which can be a segment in the sentence) during training. And during prediction, traverse all s in the encoded sequences of the same batch one by one.

[0091] According to the labeled entities obtained by the random sampling, extract the encoded vectors corresponding to the head and tail of the encoded sequence. Specifically, according to the currently randomly sampled s, extract the encoded vectors corresponding to the head and tail of s from the encoded sequence.

[0092] Taking the encoded vectors as conditions, perform normalization processing on the encoded sequence. Specifically, the feature vector of the head entity (i.e., the encoded vector corresponding to the head of s) can be used as the conditional feature to dynamically adjust the feature vector representation of the tail entity (i.e., the encoding corresponding to the tail of s), that is, to achieve conditional normalization. Conditional normalization can be performed on the word vector or sample scale. By dynamically adjusting the normalization parameters (such as the bias β and the scaling factor γ), the distribution of features can be made more stable and consistent, and at the same time, the convergence speed of the algorithm can be improved.

[0093] In some embodiments, the fine-tuning training may further include step S324:

[0094] Test the pre-trained BERT model that has been trained using the validation dataset every preset time duration, and calculate the evaluation metrics of the pre-trained BERT model that has been trained to obtain the evaluation metrics of multiple tests. Usually, the evaluation metrics may include at least one of accuracy, recall, and F1 value. Here, the preset time duration can be understood as a fixed time duration. In this way, testing every preset time duration can be understood as regularly.

[0095] In response to determining that the difference between the evaluation metrics of the multiple tests does not exceed the preset threshold, determine that the pre-trained BERT model is obtained. That is, stop the training of the BERT model when the performance of the BERT model no longer improves. Or in response to determining that the difference between the evaluation metrics of the multiple tests does not exceed the preset threshold, adjust the parameters and return to step S323 to start training the BERT model again.

[0096] In this way, it can make the BERT model efficient and accurate in the knowledge extraction task.

[0097] In some of these embodiments, in step S400, before performing clustering-based knowledge fusion on the knowledge extraction result data, post-processing may also be performed on the output of the bert model (i.e., the knowledge extraction result data). That is, the clustering-based knowledge fusion of the knowledge extraction result data may include: performing duplicate removal processing and correction processing on the knowledge extraction result data. Generally, error detection and correction of the knowledge extraction result data can be performed through an external knowledge base or manual review. In this way, the accuracy and integrity of the knowledge extraction result data (i.e., knowledge) can be improved.

[0098] In some of these embodiments, as Figure 2 shown, a knowledge fusion method based on the clustering SinglePass algorithm can be used to perform knowledge fusion on the extracted knowledge results (i.e., knowledge extraction result data). The SinglePass algorithm is a streaming clustering algorithm. Each sample only participates in sample clustering once and has a certain dependence on the order of the samples. Its basic principle is that for an unknown new sample, if it is similar enough to an existing class (such as a person's name, a place name, an organization), then it is put into this class, otherwise it forms a new class by itself.

[0099] In some of these embodiments, the knowledge extraction result data may include knowledge extraction result data with labeled entities. The clustering-based knowledge fusion of the knowledge extraction result data may include:

[0100] For the knowledge extraction result data of multiple labeled entities for the same task, respectively determine the entity category of the knowledge extraction result data of the labeled entity according to the similarity between the knowledge extraction result data of each labeled entity and the existing entity categories. In this way, the knowledge extraction result data of each labeled entity under the same task id can be obtained.

[0101] Determine whether the knowledge extraction result data of multiple labeled entities corresponding to the same entity category correspond to the same object. It can be understood as judging whether two or more entities from different information sources point to the same object.

[0102] In response to determining that the knowledge extraction result data of multiple labeled entities corresponding to the same entity category correspond to the same object, as Figure 3 shown, construct an alignment relationship for the knowledge extraction result data of the multiple labeled entities and perform knowledge fusion. Knowledge fusion is to collect, process, represent, integrate, and reason about knowledge from different sources and different fields to achieve in-depth understanding and application of knowledge.

[0103] In this way, it is possible to fuse and aggregate information among multiple entities representing the same object. For example, various forms of knowledge (knowledge extraction result data) can be transformed and mapped so that knowledge from different sources (knowledge extraction result data) can be related to each other.

[0104] In some embodiments, determining the entity category of the knowledge extraction result data according to the similarity between each piece of knowledge extraction result data and existing entity categories may include:

[0105] Based on the first database, construct a candidate set of entity categories for the knowledge extraction result data of each labeled entity; the candidate set of entity categories includes multiple entity categories. Generally, the first database can be an alias library, which may include entities and various names corresponding to the entities (such as person names or place names, etc.).

[0106] Based on the second database, respectively obtain the original text data corresponding to multiple entity entries in the candidate set of entity categories. Generally, the second database can be a knowledge base, which may contain entity entries and the original text data (i.e., sentences, etc.) corresponding to the entity entries.

[0107] Respectively calculate the similarity between the knowledge extraction result data of the labeled entity and the original text data corresponding to each entity entry in the candidate set of entity categories to obtain multiple similarities. The multiple similarities correspond one by one to the multiple entity categories.

[0108] Determine multiple target similarities among the multiple similarities that are not less than the threshold, and determine the entity category (id) corresponding to the target similarity with the largest value as the entity category (id) of the knowledge extraction result data of the labeled entity. Or determine that all the multiple similarities are less than the threshold, and add a new category (id) for the knowledge extraction result data of this labeled entity.

[0109] In this way, the entity ids of all knowledge extraction result data of unlabeled categories can be determined according to the similarity between all knowledge extraction result data of unlabeled categories and existing entity ids.

[0110] In some embodiments, determining the entity category of the knowledge extraction result data of the labeled entity according to the similarity between each piece of knowledge extraction result data of the labeled entity and existing entity categories may further include:

[0111] In response to determining that all the multiple similarities are less than the threshold, determine the entity category of the knowledge extraction result data of the labeled entity as an unknown entity category, and add a new entity category to the first database;

[0112] Based on the knowledge extraction result data of multiple labeled entities of the unknown entity category, calculate the pairwise similarity of the knowledge extraction result data of multiple labeled entities of each unknown entity category;

[0113] In response to determining that the pairwise similarity is greater than the similarity threshold, merge the knowledge extraction result data of the corresponding two labeled entities into the same newly added entity category.

[0114] In some embodiments, in step S500, the relational database may have a designed table structure, indexes, etc. The knowledge query interface may have visual controls, and by triggering the visual controls, corresponding knowledge pop-ups can be popped up, thus facilitating users to use knowledge.

[0115] It should be noted that the method of the embodiment of the present application can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiment of the present application, and these multiple devices will interact with each other to complete the described method.

[0116] It should be noted that some embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0117] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the knowledge extraction method based on the BERT algorithm described in any one of the above embodiments.

[0118] Figure 4 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0119] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0120] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0121] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0122] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or can also achieve communication through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).

[0123] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0124] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification and does not necessarily include all the components shown in the figure.

[0125] The electronic device of the above embodiment is used to implement the corresponding knowledge extraction method based on the BERT algorithm in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0126] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application further provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the knowledge extraction method based on the BERT algorithm as described in any of the foregoing embodiments.

[0127] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0128] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the knowledge extraction method based on the BERT algorithm as described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0129] Based on the same inventive concept, corresponding to the knowledge extraction method based on the BERT algorithm described in any of the above embodiments, the present disclosure further provides a computer program product including computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of the computer to cause the computer and / or the processor to execute the knowledge extraction method based on the BERT algorithm. Corresponding to the execution subject corresponding to each step in each embodiment of the knowledge extraction method based on the BERT algorithm, the processor executing the corresponding step can belong to the corresponding execution subject.

[0130] The computer program product of the above embodiment is used to cause the computer and / or the processor to execute the knowledge extraction method based on the BERT algorithm as described in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0131] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and for the sake of brevity, they are not provided in detail.

[0132] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application may be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0133] Although the present application has been described in connection with specific embodiments of the present application, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0134] The embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of the present application shall be included within the protection scope of the present application.

Claims

1. A knowledge extraction method based on BERT algorithm, characterized in that: include: Obtain text data of knowledge to be extracted; The text data includes any one of structured, semi-structured or unstructured data; Preprocessing, word segmentation, tagging and format conversion are performed on the text data of the knowledge to be extracted to obtain vector data of the knowledge to be extracted; the vector data of the knowledge to be extracted includes word vector data and character vector data; Input the vector data of the knowledge to be extracted into the pre-trained BERT model, and output the knowledge extraction result data; the pre-trained BERT model is a BERT model obtained through fine-tuning training; Performing clustering-based knowledge fusion on the knowledge extraction result data; Storing the knowledge-fused data in a relational database; The relational database has a knowledge query interface.

2. The knowledge extraction method based on BERT algorithm according to claim 1 is characterized in that: The knowledge extraction result data includes knowledge extraction result data of the labeled entity; The clustering-based knowledge fusion of the knowledge extraction result data includes: For multiple pieces of knowledge extraction result data of annotated entities of the same task, determine the entity category of the knowledge extraction result data of the annotated entity according to the similarity between each piece of knowledge extraction result data of the annotated entity and the existing entity category; Determine whether the knowledge extraction result data of multiple labeled entities corresponding to the same entity category correspond to the same object; In response to determining that multiple pieces of knowledge extraction result data of labeled entities corresponding to the same entity category correspond to the same object, an alignment relationship is established for the multiple pieces of knowledge extraction result data of labeled entities, and knowledge fusion is performed.

3. The knowledge extraction method based on BERT algorithm according to claim 2 is characterized in that: Determining the entity category of the knowledge extraction result data according to the similarity between each piece of knowledge extraction result data and the existing entity category includes: Based on the first database, construct an entity category candidate set for each knowledge extraction result data of the labeled entity; the entity category candidate set includes multiple entity categories; Based on the second database, respectively obtaining original text data corresponding to a plurality of entity terms in the entity category candidate set; Calculating the similarity between the knowledge extraction result data of the annotated entity and the original text data corresponding to each entity term in the entity category candidate set respectively, and obtaining multiple similarities; the multiple similarities correspond to the multiple entity categories one by one; Determine multiple target similarities that are not less than a threshold value among the multiple similarities, and determine that the entity category corresponding to the target similarity with the largest value is the entity category of the knowledge extraction result data of the labeled entity.

4. The knowledge extraction method based on BERT algorithm according to claim 1 is characterized in that: The method also includes training the pre-trained BERT model by: Obtain a training data set of knowledge to be extracted; Preprocessing, word segmentation, annotation and format conversion are performed on the text data in the training data set to obtain a training vector data set of knowledge to be extracted; the training vector data set of knowledge to be extracted includes training vector data of the knowledge to be extracted of multiple annotated entities; Fine-tune the pre-trained BERT model based on the vector data set of the knowledge to be extracted; Wherein, the fine-tuning training includes: Obtaining a coding sequence of training vector data of knowledge to be extracted of multiple labeled entities; For each coding sequence: Randomly sample annotated entities in the same batch of encoded sequences; Extracting the encoding vectors corresponding to the first and last part of the encoding sequence according to the labeled entities obtained by random sampling; The encoding sequence is normalized based on the encoding vector.

5. The knowledge extraction method based on BERT algorithm according to claim 4 is characterized in that: The fine-tuning training also includes: Using a validation data set to test the trained pre-trained BERT model at preset intervals, and calculating evaluation indicators of the trained pre-trained BERT model to obtain evaluation indicators of multiple tests; the evaluation indicators include at least one of accuracy, recall and F1 value; In response to determining that the difference between the evaluation indicators of the multiple tests does not exceed a preset threshold, it is determined that the pre-trained BERT model is obtained.

6. The knowledge extraction method based on BERT algorithm according to claim 1 is characterized in that: The preprocessing of the text data of the knowledge to be extracted includes: removing HTML tags, special characters and repeated contents in the text data of the knowledge to be extracted; The annotation processing includes: annotating at least one of entities, relations and events; The format conversion process includes converting words in the annotated text data into word vectors, and performing at least one of a padding process and a truncation process so that the text data of the knowledge to be extracted has a preset length; Before performing clustering-based knowledge fusion on the knowledge extraction result data, the method further includes: The knowledge extraction result data is deduplicated and corrected.

7. The knowledge extraction method based on BERT algorithm according to claim 3 is characterized in that: The step of determining the entity category of the knowledge extraction result data of the annotated entity according to the similarity between each piece of knowledge extraction result data of the annotated entity and the existing entity category further comprises: In response to determining that the multiple similarities are all less than a threshold, determining that the entity category of the knowledge extraction result data of the labeled entity is an unknown entity category, and adding a new entity category in the first database; Based on the knowledge extraction result data of the plurality of labeled entities of the unknown entity category, calculating the pairwise similarity of the knowledge extraction result data of the plurality of labeled entities of each unknown entity category; In response to determining that the pairwise similarities are greater than a similarity threshold, the knowledge extraction result data of the corresponding two labeled entities are merged into the same newly added entity category.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.

10. A computer program product, comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Multi-modal data extraction method and system based on knowledge graph

    CN120632163A

  • Courseware script processing method and device, electronic equipment and computer storage medium

    CN120764490A

  • Generative knowledge object extraction method and system based on context learning

    CN120930750A