Information extraction method and device based on bidirectional graph convolutional network

Through the information extraction method based on bidirectional graph convolutional network, by encoding text words and screening relationship weighted graphs, the relationship confusion problem in entity recognition and relationship extraction is solved, and the accuracy of information extraction is improved.

CN120706528APending Publication Date: 2025-09-26709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510736021.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When jointly performing entity recognition and relationship extraction, the existing technology has the problem of confusing the relationships between entities, resulting in low information extraction accuracy.

Method used

An information extraction method based on a bidirectional graph convolutional network is adopted. By encoding each word in the text, the relevant probability of word combinations is predicted, a relationship weighted graph is constructed, and strongly associated word combinations are screened out. The bidirectional graph convolutional network is then used to extract neighborhood features and obtain the relationship labels of the word combinations.

Benefits of technology

It improves the accuracy of information extraction, avoids confusion of word combination relationships, and ensures the accuracy of entity recognition and relationship extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706528A_ABST
    Figure CN120706528A_ABST
Patent Text Reader

Abstract

The invention provides an information extraction method and device based on a bidirectional graph convolutional network, and the method comprises the steps: coding each word in a text, taking two different words in the same sentence as word combinations, predicting the correlation probability between the two words in each word combination, and obtaining the correlation probability of the two words in each word combination; meanwhile, entity recognition of words and relation extraction of word combinations are carried out; the probability of each word relative to each entity label is obtained, so that the entity label corresponding to each word is judged, and entity recognition of the words is completed; obtaining a relation weighted graph through the word combinations and the correlation probabilities, extracting neighborhood features of each word according to the relation weighted graph, screening out strong correlation word combinations according to the relation weighted graph, obtaining the probability of each strong correlation word combination relative to each relation label according to the neighborhood features, and completing relation extraction of the word combinations; by fully extracting the ontology features and the neighborhood features of the words, confusion of word combination relationships is avoided, and the accuracy of information extraction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to an information extraction method and device based on a bidirectional graph convolutional network. Background Art

[0002] Entity recognition and relationship extraction are crucial components of text information analysis. This task involves detecting entities with specific meanings within text and determining their boundaries and categories using appropriate methods. In practical applications, high accuracy is required in the model's output of entity and relationship predictions. However, in traditional pipeline models, entity recognition is typically considered a sequence labeling problem solved by conditional random fields. Classification models such as support vector machines are used to examine each entity pair to determine whether they have a task-specific relationship. This approach can further propagate errors in entity recognition to lower-level subtasks, leading to the proposal of joint entity-relationship extraction.

[0003] The main approach to joint entity-relationship extraction is parameter sharing. This method allows a single model to simultaneously perform both entity extraction and relationship classification. The relationship classification model shares the attributes and parameters of the entity extraction model and captures many semantic and grammatical properties of both entities and text. While this joint approach can simultaneously perform entity recognition and relationship extraction from textual information using a single model, since a sentence often contains multiple entities, relationships between these entities may become confused, leading to inaccurate relationship extraction.

[0004] In view of this, overcoming the defects of the prior art is an urgent problem to be solved in this technical field. Summary of the Invention

[0005] The technical problem to be solved by the present invention is how to improve the accuracy of information extraction and avoid confusion between the relationships between entities when jointly performing entity recognition and relationship extraction.

[0006] The present invention adopts the following technical solutions: In a first aspect, a method for extracting information based on a bidirectional graph convolutional network is provided, comprising: Encode each word in the text to obtain a word code, and take two different words in the same sentence as a word combination, and predict the correlation probability between the two words in each word combination; Input all word encodings into the entity extraction model to obtain the probability of each word relative to each entity label to determine the corresponding entity label for each word; A relationship weighted graph is obtained based on the word combinations and the relevant probabilities, domain features of each word are extracted based on the relationship weighted graph, strongly associated word combinations with correlation are screened out based on the relationship weighted graph, and the probability of each strongly associated word combination relative to the corresponding relationship label is obtained based on the domain features to determine the corresponding relationship label for each strongly associated word combination.

[0007] Preferably, encoding each word in the text to obtain a word code specifically includes: Train each word and obtain the corresponding part-of-speech embedding pos for each word, concatenate the part-of-speech embedding pos and the trained word to obtain the preprocessed word; Obtaining a preceding word of the word in its corresponding sentence, and encoding the preprocessed word according to the preceding word to obtain a forward code; Obtaining a subsequent word of the word in its corresponding sentence, and encoding the preprocessed word according to the subsequent word to obtain a backward code; The forward encoding and the backward encoding are combined to obtain the word encoding of the corresponding word.

[0008] Preferably, obtaining a weighted relationship graph according to the word combination and the related probability specifically includes: The words are taken as nodes, and the nodes of the words that are the same word combination are connected to each other, and the connection positions are represented by the related probabilities to obtain a relationship weighted graph.

[0009] Preferably, extracting the domain features of each word according to the weighted relationship graph and screening out strongly associated word combinations with correlation according to the weighted relationship graph specifically includes: The relationship weighted graph is input into the bidirectional graph convolutional network to extract the neighborhood features of each node; All word combinations whose correlation probabilities exceed the preset probability are regarded as strongly associated word combinations.

[0010] Preferably, the step of inputting the weighted relationship graph into a bidirectional graph convolutional network to extract neighborhood features of each node specifically includes: The expression of the neighborhood characteristics of each node is: ; in, is the neighborhood feature of the l+1 node, is the bias adjustment value, ReLU(*) is the activation function, u is the relationship between adjacent entities, is the weight matrix and v is the entity.

[0011] Preferably, obtaining the probability of each strongly associated word combination relative to each relationship label according to the neighborhood features to determine the relationship label corresponding to each strongly associated word combination specifically includes: Calculate the score of the strongly associated word combination relative to each entity label based on the degree of association between the strongly associated word combination and each relationship label; According to the score of the strongly associated word combination relative to each relationship label, the probability of the strongly associated word combination relative to each relationship label is obtained; The relationship tag with the highest probability is used as the relationship tag corresponding to the strongly associated word combination.

[0012] Preferably, the step of obtaining the probability of each strongly associated word combination relative to each relationship label according to the neighborhood features further includes: Obtain the entity label with the highest score for each word from the entity recognition model, and take the entity labels with the highest scores for all words in each sentence as the label sequence for the sentence; According to the neighborhood features of two words in the strongly associated words and the label sequence of the corresponding sentence, the probability of the strongly associated words relative to each relationship label is obtained.

[0013] Preferably, obtaining the probability of each word relative to each entity tag to determine the entity tag corresponding to each word specifically includes: Based on the degree of association between the word and each entity tag, the score of the word relative to each entity tag is calculated; The probability of a word relative to each entity label is obtained based on the score of the word relative to each entity label; The entity tag with the highest probability is used as the entity tag corresponding to the word. In a second aspect, an information extraction device based on a bidirectional graph convolutional network is provided, comprising at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to perform the information extraction method based on the bidirectional graph convolutional network.

[0014] In a third aspect, the present invention further provides a non-volatile computer storage medium, wherein the computer storage medium stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors to complete the method described in the first aspect.

[0015] In a fourth aspect, a chip is provided, comprising: a processor and an interface, for calling and running a computer program stored in a memory to execute the method of the first aspect.

[0016] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computer or a processor, causes the computer or the processor to execute the method of the first aspect.

[0017] In the sixth aspect, an information extraction system based on a bidirectional graph convolutional network is provided, comprising an information extraction device based on a bidirectional graph convolutional network as in the second aspect, and using an information extraction method based on a bidirectional graph convolutional network as in the first aspect.

[0018] The present invention provides an information extraction method and device based on a bidirectional graph convolutional network. The method encodes each word in a text, takes two different words in the same sentence as a word combination, predicts the correlation probability between the two words in each word combination, and simultaneously performs word entity recognition and word combination relationship extraction. By obtaining the probability of each word relative to each entity label, the corresponding entity label of each word is determined, and word entity recognition is completed. A relationship weighted graph is obtained through word combinations and correlation probabilities, and neighborhood features of each word are extracted according to the relationship weighted graph. Strongly associated word combinations with correlation are screened out according to the relationship weighted graph, and the probability of each strongly associated word combination relative to each relationship label is obtained according to the neighborhood features, and word combination relationship extraction is completed. By fully extracting the ontological features and neighborhood features of the words, confusion of the relationships of the word combinations is avoided, and the accuracy of information extraction is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0020] Figure 1 This is a flow chart of a method for information extraction based on a bidirectional graph convolutional network provided by an embodiment of the present invention; Figure 2 This is a flow chart of a method for word encoding in an information extraction method based on a bidirectional graph convolutional network provided by an embodiment of the present invention; Figure 3 A method for obtaining the score and probability of each word relative to an entity label in an information extraction method based on a bidirectional graph convolutional network provided by an embodiment of the present invention; Figure 4 It is a relational weighted graph in an information extraction method based on a bidirectional graph convolutional network provided by an embodiment of the present invention; Figure 5A method for obtaining scores and probabilities of word combinations relative to relationship labels in an information extraction method based on a bidirectional graph convolutional network provided by an embodiment of the present invention; Figure 6 Schematic diagram of the structure of an information extraction device based on a bidirectional graph convolutional network provided by an embodiment of the present invention; Figure 7 It is a structural diagram of another information extraction device based on a bidirectional graph convolutional network provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0022] In the description of the present invention, the terms "inside", "outside", "longitudinal", "lateral", "upper", "lower", "top", "bottom", etc. indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and do not require that the present invention must be constructed and operated in a specific orientation. Therefore, they should not be understood as limitations on the present invention.

[0023] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0024] Embodiment 1: Embodiment 1 of the present invention provides an information extraction method based on a bidirectional graph convolutional network, such as Figure 1 As shown, the method flow includes: In step 101, each word in the text is encoded to obtain a word code, and two different words in the same sentence are taken as a word combination, and the correlation probability between the two words in each word combination is predicted.

[0025] In this embodiment, the information extraction method based on the bidirectional graph convolutional network is used to classify the entity category of each word in the text, that is, entity recognition; and classify the relationship between each entity, that is, relationship extraction; thereby achieving effective acquisition of information in the text.

[0026] By encoding each word in the text and taking the context of the word into full consideration during the encoding process, the word is encoded according to the context of the word, so that the word encoding can carry both the meaning of the word itself and the meaning of its context, so that the word can be more accurately labeled and information extracted during subsequent entity recognition.

[0027] The word combination includes any two words in the same sentence. By identifying and extracting the relationship between different word combinations, the relationship between the two words in the word combination is obtained; the related probability is: the probability that there is a relationship between the two words in a word combination. This is because there are multiple words in a sentence, and these words may be related or unrelated to each other. The purpose of predicting the related probability for each word combination here is to obtain other words related to each word in a sentence, so that when the relationship type of the word combination is judged later, the neighborhood characteristics of each word can be fully considered, so that the result of the relationship extraction can be more accurate and avoid relationship confusion; and it can also ensure that only the word combinations with association are subsequently extracted, while those word combinations without association are not extracted, thereby avoiding unnecessary calculation processes.

[0028] Entity recognition and relationship extraction are performed separately, where: The entity recognition includes: in step 102, all word codes are input into the entity extraction model, and the probability of each word relative to the corresponding entity label is obtained to determine the entity label corresponding to each word.

[0029] The entity tag refers to the type to which a word can be attributed, that is, multiple entity tags are set, each entity tag represents a feature category of an object, and each word in the text is attributed to a corresponding entity tag based on the meaning of the word itself and the context in which the word is located. Each word ultimately needs to correspond to an entity tag. For example, in this embodiment, the entity tags may include: object names, location names, and personal names. For example, "cup" can be classified as an entity tag for an object name, and "second floor" can be classified as an entity tag for a location name.

[0030] The "probability of each word relative to each entity label" refers to the probability of a word belonging to each entity label, where the entity label with the highest probability is the entity label corresponding to the word. The entity extraction model can be a conditional random field (CRF) model.

[0031] The relationship extraction includes: in step 103, obtaining a relationship weighted graph based on the word combinations and the relevant probabilities, extracting the domain features of each word based on the relationship weighted graph, screening out strongly associated word combinations with correlation based on the relationship weighted graph, and obtaining the probability of each strongly associated word combination relative to the corresponding relationship label based on the domain features to determine the corresponding relationship label for each strongly associated word combination.

[0032] The relationship tag refers to the type of relationship between words that can be attributed. Each relationship tag represents a relationship feature category. According to the meaning of the two words in the word combination and the association relationship between the two words, the word combination in the text is attributed to the corresponding relationship tag.

[0033] The relationship weighted graph is: each word that may be related is connected to each other as a node, and the connection between each node (i.e., word) in the relationship weighted graph is represented by the relevant probability, so as to more intuitively show the relationship status between each word, quickly obtain each word and other words related to it, and thus obtain the neighborhood features of each word, ensuring the accuracy of subsequent relationship extraction.

[0034] On the other hand, since the relationship weighted graph carries the characteristics of the correlation probability between each word, it is possible to determine which words are related to each other and which words are not related to each other. The strongly related word combination is a word combination composed of two related words. Subsequent relationship extraction is only performed on strongly related words, thereby avoiding relationship extraction on unrelated word combinations and avoiding the problem of relationship confusion between multiple words.

[0035] For example, a sentence includes "tiger", "ant" and "insect", among which the word combination of "ant" and "insect" is classified as "belonging to" in the relationship category, i.e. "ants belong to insects". However, if relationship extraction is performed on all word combinations, the word combination of "tiger" and "insect" may also be classified as "belonging to" in the relationship category, i.e. "tigers belong to insects". This is obviously problematic. Therefore, by selecting strongly associated word combinations through the relationship weighted graph and calculating the relationship labels, the problem of relationship category confusion in word combinations can be better avoided. At the same time, the neighborhood features of each word are introduced in the calculation process to ensure the accuracy of relationship extraction.

[0036] In this embodiment, when encoding words in a text, it is necessary to encode them in combination with the context of the words, and at the same time, it is necessary to introduce corresponding word embeddings to ensure that the resulting word encoding can more accurately carry the meaning of the word itself. Therefore, this embodiment also involves the following designs: The encoding of each word in the text is performed to obtain a word encoding, such as Figure 2 As shown, the method flow includes: In step 201, each word is trained, and the part-of-speech embedding pos corresponding to each word is obtained, and the part-of-speech embedding pos and the trained word are concatenated to obtain the preprocessed word.

[0037] In this embodiment, the XLNet training model can be used as a word vector encoder to train words using the XLNet training model, where: For each sentence in the text D, represented by s, there is , that is, the text D includes n sentences, and the words in each sentence are Indicates that ,Right now Represents the jth word in the i-th sentence. Each word in each sentence in the text D is trained through the XLNet training model to obtain the representation of each word: .

[0038] On the other hand, in order to enrich the text features, each word The corresponding part of speech is embedded into POS and concatenated. The concatenated word is represented as follows: , the corresponding sentence It can be represented by a sequence of all the words in it: ,Right now There are m words in total.

[0039] In step 202, the preceding words of the word in its corresponding sentence are obtained, and the preprocessed word is encoded according to the preceding words to obtain a forward code.

[0040] In this embodiment, the preceding word is: all other words before the word in the corresponding sentence, such as The forward encoding needs to be based on arrive All word information to Encoding is performed. In this embodiment, the forward encoding is performed using The forward coding Need to be based on the hidden vector , the previous unit vector and To calculate, as follows: .

[0041] In step 203, the following words of the word in its corresponding sentence are obtained, and the preprocessed word is encoded according to the following words to obtain a backward code.

[0042] In this embodiment, the following words are: all other words after the word in the corresponding sentence, such as The backward encoding needs to be based on arrive All word information to Encoding is performed. In this embodiment, the backward encoding is performed using The backward coding Need to be based on the hidden vector , the previous unit vector and To calculate, as follows: .

[0043] In step 204, the forward encoding and the backward encoding are combined to obtain the word encoding of the corresponding word.

[0044] The combined word encoding is ,in, , That is, the word encoding with word features and context features extracted by the BiLSTM layer.

[0045] It should be noted that after encoding, all word encodings need to be processed by the softmax function. The softmax function is mainly used to convert each vector in the received word encoding into a positive probability distribution. Given a coding vector, the Softmax function will convert it into a vector whose all elements are in the range of (0, 1) and the sum is 1, so it can be interpreted as a probability distribution.

[0046] In this embodiment, the probability of each word relative to each entity tag is essentially obtained by scoring the word relative to each entity tag, and then obtaining the corresponding probability. For each word, the entity tag with a higher score represents a greater probability that the word belongs to the entity tag, and the entity tag with a lower score represents a smaller probability that the word belongs to the entity tag. Therefore, this embodiment also involves the following designs, such as Figure 3 As shown, the method flow includes: In step 301, the score of the word relative to each entity tag is calculated based on the degree of association between the word and each entity tag.

[0047] In this embodiment, a CRF model can be used to complete the scoring of words relative to various entity tags. In the CRF model, B-PER is assigned as the start marker of a word and E-PER is assigned as the end marker of a word. The score of a word relative to each entity tag is represented in the form of a matrix vector as a score vector sequence. The encoding of the entity tag is represented in the form of a matrix vector. The score vector sequence is And the label prediction vector is , correspondingly, the sequence of scores of words corresponding to each entity label, that is, the score of the linear chain conditional random field is .

[0048] In step 302, the probability of the word relative to each entity tag is obtained based on the score of the word relative to each entity tag, where the entity tag with the highest probability is used as the entity tag corresponding to the word.

[0049] Based on the score of each word relative to each entity label, the probability of the word relative to each entity label is obtained. The expression is as follows: ; Where Pr(*) is the probability.

[0050] In this embodiment, for the relationship extraction part, it is necessary to obtain a relationship weighted graph based on the relevant probabilities corresponding to each word combination, such as Figure 4 As shown, words are taken as nodes, the nodes of words that are the same word combination are connected to each other, and the connection positions are represented by the relevant probabilities.

[0051] It should be noted that, since only words in a sentence can be combined with each other, a sentence can be represented by a relational weighted graph, in which each word is connected to each other as a node, and the connection line between the nodes is represented by the related probability. Figure 4 For example, node 1 and node 2 are connected, and the correlation probability on the connecting line is 0.7, which means that the correlation probability between node 1 and node 2 as a word combination is 0.7, that is, the probability of association between node 1 and node 2 is 0.7.

[0052] In this embodiment, the method for obtaining the relevant probability is as follows: The probability prediction of word combinations is performed simultaneously when BiLSTM extracts and encodes the word features. For each word combination relationship r, the weight matrix needs to be learned. 、 and , and the correlation score , h w1 is the vector of word 1, h w2 is the vector of word 2, where RELU(*) is the activation function, and the relevant probability of the word combination is obtained according to the relevant score.

[0053] The method extracts the domain features of each word according to the relationship weighted graph, and screens out strongly associated word combinations with correlation according to the relationship weighted graph, specifically including: inputting the relationship weighted graph into a bidirectional graph convolutional network to extract the neighborhood features of each node; and treating all word combinations whose correlation probabilities exceed a preset probability as strongly associated word combinations.

[0054] The preset probability is set by those skilled in the art according to actual conditions. In this embodiment, the preset probability may be 50%.

[0055] The probability of each strongly associated word combination relative to each relationship label is obtained according to the neighborhood features, so as to determine the corresponding relationship label of each strongly associated word combination, such as Figure 5 As shown, the method flow includes: In step 401, a score of the strongly associated word combination relative to each entity tag is calculated based on the degree of association between the strongly associated word combination and each relationship tag.

[0056] It should be noted that in this embodiment, since only the relevant probabilities of word combinations are predicted when encoding words, and the relationship types of each word combination are not determined, it is necessary to input the relationship weighted graph into the Bi-GCN module, and extract the neighborhood features of each node according to the relationship weighted graph through the bidirectional graph convolutional network, thereby aggregating the comprehensive features of each word, and fully considering the influence of different relationship types on each other. All strongly associated word combinations in the relationship weighted graph are obtained through the Bi-GCN module, and the scores of each strongly associated word combination relative to each entity label are preliminarily calculated. The scoring is mainly based on the ontological features and neighborhood features of the two words in the strongly associated word combination, thereby improving the accuracy of the scoring.

[0057] The Bi-GCN module sends the selected strongly associated word combinations and their corresponding scores to the softmax module for optimization. At the same time, the softmax module also obtains label embeddings from the entity recognition model. The label embeddings are a label sequence composed of the entity labels with the highest corresponding scores for all words in each sentence. The softmax module optimizes the scores of all strongly associated words based on the label embeddings, strongly associated word combinations, and their corresponding scores, thereby obtaining the optimized score of each strongly associated word relative to all relationship labels.

[0058] In this embodiment, the features obtained by the Bi-GCN module for each word can be expressed as: ,in , ; b is the bias adjustment value, N(u) is the convolution model, and v is the entity.

[0059] The step of inputting the relationship weighted graph into the bidirectional graph convolutional network to extract the neighborhood features of each node specifically includes: The expression of the neighborhood characteristics of each node is: ; in, is the neighborhood feature of the l+1 node, is the bias adjustment value, ReLU(*) is the activation function, u is the relationship between adjacent entities, is the weight matrix and v is the entity.

[0060] In step 402, the probability of the strongly associated word combination relative to each relationship tag is obtained based on the score of the strongly associated word combination relative to each relationship tag; the relationship tag with the highest probability is used as the relationship tag corresponding to the strongly associated word combination.

[0061] The model proposed in this method uses two kinds of losses in training: entity loss and relation loss, both of which belong to classification loss. , the model uses the conventional (Begin, Inside, End, Single, Out) tagging method to mark the true labels, each word belongs to one of the five categories, and cross entropy is used as the classification loss function during model training.

[0062] For relationship loss , the model predicts relations based on word pairs, so the relationship types between word pairs should also be based on word pairs. and second round of relationship losses The basic relationship vectors of are the same, so the cross entropy loss function is used during model training.

[0063] for and , where NER is named entity recognition, RE is relation recognition, c is the normality parameter, eloss is the entity loss, and rloss is the relation loss. During the model training process, additional weights are added to those in-class entity or relation terms. , the total loss is calculated as the sum of entity loss and two-round relationship loss, that is ,in is the total loss, For the first round of entity losses, For the first round of relationship loss, is the second-round relation loss. This method minimizes the loss and trains the proposed model in an end-to-end manner.

[0064] Example 2: This embodiment 2 provides an information extraction device based on a bidirectional graph convolutional network on the basis of embodiment 1, including an input module, an XLNet module, a BiLSTM module, a first softmax module, a CRF module, a Bi-GCN module, a second softmax module and an output module, wherein: The input module is used to input the corresponding text to the XLNet module.

[0065] The XLNet module is used to train and extract features from words in the text, and perform corresponding word embedding for each word.

[0066] The BiLSTM (full name: Bi - Long Short-Term Memory) module is a bidirectional LSTM model used to extract word features and encode the preceding and following contexts of each word separately. Figure 6 The upward arrows of the BiLSTM module represent the forward code and backward code corresponding to each word, and then the forward code and backward code are merged to obtain the word code, and the word code is sent to the first softmax module. At the same time, the BiLSTM module also needs to predict the relevant probability of each word combination for use in subsequent relationship extraction.

[0067] The first softmax module is mainly used to convert each vector in the received word encoding into a positive probability distribution. Given a coding vector, the Softmax function converts it into a vector whose elements are in the range of (0, 1) and sum to 1, so it can be interpreted as a probability distribution. The first softmax module sends the processed word encoding and the related probabilities of each word combination to the CRF module and Bi-GCN module respectively.

[0068] The CRF module is used to complete the scoring of words relative to various entity tags based on word encoding, and obtain the probability of words relative to each entity tag based on the score of each word relative to each entity tag, thereby completing entity recognition of words in the text.

[0069] The Bi-GCN module is used to input a relational weighted graph consisting of word combinations and related probabilities, and screen out strongly correlated word combinations based on the relational weighted graph, calculate the scores of the strongly correlated word combinations relative to each relationship label, and send the scores to the second softmax module.

[0070] The second softmax module is used to receive the scores of strongly correlated word combinations relative to each relationship label, and the label embeddings from the CRF module. The label embeddings are a label sequence composed of the entity labels with the highest corresponding scores for all words in each sentence, thereby optimizing the scores of strongly correlated word combinations relative to each relationship label to obtain relatively more accurate scores, and calculating the probability of strongly correlated word combinations relative to each relationship label, thereby completing the relationship extraction of word combinations in the text.

[0071] Example 3: like Figure 7, which is a schematic diagram of the structure of an information extraction device based on a bidirectional graph convolutional network according to an embodiment of the present invention. The information extraction device based on a bidirectional graph convolutional network according to this embodiment includes one or more processors 71 and a memory 72.

[0072] The processor 71 and the memory 72 may be connected via a bus or other means. Figure 7 The bus connection is taken as an example.

[0073] Memory 72, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the information extraction method based on a bidirectional graph convolutional network in the above embodiment. Processor 71 executes the information extraction method based on a bidirectional graph convolutional network by running the non-volatile software program and instructions stored in memory 72.

[0074] The memory 72 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 72 may optionally include a memory remotely located relative to the processor 71, and such remote memory may be connected to the processor 71 via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0075] The program instructions / modules are stored in the memory 72, and when executed by the one or more processors 71, the information extraction method based on the bidirectional graph convolutional network in the above embodiment is executed, for example, the above described Figures 1 to 6 The steps shown.

[0076] An embodiment of the present invention further provides a computer storage medium having computer program instructions stored thereon; when the computer program instructions are executed by a processor, the multi-vessel tracking method based on multimodal information fusion provided by an embodiment of the present invention is implemented.

[0077] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An information extraction method based on a bidirectional graph convolutional network, characterized in that: include: Encode each word in the text to obtain a word code, and take two different words in the same sentence as a word combination, and predict the correlation probability between the two words in each word combination; Input all word encodings into the entity extraction model to obtain the probability of each word relative to each entity label to determine the corresponding entity label for each word; A relationship weighted graph is obtained based on the word combinations and the relevant probabilities, domain features of each word are extracted based on the relationship weighted graph, strongly associated word combinations with correlation are screened out based on the relationship weighted graph, and the probability of each strongly associated word combination relative to the corresponding relationship label is obtained based on the domain features to determine the corresponding relationship label for each strongly associated word combination.

2. The information extraction method based on a bidirectional graph convolutional network according to claim 1, characterized in that The encoding of each word in the text to obtain a word code specifically includes: Train each word and obtain the corresponding part-of-speech embedding pos for each word, concatenate the part-of-speech embedding pos and the trained word to obtain the preprocessed word; Obtaining a preceding word of the word in its corresponding sentence, and encoding the preprocessed word according to the preceding word to obtain a forward code; Obtaining a subsequent word of the word in its corresponding sentence, and encoding the preprocessed word according to the subsequent word to obtain a backward code; The forward encoding and the backward encoding are combined to obtain the word encoding of the corresponding word.

3. The information extraction method based on a bidirectional graph convolutional network according to claim 1, characterized in that The step of obtaining a weighted relationship graph based on the word combination and the related probabilities specifically includes: The words are taken as nodes, and the nodes of the words that are the same word combination are connected to each other, and the connection positions are represented by the related probabilities to obtain a relationship weighted graph.

4. The information extraction method based on a bidirectional graph convolutional network according to claim 1, characterized in that The extracting of domain features of each word according to the weighted relationship graph and screening out strongly associated word combinations having correlation according to the weighted relationship graph specifically includes: The relationship weighted graph is input into the bidirectional graph convolutional network to extract the neighborhood features of each node; All word combinations whose correlation probabilities exceed the preset probability are regarded as strongly associated word combinations.

5. The information extraction method based on a bidirectional graph convolutional network according to claim 4, characterized in that The step of inputting the relationship weighted graph into the bidirectional graph convolutional network to extract the neighborhood features of each node specifically includes: The expression of the neighborhood characteristics of each node is: ; in, is the neighborhood feature of the l+1 node, is the bias adjustment value, ReLU(*) is the activation function, u is the relationship between adjacent entities, is the weight matrix and v is the entity.

6. The information extraction method based on a bidirectional graph convolutional network according to claim 1, characterized in that Obtaining the probability of each strongly associated word combination relative to each relationship label according to the neighborhood features to determine the relationship label corresponding to each strongly associated word combination, specifically including: Calculate the score of the strongly associated word combination relative to each entity label based on the degree of association between the strongly associated word combination and each relationship label; According to the score of the strongly associated word combination relative to each relationship label, the probability of the strongly associated word combination relative to each relationship label is obtained; The relationship tag with the highest probability is used as the relationship tag corresponding to the strongly associated word combination.

7. The information extraction method based on a bidirectional graph convolutional network according to claim 6, characterized in that The step of obtaining the probability of each strongly associated word combination relative to each relationship label according to the neighborhood features further includes: Obtain the entity label with the highest score for each word from the entity recognition model, and take the entity labels with the highest scores for all words in each sentence as the label sequence for the sentence; According to the neighborhood features of two words in the strongly associated words and the label sequence of the corresponding sentence, the probability of the strongly associated words relative to each relationship label is obtained.

8. The information extraction method based on a bidirectional graph convolutional network according to claim 1, characterized in that Obtaining the probability of each word relative to each entity tag to determine the entity tag corresponding to each word specifically includes: Based on the degree of association between the word and each entity tag, the score of the word relative to each entity tag is calculated; The probability of a word relative to each entity label is obtained based on the score of the word relative to each entity label; The entity tag with the highest probability is used as the entity tag corresponding to the word.

9. An information extraction device based on a bidirectional graph convolutional network, characterized in that: It includes at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to execute the information extraction method based on a bidirectional graph convolutional network according to any one of claims 1 to 8.

10. A non-volatile computer storage medium, characterized in that The non-volatile computer storage medium stores computer program instructions, which, when executed by one or more processors, are used to complete the information extraction method based on a bidirectional graph convolutional network as described in any one of claims 1 to 8.