Entity Information Processing Method, Apparatus, Computer Device, and Storage Medium
By constructing a vocabulary co-occurrence graph and using the target graph convolution neural network to generate the target word vector of vocabulary, the problem of unsatisfactory entity uniformity in traditional technology is solved, and the accurate judgment of the same entity vocabulary is achieved.
Patent Information
- Application Number
- CN202111380860.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-20
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-11-20
AI Technical Summary
In traditional technology, the effect of entity unification is not ideal, and it is impossible to accurately determine whether different vocabulary belongs to the same entity, mainly because there is too little labeling data.
By constructing a vocabulary co-occurrence graph, a target word vector of vocabulary is generated using the target graph convolution neural network, and the vocabulary belonging to the same entity is determined based on the similarity of these vectors.
It realizes accurate judgment of the vocabulary of the same entity, improves the effect of entity unity, and avoids the difficulty of identification caused by too little labeling data.
Smart Images

Figure CN114091445B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and particularly to a method, apparatus, computer device, storage medium, and computer program product for entity information processing. Background Art
[0002] In recent years, artificial intelligence technology has developed rapidly and has very broad application prospects in fields such as education, medical care, agriculture, and transportation.
[0003] Artificial intelligence technology can also be applied to the field of natural language processing, and can use artificial intelligence technology to analyze, understand, and process natural language, enabling effective communication between humans and computers in natural language. For example, in real life, for the same objective entity, there are often different names or expressions in different scenarios. For example, "Industrial and Commercial Bank of China", "ICBC", "Industrial and Commercial Bank", and "ICBC" all refer to the same objective entity, Industrial and Commercial Bank of China. Such diverse expressions often cause ambiguity in the computer's understanding of natural language. In traditional technologies, the machine is often trained by relying on the method of manual data annotation. However, the annotated data is often very small, resulting in the problem that it is impossible to accurately determine whether different words belong to the same entity.
[0004] Therefore, in traditional technologies, there is a problem that the entity unification effect is not ideal. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for entity information processing that can accurately determine words belonging to the same entity.
[0006] In a first aspect, the present application provides a method for entity information processing. The method includes:
[0007] Obtain a corpus to be processed;
[0008] Construct a word co-occurrence graph with each word in the corpus to be processed as a node and the co-occurrence relationship between any two words as the edge of the node; the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the words corresponding to the nodes connected by the edge in the corpus to be processed;
[0009] Input the word co-occurrence graph into a target graph convolutional neural network to obtain target word vectors corresponding to each word; the target word vectors are obtained by the target graph convolutional neural network encoding and processing the word co-occurrence graph;
[0010] Determine target words belonging to the same entity according to the target word vectors corresponding to each word; the similarity between the target word vectors corresponding to the target words is greater than a preset similarity threshold.
[0011] In one embodiment, the method further includes:
[0012] Construct a sample co-occurrence graph with each sample word in the sample corpus to be processed as a node and the co-occurrence relationship between any two of the sample words as the edge of the node; the edge has a corresponding weight; the weight is used to characterize the co-occurrence times of the sample words corresponding to the nodes connected by the edge in the sample corpus to be processed; the sample words corresponding to the nodes connected by the edge appear simultaneously in the same sentence in the sample corpus to be processed;
[0013] Input the sample co-occurrence graph into an initial graph convolutional neural network to obtain initial word vectors corresponding to each of the sample words; the initial word vectors are obtained by the initial graph convolutional neural network encoding the sample co-occurrence graph;
[0014] Train the initial graph convolutional neural network according to the initial word vectors corresponding to each of the sample words to obtain the target graph convolutional neural network.
[0015] In one embodiment, the training the initial graph convolutional neural network according to the initial word vectors corresponding to each of the sample words to obtain the target graph convolutional neural network includes:
[0016] Determine a word to be predicted among each of the sample words;
[0017] Determine the adjacent sample words of the word to be predicted; wherein, in the sample co-occurrence graph, the nodes corresponding to the adjacent sample words are connected to the node corresponding to the word to be predicted;
[0018] Obtain a prediction distribution result of the initial word vector of the word to be predicted according to the initial word vectors corresponding to the adjacent sample words;
[0019] Train the initial graph convolutional neural network according to the difference between the prediction distribution result of the initial word vector of the word to be predicted and the prior distribution result to obtain the target graph convolutional neural network.
[0020] In one embodiment, the training the initial graph convolutional neural network according to the difference between the prediction distribution result of the initial word vector of the word to be predicted and the prior distribution result to obtain the target graph convolutional neural network includes:
[0021] Obtain the prior distribution result of the initial word vector of the word to be predicted;
[0022] Perform cross-entropy calculation on the prediction distribution result and the prior distribution result to obtain a cross-entropy calculation result;
[0023] According to the cross - entropy calculation result, update the weight parameters of the initial graph convolutional neural network in the reverse gradient direction;
[0024] Retrain the graph convolutional neural network with the updated weight parameters until the trained graph convolutional neural network meets the preset training end condition, and obtain the target graph convolutional neural network.
[0025] In one embodiment, after the step of obtaining the corpus to be processed, the method further includes:
[0026] Perform word segmentation on the corpus to be processed to obtain the word - segmented corpus;
[0027] Perform redundancy removal on the word - segmented corpus to obtain the redundancy - removed corpus.
[0028] In one embodiment, the step of determining the target words belonging to the same entity according to the target word vectors corresponding to each of the words includes:
[0029] Determine the words to be matched among each of the words;
[0030] Determine the cosine similarity between the target word vector corresponding to the word to be matched and the target word vectors corresponding to the candidate words; the candidate words are the words in the corpus to be processed other than the word to be matched;
[0031] If the cosine similarity is greater than the preset cosine similarity threshold, determine that the word to be matched and the candidate words are the target words belonging to the same entity.
[0032] In a second aspect, the present application further provides an entity information processing device. The device includes:
[0033] A corpus acquisition module, configured to acquire a corpus to be processed;
[0034] A co - occurrence graph construction module for constructing a word co - occurrence graph with each word in the corpus to be processed as a node and the co - occurrence relationship between any two words as the edge of the node; the edge has a corresponding weight; the weight is used to represent the co - occurrence times of the words corresponding to the nodes connected to the edge in the corpus to be processed;
[0035] A target word vector determination module for inputting the word co - occurrence graph into a target graph convolutional neural network to obtain the target word vectors corresponding to each of the words; the target word vectors are obtained by the target graph convolutional neural network encoding the word co - occurrence graph;
[0036] A target vocabulary determination module, configured to determine target vocabularies belonging to the same entity according to the target word vectors corresponding to each of the vocabularies; the similarity between the target word vectors corresponding to the target vocabularies is greater than a preset similarity threshold.
[0037] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0038] Obtain a corpus to be processed;
[0039] Construct a vocabulary co-occurrence graph with each vocabulary in the corpus to be processed as a node and the co-occurrence relationship between any two vocabularies as the edge of the node; the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the vocabulary corresponding to the node connected to the edge in the corpus to be processed;
[0040] Input the vocabulary co-occurrence graph into a target graph convolutional neural network to obtain target word vectors corresponding to each of the vocabularies; the target word vectors are obtained by the target graph convolutional neural network encoding and processing the vocabulary co-occurrence graph;
[0041] Determine target vocabularies belonging to the same entity according to the target word vectors corresponding to each of the vocabularies; the similarity between the target word vectors corresponding to the target vocabularies is greater than a preset similarity threshold.
[0042] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0043] Obtain a corpus to be processed;
[0044] Construct a vocabulary co-occurrence graph with each vocabulary in the corpus to be processed as a node and the co-occurrence relationship between any two vocabularies as the edge of the node; the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the vocabulary corresponding to the node connected to the edge in the corpus to be processed;
[0045] Input the vocabulary co-occurrence graph into a target graph convolutional neural network to obtain target word vectors corresponding to each of the vocabularies; the target word vectors are obtained by the target graph convolutional neural network encoding and processing the vocabulary co-occurrence graph;
[0046] Determine target vocabularies belonging to the same entity according to the target word vectors corresponding to each of the vocabularies; the similarity between the target word vectors corresponding to the target vocabularies is greater than a preset similarity threshold.
[0047] Fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the following steps:
[0048] Obtain the corpus to be processed;
[0049] Construct a lexical co-occurrence graph with each word in the corpus to be processed as a node and the co-occurrence relationship between any two words as the edge of the node; the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the words corresponding to the nodes connected by the edge in the corpus to be processed;
[0050] Input the lexical co-occurrence graph into a target graph convolutional neural network to obtain target word vectors corresponding to each word; the target word vectors are obtained by the target graph convolutional neural network encoding the lexical co-occurrence graph;
[0051] Determine the target words belonging to the same entity according to the target word vectors corresponding to each word; the similarity between the target word vectors corresponding to the target words is greater than a preset similarity threshold.
[0052] The above entity information method, device, computer device, storage medium and computer program product, by obtaining the corpus to be processed; then, using each word in the corpus to be processed as a node and the co-occurrence relationship between any two words as the edge of the node, construct a lexical co-occurrence graph; wherein, the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the words corresponding to the nodes connected by the edge in the corpus to be processed; then, input the lexical co-occurrence graph into a target graph convolutional neural network to obtain target word vectors corresponding to each word; wherein, the target word vectors are obtained by the target graph convolutional neural network encoding the lexical co-occurrence graph; finally, determine the target words belonging to the same entity according to the target word vectors corresponding to each word; wherein, the similarity between the target word vectors corresponding to the target words is greater than a preset similarity threshold; thus, by using each word in the corpus to be processed as a node and connecting the nodes corresponding to the words with co-occurrence relationships by edges, construct a lexical co-occurrence graph; wherein, the weight corresponding to the edge represents the co-occurrence times of the words corresponding to the nodes connected by the edge in the corpus to be processed, so that the lexical co-occurrence graph not only reflects the co-occurrence relationship of the words corresponding to each node in the corpus to be processed, which can correspond to the local features of the words; but also reflects the degree of association of the words with co-occurrence relationships in the corpus to be processed, which can correspond to the global features of the words; then input the above lexical co-occurrence graph into the target graph convolutional neural network, extract the feature information of each node of the lexical co-occurrence graph, and obtain target word vectors corresponding to each word that integrate local feature information and global feature information; furthermore, according to the target word vectors with rich feature information corresponding to each word, the target words belonging to the same entity in the corpus to be processed can be accurately determined. Description of the Drawings
[0053] Figure 1 It is a schematic flowchart of a method for processing entity information in an embodiment;
[0054] Figure 2 It is a diagram of the node connection relationship in the sample co-occurrence graph in an embodiment;
[0055] Figure 3 It is a schematic flowchart of the steps for obtaining a target graph convolutional neural network in an embodiment;
[0056] Figure 4 It is a schematic flowchart of the steps for obtaining a target graph convolutional neural network in another embodiment;
[0057] Figure 5 It is a schematic flowchart of another method for processing entity information in another embodiment;
[0058] Figure 6 It is a structural block diagram of an entity information processing device in an embodiment;
[0059] Figure 7 It is the internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0060] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. It should be noted that the entity information methods, devices, computer devices, storage media and computer program products disclosed in the present application can be applied to the field of fintech, and can also be used in any field other than the field of fintech.
[0061] In one embodiment, as Figure 1 shown, a method for processing entity information is provided. In this embodiment, it is exemplified that the method is applied to a terminal. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0062] Step S110, obtain the corpus to be processed.
[0063] In specific implementation, the terminal can receive the corpus to be processed uploaded by other terminals, or can also perform data crawling on the website to obtain the corpus to be processed. By preprocessing the corpus to be processed, for example, performing word segmentation processing and / or redundancy removal processing on the corpus to be processed, each vocabulary in the corpus to be processed is obtained; among them, the corpus to be processed can be from the corpus within different industries, or can also be from journal articles, etc.
[0064] Step S120: Construct a lexical co-occurrence graph with each word in the corpus to be processed as a node and the co-occurrence relationship between any two words as the edge of the node.
[0065] Among them, the edge has a corresponding weight.
[0066] Among them, the weight is used to represent the number of co-occurrences of the words corresponding to the nodes connected by the edge in the corpus to be processed.
[0067] Among them, the words corresponding to the nodes connected by the edge appear simultaneously in the same sentence in the corpus to be processed.
[0068] In specific implementation, each word in the corpus to be processed will be used as a node. If any two words appear simultaneously in the same sentence in the corpus to be processed, it indicates that the above two words have a co-occurrence relationship, and the nodes corresponding to the above two words will be connected by an edge. Each edge has a corresponding weight, and the weight is used to represent the number of times the words corresponding to the nodes connected by the edge appear simultaneously in the same sentence in the corpus to be processed. For example, if there are three sentences in the corpus to be processed, and the words in the corpus to be processed are obtained after preprocessing the three sentences: Financial Street, Xicheng District, Beijing; Fuxingmen, Xicheng District, Beijing; Industrial and Commercial Bank of China, Fuxingmen, Beijing. Then, the number of times "Beijing" and "Xicheng District" appear simultaneously in the same sentence in the corpus to be processed is 2, so the weight of the connecting edge between the nodes corresponding to "Beijing" and "Xicheng District" is 2; the number of times "Beijing" and "Fuxingmen" appear simultaneously in the same sentence in the corpus to be processed is 2, so the weight of the connecting edge between the nodes corresponding to "Beijing" and "Fuxingmen" is 2. In this way, by connecting the nodes corresponding to the words with co-occurrence relationships with edges and assigning corresponding weights to each edge, the lexical co-occurrence graph of the corpus to be processed can be constructed.
[0069] Step S130: Input the lexical co-occurrence graph into the target graph convolutional neural network to obtain the target word vectors corresponding to each word.
[0070] Among them, the target word vector is obtained by the target graph convolutional neural network encoding the lexical co-occurrence graph.
[0071] In specific implementation, input the lexical co-occurrence graph into the trained target graph convolutional neural network. The target graph convolutional neural network can encode the lexical co-occurrence graph according to the edge connection relationship between each node in the lexical co-occurrence graph and the weight of the edge, that is, the co-occurrence relationship of the words corresponding to each node in the corpus to be processed and the degree of association of the words with co-occurrence relationships, extract the local features and global features of each node, and output the target word vectors corresponding to the words of each node in the lexical co-occurrence graph.
[0072] Step S140: Determine the target words belonging to the same entity according to the target word vectors corresponding to each word.
[0073] Among them, the similarity between the target word vectors corresponding to the target words is greater than a preset similarity threshold.
[0074] In a specific implementation, the terminal calculates the similarity between the target word vector corresponding to the word to be solved and the target word vector corresponding to the candidate word according to the target word vectors corresponding to each word output by the graph convolutional neural network; among them, the candidate word is a word other than the word to be solved in the corpus to be processed; if there is a target word vector corresponding to a candidate word whose similarity with the target word vector corresponding to the word to be solved is greater than the preset similarity threshold, it is determined that the candidate word and the word to be solved belong to the same entity.
[0075] In the above entity information processing method, the corpus to be processed is obtained; then, each word in the corpus to be processed is used as a node, and the co-occurrence relationship between any two of the words is used as the edge of the node to construct a word co-occurrence graph; among them, the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the words corresponding to the nodes connected to the edge in the corpus to be processed; then, the word co-occurrence graph is input into the target graph convolutional neural network to obtain the target word vectors corresponding to each word; among them, the target word vector is obtained by the target graph convolutional neural network encoding and processing the word co-occurrence graph; finally, according to the target word vectors corresponding to each word, the target words belonging to the same entity are determined; among them, the similarity between the target word vectors corresponding to the target words is greater than a preset similarity threshold; in this way, each word in the corpus to be processed is used as a node, and the nodes corresponding to the words with co-occurrence relationships are connected by edges to construct a word co-occurrence graph; among them, the weight corresponding to the edge represents the co-occurrence times of the words corresponding to the nodes connected by the edge in the corpus to be processed, so that the word co-occurrence graph not only reflects the co-occurrence relationship of the words corresponding to each node in the corpus to be processed, which can correspond to the local features of the words; but also reflects the degree of association of the words with co-occurrence relationships in the corpus to be processed, which can correspond to the global features of the words; then the above word co-occurrence graph is input into the target graph convolutional neural network, and the feature information of each node in the word co-occurrence graph is extracted to obtain the target word vectors corresponding to each word that integrate local feature information and global feature information; furthermore, according to the target word vectors corresponding to each word with rich feature information, the target words belonging to the same entity in the corpus to be processed can be accurately determined.
[0076] In one embodiment, the method further includes: constructing a sample co-occurrence graph with each sample word in the sample corpus to be processed as a node and the co-occurrence relationship between any two sample words as the edge of the node; inputting the sample co-occurrence graph into an initial graph convolutional neural network to obtain initial word vectors corresponding to each of the sample words; and training the initial graph convolutional neural network according to the initial word vectors corresponding to each sample word to obtain a target graph convolutional neural network.
[0077] Among them, the edge has a corresponding weight.
[0078] Among them, the weight is used to represent the co-occurrence times of the sample words corresponding to the nodes connected by the edge in the sample corpus to be processed.
[0079] Among them, the sample words corresponding to the nodes connected by the edge appear simultaneously in the same sentence of the sample corpus to be processed.
[0080] Among them, the initial word vector is obtained by the initial graph convolutional neural network encoding the sample co-occurrence graph.
[0081] Among them, the sample corpus to be processed can come from the corpora within different industries or from journal articles, etc.
[0082] In a specific implementation, the terminal preprocesses the obtained sample corpus to be processed, such as word segmentation processing and / or redundancy removal processing, to obtain each sample word in the sample corpus to be processed; then, according to each sample word in the sample corpus to be processed, constructs a sample co-occurrence graph, and the specific construction method and steps of the sample co-occurrence graph are the same as those in step S120, which will not be elaborated here. Then, input the sample co-occurrence graph into the initial graph convolutional neural network, and the initial graph convolutional neural network encodes the sample co-occurrence graph according to the edge connection relationship and the edge weight of each node in the sample co-occurrence graph to obtain the initial word vectors of each sample word; among them, the specific expression of the initial word vector is shown in the following formula:
[0083]
[0084] Among them, i is the node corresponding to the sample word to be solved, H i is the initial word vector corresponding to the sample word to be solved, j is the first-order adjacent node connected to i by an edge, A ij is the adjacency matrix of node i, and H j is the weight of the connection edge between node i and node j.
[0085] Among them, the connection relationship graph between node i and node j is as Figure 2 shown.
[0086] Then, based on the initial word vectors corresponding to each sample word, the initial graph convolutional neural network is trained until the initial graph convolutional neural network meets the training end condition, thereby obtaining the target graph convolutional neural network.
[0087] The technical solution of this embodiment inputs the sample co-occurrence graph constructed based on sample words into the initial graph convolutional neural network to obtain the initial word vectors corresponding to each sample word; then, based on the initial word vectors corresponding to each sample word, the initial graph convolutional neural network is trained, so that the trained target graph convolutional neural network can accurately generate the target word vectors corresponding to each word based on the word co-occurrence graph; thus, the target words belonging to the same entity can be accurately determined based on the target word vectors corresponding to each word.
[0088] In one embodiment, as Figure 3 shown, training the initial graph convolutional neural network based on the initial word vectors corresponding to each of the sample words to obtain a target graph convolutional neural network includes the following steps:
[0089] Step S310, determine the word to be predicted among each sample word.
[0090] In specific implementation, the terminal needs to determine the node to be predicted in the sample co-occurrence graph, that is, determine the word to be predicted among each sample word.
[0091] Step S320, determine the adjacent sample words of the word to be predicted.
[0092] Among them, in the sample co-occurrence graph, the nodes corresponding to the adjacent sample words are connected to the nodes corresponding to the word to be predicted.
[0093] In specific implementation, after the terminal determines the node to be predicted, it can determine the nodes directly connected to the node to be predicted as the adjacent nodes of the node to be predicted, that is, the adjacent sample words corresponding to the word to be predicted.
[0094] Step S330, obtain the prediction distribution result of the initial word vector of the word to be predicted according to the initial word vectors corresponding to the adjacent sample words.
[0095] In specific implementation, after the terminal determines the adjacent sample words of the word to be predicted, it inputs the initial word vectors corresponding to the adjacent sample words of the word to be predicted into the initial graph convolutional neural network. Through the iterative operation of the initial graph convolutional neural network, an output vector is obtained, and the output vector is subjected to a softmax (logistic regression) function operation to obtain the prediction distribution result of the initial word vector of the word to be predicted.
[0096] Step S340, training the initial graph convolutional neural network according to the difference between the predicted distribution result and the prior distribution result of the initial word vector of the vocabulary to be predicted, to obtain a target graph convolutional neural network.
[0097] In the specific implementation, the terminal needs to calculate the difference between the predicted distribution result of the initial word vector of the vocabulary to be predicted and the prior distribution result of the initial word vector of the vocabulary to be predicted, and update the parameters in the initial graph convolutional neural network according to the difference between the predicted distribution result and the prior distribution result, until the graph convolutional neural network after the parameter update meets the training end condition, so that the target graph convolutional neural network can be obtained.
[0098] The technical solution of this embodiment is to determine the vocabulary to be predicted in each sample vocabulary and determine the adjacent sample vocabulary of the vocabulary to be predicted; then, according to the initial word vector corresponding to the adjacent sample vocabulary, obtain the predicted distribution result of the initial word vector of the vocabulary to be predicted; finally, based on the difference between the predicted distribution result and the prior distribution result of the initial word vector of the vocabulary to be predicted, train the initial graph convolutional neural network to obtain the target graph convolutional neural network; in this way, the initial graph convolutional neural network can be trained based on the difference between the predicted distribution result and the prior distribution result of the initial word vector of the vocabulary to be predicted without introducing external information to obtain the target graph convolutional neural network; thus, there is no need to rely on manually labeled data, avoiding the problem of inaccurate entity unified processing due to too little labeled data.
[0099] In one embodiment, if Figure 4 As shown, step S340 includes:
[0100] Step S410, obtaining the prior distribution result of the initial word vector of the vocabulary to be predicted.
[0101] In the specific implementation, the terminal needs to obtain the prior distribution result of the initial word vector of the vocabulary to be predicted to evaluate the degree of difference between the predicted distribution result of the initial word vector of the vocabulary to be predicted output by the initial graph convolutional neural network and the prior distribution result.
[0102] Step S420, performing cross entropy calculation on the predicted distribution result and the prior distribution result to obtain a cross entropy calculation result.
[0103] In the specific implementation, the terminal performs a cross entropy loss function operation on the predicted distribution result of the initial word vector of the vocabulary to be predicted and the prior distribution result to obtain a cross entropy calculation result to measure the difference between the predicted distribution result of the initial word vector of the vocabulary to be predicted and the prior distribution result.
[0104] Among them, the formula of the cross entropy loss function is as follows:
[0105]
[0106] Among them, N is the dimension of the initial word vector of the word to be predicted, and y i is the prior distribution result of the initial word vector of the word to be predicted, and p i is the predicted distribution result of the initial word vector of the word to be predicted.
[0107] Step S430, according to the cross-entropy calculation result, update the weight parameters of the initial graph convolutional neural network by backpropagation gradient.
[0108] In specific implementation, the terminal updates the weight parameters in the initial graph convolutional neural network by backpropagation gradient based on the cross-entropy calculation result between the predicted distribution result and the prior distribution result of the initial word vector of the word to be predicted. The gradient descent method is a first-order optimization algorithm, usually also called the steepest descent method. To use the gradient descent method to find the local minimum of a function, the terminal must perform iterative search at a specified step size distance point in the opposite direction of the gradient or approximate gradient corresponding to the current point on the function. The formula is as follows:
[0109]
[0110] Among them, the function f(x) is differentiable and defined at the point x 1 and γ is the step size. When γ > 0 and is a sufficiently small value, there is f(x 1 ) ≥ f(x 2 ). Preferably, an improved gradient descent method can also be used to update the weight parameters of the initial graph convolutional neural network, such as the stochastic gradient descent method, the batch gradient descent method, the Adam gradient descent method, etc.
[0111] Step S440, retrain the graph convolutional neural network with updated weight parameters until the trained graph convolutional neural network meets the preset training end condition, and obtain the target graph convolutional neural network.
[0112] In specific implementation, the terminal based on the graph convolutional neural network with updated weight parameters, performs the process of steps 310 - 330 again until the cross-entropy calculation result between the vector predicted distribution result and the prior distribution result of the word to be predicted meets the condition of convergence of the cross-entropy loss function, that is, the trained graph convolutional neural network meets the preset training end condition, so as to use the graph convolutional neural network that meets the preset training end condition as the target graph convolutional neural network.
[0113] The technical solution of this embodiment is to obtain the prior distribution result of the initial word vector of the word to be predicted; then, calculate the cross entropy between the prediction distribution result and the prior distribution result to obtain the cross entropy calculation result; then, according to the cross entropy calculation result, update the weight parameters of the initial graph convolutional neural network by backpropagation; finally, retrain the graph convolutional neural network with the updated weight parameters until the trained graph convolutional neural network meets the preset training end condition to obtain the target graph convolutional neural network; in this way, the target word vector corresponding to each word can be accurately generated by the target graph convolutional neural network; furthermore, the target words belonging to the same entity can be accurately determined based on the target word vectors corresponding to each word.
[0114] In one embodiment, after the step of obtaining the corpus to be processed, the method further includes: performing word segmentation on the corpus to be processed to obtain the segmented corpus; and performing redundancy removal processing on the segmented corpus to obtain the corpus after redundancy removal.
[0115] Among them, the corpus to be processed can come from the corpora within different industries or from journal articles, etc.
[0116] Among them, the redundancy removal processing includes at least one of stop word removal processing, punctuation mark removal processing, and invisible character removal processing.
[0117] In specific implementation, the terminal can receive the corpus to be processed uploaded by other terminals or can crawl data on the website to obtain the corpus to be processed; then, the terminal needs to perform word segmentation on the corpus to be processed, and the word segmentation method can be the forward maximum matching method, the backward maximum matching method, the bidirectional matching word segmentation method, etc., which is not limited here, so as to obtain the segmented corpus. For example, for a sentence in the corpus to be processed, "The head office of Industrial and Commercial Bank of China is located in Xicheng District, Beijing.", after word segmentation, it gets "Industrial and Commercial Bank of China / head office / located in / Beijing / Xicheng District / ."
[0118] Then, the terminal performs redundancy removal processing on the corpus after word segmentation, such as stop word removal processing, punctuation mark removal processing, and invisible character removal processing, so as to obtain the corpus after redundancy removal. For example, for the corpus after word segmentation in the corpus to be processed, "Industrial and Commercial Bank of China / head office / located in / Beijing / Xicheng District / .", after redundancy removal processing, it gets the corpus after redundancy removal, "Industrial and Commercial Bank of China / head office / / Beijing / Xicheng District"
[0119] In the technical solution of this embodiment, by performing word segmentation on the corpus to be processed, the segmented corpus is obtained; then, the segmented corpus is processed to remove redundancy to obtain the corpus after redundancy removal. In this way, the target graph convolutional network does not need to encode the words without actual meaning in the corpus to be processed, but only needs to encode the words with actual meaning in the corpus to be processed to obtain the corresponding target word vectors to determine the target words belonging to the same entity; thus, the target words belonging to the same entity can be determined more efficiently.
[0120] In one embodiment, determining the target words belonging to the same entity according to the target word vectors corresponding to each word includes: determining the word to be matched among each word; determining the cosine similarity between the target word vector corresponding to the word to be matched and the target word vector corresponding to the candidate word; if the cosine similarity is greater than the preset cosine similarity threshold, it is determined that the word to be matched and the candidate word belong to the target words of the same entity.
[0121] Wherein, the candidate word is a word in the corpus to be processed other than the word to be matched.
[0122] In specific implementation, after the terminal obtains the target word vectors corresponding to each word, at least one word can be selected from each word in the corpus to be processed as the word to be matched and matched with the candidate word; the specific process is as follows: performing a cosine similarity operation on the target word vector corresponding to the word to be matched and the target word vector corresponding to the candidate word one by one to obtain the cosine similarity operation result; wherein, the candidate word is a word in the corpus to be processed other than the word to be matched; if there is a cosine similarity greater than the preset cosine similarity threshold in the above cosine similarity operation result, it is determined that the candidate word corresponding to the cosine similarity greater than the preset cosine similarity threshold matches successfully with the word to be matched and belongs to the same entity; optionally, the cosine distance between the target word vector corresponding to the word to be matched and the target word vector corresponding to the candidate word can also be determined, if the cosine distance is less than the preset cosine distance, it is determined that the word to be matched and the candidate word belong to the target words of the same entity; thus, all the words belonging to the same entity as the above word to be matched can be determined in the corpus to be processed; furthermore, the entity unification result of each word in the corpus to be processed is obtained.
[0123] In the technical solution of this embodiment, the to-be-matched vocabulary is determined in each vocabulary; then, the cosine similarity between the target word vector corresponding to the to-be-matched vocabulary and the target word vector corresponding to the candidate vocabulary is determined; wherein, the candidate vocabulary is the vocabulary other than the to-be-matched vocabulary in the to-be-processed corpus; finally, if the cosine similarity is greater than the preset cosine similarity threshold, it is determined that the to-be-matched vocabulary and the candidate vocabulary are the target vocabulary belonging to the same entity; thus, by determining the cosine similarity between the target word vector corresponding to the to-be-matched vocabulary and the target word vector corresponding to the candidate vocabulary, it is determined that the candidate vocabulary corresponding to the cosine similarity greater than the preset cosine similarity threshold belongs to the same entity as the to-be-matched vocabulary; thereby, all the vocabulary belonging to the same entity as the to-be-matched vocabulary in the to-be-processed corpus can be accurately determined; furthermore, the entity unified result of each vocabulary in the to-be-processed corpus can be accurately obtained.
[0124] In another embodiment, as Figure 5 shown, a method for processing entity information is provided. Taking the application of this method to a terminal as an example for illustration, the method includes the following steps:
[0125] Step S502: Taking each sample vocabulary in the to-be-processed sample corpus as a node, and the co-occurrence relationship between any two of the sample vocabularies as the edge of the node, a sample co-occurrence graph is constructed; the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the sample vocabulary corresponding to the node connected to the edge in the to-be-processed sample corpus; the sample vocabularies corresponding to the nodes connected by the edge appear simultaneously in the same sentence of the to-be-processed sample corpus.
[0126] Step S504: Inputting the sample co-occurrence graph into an initial graph convolutional neural network to obtain the initial word vectors corresponding to each of the sample vocabularies; the initial word vectors are obtained by the initial graph convolutional neural network encoding and processing the sample co-occurrence graph.
[0127] Step S506: Determine the to-be-predicted vocabulary among each of the sample vocabularies.
[0128] Step S508: Determine the adjacent sample vocabularies of the to-be-predicted vocabulary; wherein, in the sample co-occurrence graph, the nodes corresponding to the adjacent sample vocabularies are connected to the node corresponding to the to-be-predicted vocabulary.
[0129] Step S510: According to the initial word vectors corresponding to the adjacent sample vocabularies, obtain the prediction distribution result of the initial word vector of the to-be-predicted vocabulary.
[0130] Step S512: Obtain the prior distribution result of the initial word vector of the to-be-predicted vocabulary.
[0131] Step S514: Perform cross-entropy calculation on the prediction distribution result and the prior distribution result to obtain the cross-entropy calculation result.
[0132] Step S516: Update the weight parameters of the initial graph convolutional neural network by backpropagating the gradient according to the cross-entropy calculation result.
[0133] Step S518: Retrain the graph convolutional neural network with updated weight parameters until the trained graph convolutional neural network meets the preset training end condition, and obtain the target graph convolutional neural network.
[0134] Step S520: Obtain the corpus to be processed.
[0135] Step S522: Perform word segmentation on the corpus to be processed to obtain the segmented corpus.
[0136] Step S524: Perform redundancy removal on the segmented corpus to obtain the corpus after redundancy removal.
[0137] Step S526: Construct a word co-occurrence graph with each word in the corpus to be processed as a node and the co-occurrence relationship between any two words as the edge of the node; the edge has a corresponding weight; the weight is used to represent the number of co-occurrences of the words corresponding to the nodes connected to the edge in the corpus to be processed.
[0138] Step S528: Input the word co-occurrence graph into the target graph convolutional neural network to obtain the target word vectors corresponding to each word; the target word vectors are obtained by the target graph convolutional neural network encoding the word co-occurrence graph.
[0139] Step S530: Determine the target words belonging to the same entity according to the target word vectors corresponding to each word; the similarity between the target word vectors corresponding to the target words is greater than the preset similarity threshold.
[0140] It should be noted that the specific limitations of the above steps can refer to the specific limitations of a method for processing entity information above.
[0141] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0142] Based on the same inventive concept, an embodiment of this application further provides an entity information processing device for implementing an entity information processing method involved above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the entity information processing device provided below can refer to the limitations on an entity information processing method in the above text, and will not be repeated here.
[0143] In one embodiment, as Figure 6 shown, an entity information processing device is provided, including: a corpus acquisition module 610, a lexical co-occurrence graph construction module 620, a target word vector determination module 630, and a target vocabulary determination module 640, where:
[0144] The corpus acquisition module 610 is configured to acquire a corpus to be processed.
[0145] The lexical co-occurrence graph construction module 620 is configured to construct a lexical co-occurrence graph with each word in the corpus to be processed as a node and the co-occurrence relationship between any two words as the edge of the node; the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the words corresponding to the nodes connected by the edge in the corpus to be processed.
[0146] The target word vector determination module 630 is configured to input the lexical co-occurrence graph into a target graph convolutional neural network to obtain target word vectors corresponding to each word; the target word vectors are obtained by the target graph convolutional neural network encoding and processing the lexical co-occurrence graph.
[0147] The target vocabulary determination module 640 is configured to determine target vocabulary belonging to the same entity according to the target word vectors corresponding to each word; the similarity between the target word vectors corresponding to the target vocabulary is greater than a preset similarity threshold.
[0148] In one embodiment, the entity information processing device further includes: a sample co-occurrence graph construction module, configured to construct a sample co-occurrence graph with each sample word in the to-be-processed sample corpus as a node and the co-occurrence relationship between any two of the sample words as the edge of the node; the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the sample words corresponding to the nodes connected by the edge in the to-be-processed sample corpus; the sample words corresponding to the nodes connected by the edge appear simultaneously in the same sentence of the to-be-processed sample corpus; an initial word vector determination module, configured to input the sample co-occurrence graph into an initial graph convolutional neural network to obtain initial word vectors corresponding to the sample words; the initial word vectors are obtained by the initial graph convolutional neural network encoding the sample co-occurrence graph; a target graph convolutional neural network determination module, configured to train the initial graph convolutional neural network according to the initial word vectors corresponding to the sample words to obtain the target graph convolutional neural network.
[0149] In one embodiment, the target graph convolutional neural network determination module is specifically configured to determine a to-be-predicted word among the sample words; determine the adjacent sample words of the to-be-predicted word; wherein, in the sample co-occurrence graph, the nodes corresponding to the adjacent sample words are connected to the node corresponding to the to-be-predicted word; obtain a prediction distribution result of the initial word vector of the to-be-predicted word according to the initial word vectors corresponding to the adjacent sample words; train the initial graph convolutional neural network according to the difference between the prediction distribution result of the initial word vector of the to-be-predicted word and the prior distribution result to obtain the target graph convolutional neural network.
[0150] In one embodiment, the target graph convolutional neural network determination module is specifically configured to obtain a prior distribution result of the initial word vector of the to-be-predicted word; perform cross-entropy calculation on the prediction distribution result and the prior distribution result to obtain a cross-entropy calculation result; update the weight parameters of the initial graph convolutional neural network in the reverse gradient according to the cross-entropy calculation result; retrain the graph convolutional neural network with the updated weight parameters until the trained graph convolutional neural network meets a preset training end condition to obtain the target graph convolutional neural network.
[0151] In one embodiment, the entity information processing device further includes: a word segmentation module, configured to perform word segmentation on the to-be-processed corpus to obtain the segmented corpus; a redundancy removal module, configured to perform redundancy removal on the segmented corpus to obtain the corpus after redundancy removal.
[0152] In one embodiment, the target vocabulary determination module 640 is specifically configured to determine a to-be-matched vocabulary among all the vocabularies; determine the cosine similarity between the target word vector corresponding to the to-be-matched vocabulary and the target word vector corresponding to a candidate vocabulary, where the candidate vocabulary is a vocabulary in the to-be-processed corpus other than the to-be-matched vocabulary; and if the cosine similarity is greater than a preset cosine similarity threshold, determine that the to-be-matched vocabulary and the candidate vocabulary are target vocabularies belonging to the same entity.
[0153] Each module in the above entity information processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of the processor, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0154] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store corpus data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an entity information processing method.
[0155] Those skilled in the art can understand that Figure 7 the structure shown in
[0156] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0157] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above respective method embodiments.
[0158] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps in the above method embodiments.
[0159] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0160] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in this application can include at least one of a relational database and a non-relational database. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0161] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0162] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several variations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. An entity information processing method, characterized in that, the method includes: obtaining the corpus to be processed; constructing a lexical co-occurrence graph with each word in the corpus to be processed as a node and the co-occurrence relationship between any two words as the edge of the node; the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the words corresponding to the nodes connected by the edge in the corpus to be processed; inputting the lexical co-occurrence graph into a target graph convolutional neural network to obtain target word vectors corresponding to each word; the target word vector is obtained by the target graph convolutional neural network encoding the lexical co-occurrence graph, where the encoding process includes extracting local features and global features of each node in the lexical co-occurrence graph according to the edge connection relationship and the weight of each node in the lexical co-occurrence graph, and the target graph convolutional neural network is trained based on an initial graph convolutional neural network; determining target words belonging to the same entity according to the target word vectors corresponding to each word; the similarity between the target word vectors corresponding to the target words is greater than a preset similarity threshold.
2. The method according to claim 1, characterized in that, the method further includes: constructing a sample co-occurrence graph with each sample word in the sample corpus to be processed as a node and the co-occurrence relationship between any two sample words as the edge of the node; the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the sample words corresponding to the nodes connected by the edge in the sample corpus to be processed; the sample words corresponding to the nodes connected by the edge appear simultaneously in the same sentence of the sample corpus to be processed; inputting the sample co-occurrence graph into the initial graph convolutional neural network to obtain initial word vectors corresponding to each sample word; the initial word vector is obtained by the initial graph convolutional neural network encoding the sample co-occurrence graph; training the initial graph convolutional neural network according to the initial word vectors corresponding to each sample word to obtain the target graph convolutional neural network.
3. The method according to claim 2, characterized in that, the training the initial graph convolutional neural network according to the initial word vectors corresponding to each sample word to obtain the target graph convolutional neural network includes: determining a word to be predicted among each sample word; determining adjacent sample words of the word to be predicted; wherein, in the sample co-occurrence graph, the nodes corresponding to the adjacent sample words are connected to the node corresponding to the word to be predicted; obtaining a predicted distribution result of the initial word vector of the word to be predicted according to the initial word vectors corresponding to the adjacent sample words; training the initial graph convolutional neural network according to the difference between the predicted distribution result of the initial word vector of the word to be predicted and the prior distribution result to obtain the target graph convolutional neural network.
4. The method according to claim 3, characterized in that, Training the initial graph convolutional neural network according to the difference between the prediction distribution result and the prior distribution result of the initial word vector of the to-be-predicted word to obtain the target graph convolutional neural network includes: Obtaining the prior distribution result of the initial word vector of the to-be-predicted word; Performing cross-entropy calculation on the prediction distribution result and the prior distribution result to obtain a cross-entropy calculation result; According to the cross-entropy calculation result, reversely updating the weight parameters of the initial graph convolutional neural network by gradient; Retraining the graph convolutional neural network with updated weight parameters until the trained graph convolutional neural network meets a preset training end condition to obtain the target graph convolutional neural network.
5. The method according to claim 1, wherein, after the step of obtaining the to-be-processed corpus, the method further includes: Performing word segmentation on the to-be-processed corpus to obtain the word-segmented corpus; Performing redundancy removal on the word-segmented corpus to obtain the redundancy-removed corpus.
6. The method according to claim 1, wherein, according to the target word vectors corresponding to each of the words, determining the target words belonging to the same entity includes: Determining a to-be-matched word among each of the words; Determining the cosine similarity between the target word vector corresponding to the to-be-matched word and the target word vector corresponding to a candidate word; the candidate word is a word in the to-be-processed corpus other than the to-be-matched word; If the cosine similarity is greater than a preset cosine similarity threshold, determining that the to-be-matched word and the candidate word belong to the target words of the same entity.
7. An entity information processing device, wherein, the device includes: A corpus acquisition module, configured to acquire a to-be-processed corpus; A vocabulary co-occurrence graph construction module, configured to construct a vocabulary co-occurrence graph with each word in the to-be-processed corpus as a node and the co-occurrence relationship between any two words as the edge of the node; the edge has a corresponding weight; the weight is used to represent the co-occurrence times of the words corresponding to the nodes connected to the edge in the to-be-processed corpus; A target word vector determination module, configured to input the vocabulary co-occurrence graph into a target graph convolutional neural network to obtain the target word vectors corresponding to each of the words; the target word vector is obtained by the target graph convolutional neural network encoding and processing the vocabulary co-occurrence graph, wherein the encoding and processing includes extracting local features and global features of each node in the vocabulary co-occurrence graph according to the edge connection relationship and the weight of the edge of each node in the vocabulary co-occurrence graph, and wherein the target graph convolutional neural network is trained based on an initial graph convolutional neural network; A target word determination module, configured to determine the target words belonging to the same entity according to the target word vectors corresponding to each of the words; the similarity between the target word vectors corresponding to the target words is greater than a preset similarity threshold.
8. A computer device, including a memory and a processor, the memory stores a computer program, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, wherein, when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program, wherein, when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Topic detection or tracking method for network text big data
CN104462253A
Knowledge graph mining method, device and equipment and storage medium
CN109670051A