Abstract extraction method, device, terminal equipment and medium based on heterogeneous graph network

By connecting sentence nodes in heterogeneous graph networks and performing importance analysis, the problem of difficult to capture the remote dependence relationship of sentences in the prior art is solved, and a more efficient summary extraction effect is achieved.

CN113935314BActive Publication Date: 2025-06-20SHENZHEN PING AN SMART HEALTHCARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111231702.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-06-20
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

Existing abstract extraction models based on recursive neural networks are difficult to capture the remote dependencies of sentences in long documents or multiple documents, resulting in low accuracy of abstract extraction.

Method used

The abstract extraction method based on a heterogeneous graph network is used to obtain the sentence vector and position information of each sentence in the document to be extracted, and the similarity between sentences is calculated, and the sentences are connected as nodes in the heterogeneous graph network are analyzed to determine the abstract sentence.

Benefits of technology

Effectively capture the remote dependencies of sentences, improve the accuracy of abstract extraction, and better reflect the semantic relationship and position information of sentences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113935314B_ABST
    Figure CN113935314B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of artificial intelligence analysis, and particularly relates to a summary extraction method, device, terminal device and medium based on a heterogeneous graph network. According to the sentence vectors and position information of each sentence in the document to be extracted, the sentence similarity between corresponding sentences is obtained, and the sentences are used as nodes. According to the sentence similarity and position information, the nodes are connected to obtain a heterogeneous graph network. The importance of the information and connection relationships of each node in the heterogeneous graph network is analyzed, and it is determined that the sentences corresponding to the nodes with an importance greater than the threshold are the summary sentences of the document to be extracted, realizing summary extraction. The heterogeneous graph network constructed by combining the sentence similarity and position information of sentences can better reflect the long-distance dependence relationship of sentences, thereby improving the accuracy of summary extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence analysis, and particularly relates to a method, device, terminal device, and medium for abstract extraction based on a heterogeneous graph network. Background Art

[0002] Currently, an extractive document abstract refers to extracting relevant sentences from an original document and reorganizing them into an abstract. To effectively extract relevant sentences from a document, it is necessary to model the relationships between sentences. The existing models use recurrent neural networks to capture the relationships between sentences, but in the case of long documents or multiple documents, it is not easy for this recurrent neural network-based model to capture the long-range dependencies of sentences. Therefore, how to effectively capture the long-range dependencies of sentences during the abstract extraction process has become an urgent problem to be solved. Summary of the Invention

[0003] In view of this, embodiments of this application provide a method, device, terminal device, and medium for abstract extraction based on a heterogeneous graph network to solve the problem of how to effectively capture the long-range dependencies of sentences during the abstract extraction process.

[0004] In a first aspect, embodiments of this application provide a method for abstract extraction based on a heterogeneous graph network. The abstract extraction method includes:

[0005] Obtain the sentence vectors and position information of each sentence in the document to be extracted;

[0006] Obtain the sentence similarity between corresponding sentences according to the sentence vectors of any two sentences;

[0007] Use sentences as nodes, and connect the nodes according to the sentence similarity between sentences and the position information of each sentence to obtain a heterogeneous graph network, where the information of the nodes is the sentence vectors of the corresponding sentences;

[0008] Analyze the importance of the information and connection relationships of each node in the heterogeneous graph network, and determine the target sentence as the abstract sentence of the document to be extracted, where the target sentence is the sentence corresponding to the node with an importance greater than the threshold.

[0009] In a second aspect, embodiments of this application provide an apparatus for abstract extraction based on a heterogeneous graph network. The abstract extraction apparatus includes:

[0010] An information acquisition module for obtaining the sentence vectors and position information of each sentence in the document to be extracted;

[0011] A sentence similarity determination module for obtaining the sentence similarity between corresponding sentences according to the sentence vectors of any two sentences;

[0012] The heterogeneous graph construction module is used to take sentences as nodes and connect the nodes according to the sentence similarity between sentences and the position information of each sentence to obtain a heterogeneous graph network. Among them, the information of the nodes is the sentence vectors corresponding to the sentences.

[0013] The abstract extraction module is used to analyze the importance of the information and connection relationships of each node in the heterogeneous graph network, and determine the target sentence as the abstract sentence of the document to be extracted. The target sentence is the sentence corresponding to the node with an importance greater than the threshold.

[0014] In a third aspect, an embodiment of the present application provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the abstract extraction method described in the first aspect is implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the abstract extraction method described in the first aspect is implemented.

[0016] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device is enabled to execute the abstract extraction method described in the first aspect above.

[0017] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: According to the sentence vectors and position information of each sentence in the document to be extracted, the present application obtains the sentence similarity between the corresponding sentences, takes the sentences as nodes, and connects the nodes according to the sentence similarity and position information to obtain a heterogeneous graph network. The importance of the information and connection relationships of each node in the heterogeneous graph network is analyzed, and the sentence corresponding to the node with an importance greater than the threshold is determined as the abstract sentence of the document to be extracted, realizing abstract extraction. The heterogeneous graph network constructed by combining the sentence similarity and position information of sentences can better reflect the long-distance dependence relationship of sentences, thereby improving the accuracy of abstract extraction. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart of an abstract extraction method based on a heterogeneous graph network provided in Embodiment 1 of the present application;

[0020] Figure 2 It is a schematic flowchart of a method for abstract extraction based on a heterogeneous graph network provided in the second embodiment of the present application;

[0021] Figure 3 It is a schematic structural diagram of a device for abstract extraction based on a heterogeneous graph network provided in the third embodiment of the present application;

[0022] Figure 4 It is a schematic structural diagram of a terminal device provided in the fourth embodiment of the present application. Detailed implementation manners

[0023] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0024] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0025] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0026] As used in the specification and the appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0027] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0028] References to "one embodiment" or "some embodiments" in the description of this application mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc., which appear in different places in this specification, do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0029] The terminal device in the embodiments of this application may be a palm computer, a desktop computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud terminal device, a personal digital assistant (PDA), etc. The embodiments of this application do not impose any restrictions on the specific type of the terminal device.

[0030] The embodiments of this application may acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results.

[0031] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0032] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution is prior or subsequent. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0033] To illustrate the technical solution of this application, the following specific embodiments are used for illustration.

[0034] See Figure 1 , which is a schematic flowchart of a method for abstract extraction based on a heterogeneous graph network provided by Embodiment 1 of this application. The above method for abstract extraction is applied to a terminal device. As Figure 1As shown, the abstract extraction method may include the following steps:

[0035] Step S101, obtain the sentence vector and position information of each sentence in the document to be extracted.

[0036] Among them, a sentence may refer to the text segmented by punctuation marks such as a period, a comma, an exclamation mark, etc. The sentence vector of a sentence may refer to a vector representing the characteristics of the sentence. In this application, technologies such as the Bidirectional Encoder Representation from Transformers (BERT) model, Word2Vec, and Doc2vec can be used to calculate the sentence vector of a sentence. Of course, when calculating the sentence vectors of the sentences in the document to be extracted, the calculation of all sentences is based on one of the above technologies to ensure the consistency of the sentence vector structure.

[0037] The position information may refer to information such as the paragraph where the sentence is located in the document to be extracted. When processing the document to be extracted, each sentence can be annotated according to the defined sentence segmentation rules, and the annotation is used to reflect the paragraph where the sentence is located and the specific position in the paragraph. For example, for a document to be extracted with two paragraphs, using a period, an exclamation mark, a question mark, etc. as the sentence segmentation symbol, the first sentence of the first paragraph can be annotated as 1-1, the second sentence of the first paragraph can be annotated as 1-2, and the first sentence of the second paragraph can be annotated as 2-1.

[0038] In the terminal device of this application, corresponding software can be set to provide a corresponding service interface for configuration to provide an abstract extraction service. The user uploads the document to be extracted in the above service interface and triggers the abstract extraction service, and then the abstract of the document to be extracted can be obtained. An upload component is configured in the above service interface. Clicking on the above upload component can obtain any file stored in the terminal device. In addition, the upload component of the above service interface supports uploading multiple documents simultaneously. After performing the abstract extraction, an abstract is output for each document, thereby realizing batch abstract extraction.

[0039] Optionally, before obtaining the sentence vector and position information of each sentence in the document to be extracted, it further includes:

[0040] Perform text segmentation on the document to be extracted to obtain each sentence in the document to be extracted and the paragraph where the sentence is located, and use the paragraph where the sentence is located as the position information of the sentence;

[0041] Extract the feature vector of each sentence to obtain the sentence vector corresponding to the sentence.

[0042] Among them, most of the documents to be extracted are text documents composed of multiple sentences and paragraphs. It is necessary to perform text segmentation on the text documents to obtain each sentence and the position information of each sentence.

[0043] Calculate the feature vector of the sentence, and use the feature vector as the sentence vector representing the features of the sentence. Based on technologies such as the BERT model, Word2Vec, and Doc2vec, the feature vector of the sentence can be calculated.

[0044] The above-mentioned documents to be extracted can be documents in word format, TXT format, etc. This application can use the jieba tool and set the segmentation rules to segment the documents to be extracted into individual sentences, forming one or more sentence sets. Specifically, when segmenting the documents to be extracted, the sentences in a paragraph can be divided into a sentence set. Among them, the sentences in a paragraph of the document to be extracted can be used as a sentence set. If the position information of the sentence obtained is only the paragraph where the sentence is located, then each sentence set is marked, and each sentence in the sentence set is associated with the mark of the sentence set, so as to obtain the position information of the sentence.

[0045] In one implementation, the above-mentioned documents to be extracted can be PDF format files, pictures, etc. At this time, the PDF format files, pictures, etc. can be first converted into text format based on optical character recognition (OCR) technology, and then the abstract extraction is performed on the converted files.

[0046] The above-mentioned jieba tool is based on a prefix dictionary to achieve efficient word graph scanning, generate a directed acyclic graph composed of all possible word formation situations in the sentence, use dynamic programming to find the maximum probability path, find the maximum segmentation combination based on word frequency, and use the hidden Markov model (Hidden Markov Model, HMM) and Viterbi algorithm based on word formation ability to automatically segment the paragraph into sentences and segment the sentences into words.

[0047] Step S102, according to the sentence vectors of any two sentences, obtain the sentence similarity between the corresponding sentences.

[0048] Among them, the sentence similarity refers to the degree of similarity between two sentences. In the process of natural language processing, it is necessary to find similar sentences or approximate expressions of sentences, so that similar sentences can be grouped together. Therefore, after obtaining the sentence vector of each sentence, it is necessary to calculate the sentence similarity.

[0049] Common sentence similarity calculation methods include: edit distance calculation method, Jaccard coefficient calculation method, term frequency (TF) calculation method, inverse document frequency (TF-IDF) calculation method, and Word2Vec calculation method.

[0050] In this application, the TF calculation method can be adopted. Specifically, according to each sentence vector, a TF matrix is generated, and the similarity between two vectors in the TF matrix is calculated, that is, the cosine value of the included angle between the two vectors is solved. The larger the cosine value, the higher the sentence similarity between the sentences.

[0051] Step S103: Take sentences as nodes, and connect the nodes according to the sentence similarity between sentences and the position information of each sentence to obtain a heterogeneous graph network.

[0052] Among them, one node corresponds to one sentence, and the information of the node is the sentence vector of the corresponding sentence.

[0053] Each sentence forms a corresponding node in the heterogeneous graph network. The information of each node is the sentence vector of the corresponding sentence, and the connections between the nodes form edges. Among them, the connections between the nodes need to be based on the relationship between the corresponding sentences. In this application, the corresponding nodes in the heterogeneous graph network are connected according to the similarity between the corresponding sentences and the position information of each sentence.

[0054] Compared with a non-heterogeneous graph, the difference of a heterogeneous graph is that there can be multiple types of nodes and edges in the heterogeneous graph. Therefore, different types of nodes are allowed to have different-dimensional features or attributes, which can better express the real situation of the nodes.

[0055] In this application, there are two ways to construct edges. The first is to connect edges using sentence similarity. By setting a corresponding threshold, the sentence similarity is compared with the threshold. When the comparison result meets certain conditions, the nodes corresponding to the two sentences are connected; the second way is to connect the nodes corresponding to the two sentences whose position information meets the conditions using the above position information.

[0056] Optionally, connecting the nodes according to the sentence similarity between sentences and the position information of each sentence includes:

[0057] Connect the nodes corresponding to the two sentences whose sentence similarity is greater than the similarity threshold;

[0058] Connect the nodes corresponding to any two sentences whose position information indicates that they are in the same paragraph.

[0059] Among them, a sentence similarity greater than the similarity threshold can be used to indicate that there is a certain similarity relationship between two sentences, and there is a certain positional relationship between two sentences in the same paragraph. Therefore, there are two types of edges in the heterogeneous graph, one is the edge formed according to the sentence similarity, and the other is the edge formed according to the position information.

[0060] For example, for four sentences, namely the first sentence, the second sentence, the third sentence, and the fourth sentence, among them, the first sentence and the second sentence are in the same paragraph, the third sentence and the fourth sentence are in the same paragraph, the sentence similarities between the first sentence and the second sentence, the third sentence, and the fourth sentence are 0.5, 0.6, and 0.9 respectively, the sentence similarities between the second sentence and the third sentence, and the fourth sentence are 0.8 and 0.9 respectively, the sentence similarity between the third sentence and the fourth sentence is 0.6, the similarity threshold is set to 0.7, the nodes of the heterogeneous graph network are the first node, the second node, the third node, and the fourth node respectively, the first node corresponds to the first sentence, the information of the first node is the sentence vector of the first sentence, the second node corresponds to the second sentence, the information of the second node is the sentence vector of the second sentence, the third node corresponds to the third sentence, the information of the third node is the sentence vector of the third sentence, the fourth node corresponds to the fourth sentence, and the information of the fourth node is the sentence vector of the fourth sentence. Since the first sentence and the second sentence are in one paragraph, therefore, the first node and the second node are connected in the heterogeneous graph network. Since the sentence similarity between the first sentence and the fourth sentence is greater than 0.7, therefore, the first node and the fourth node are connected in the heterogeneous graph network. Since the third sentence and the fourth sentence are in one paragraph, therefore, the third node and the fourth node are connected in the heterogeneous graph network. Since the sentence similarity between the second sentence and the third sentence is greater than 0.7, therefore, the second node and the third node are connected in the heterogeneous graph network. Since the sentence similarity between the second sentence and the fourth sentence is greater than 0.7, therefore, the second node and the fourth node are connected in the heterogeneous graph network.

[0061] Step S104, analyze the importance of the information and connection relationships of each node in the heterogeneous graph network, and determine the target sentence as the summary sentence of the document to be extracted.

[0062] Among them, the target sentence is the sentence corresponding to the node with an importance greater than the threshold. The importance can refer to the degree to which the node participates in the heterogeneous graph network. The higher the degree of participation in the heterogeneous graph network, the higher the importance of the node.

[0063] The importance analysis can refer to the following rules: The more edges of the first type on a certain node, the more important the node is. Among them, the first type can refer to the edges formed according to the sentence similarity mentioned above. The fewer edges of the second type on a certain node, the more important the node is. Among them, the second type can refer to the edges formed according to the position information mentioned above. The information of the node is used to calculate the importance of the node at a certain layer of the graph neural network.

[0064] The extraction of abstract sentences is essentially to extract relatively important or key sentences from the document to be extracted. After the abstract sentences are extracted, they can be sorted according to their positions in the document to be extracted, and all the abstract sentences are integrated in the order of their appearance in the document to be extracted to obtain the abstract.

[0065] In this application, the analysis of the importance of nodes in the heterogeneous graph network can be achieved based on a trained Graph Convolution Networks (GCN), a trained Graph Attention Networks (GAN), a trained Graph Autoencoders (GA), a trained Graph Generative Networks (GGN), and a trained Graph Spatial-temporal Networks (GSN), etc.

[0066] Optionally, the information and connection relationships of each node in the heterogeneous graph network are analyzed for importance, and the target sentence is determined to be the abstract sentence of the document to be extracted. The target sentence is the sentence corresponding to the node with an importance greater than the threshold, including:

[0067] Use the GraphSAGE algorithm to analyze the importance of the information and connection relationships of each node in the heterogeneous graph network, and output the target nodes with an importance greater than the threshold. The sentence corresponding to the target node is the target sentence;

[0068] Determine the target sentence to be the abstract sentence of the document to be extracted.

[0069] This application uses a GCN composed of the GraphSAGE algorithm to sample and aggregate the edges in the heterogeneous graph network for training. The GraphSAGE algorithm can effectively aggregate information related to nodes. Through continuous classification iterations, the parameter data of this GCN can be obtained. Then, the parameter data of this GCN is used to predict the nodes in the above heterogeneous graph network to obtain key nodes or nodes with a relatively high degree of importance.

[0070] For example, for a document that includes two paragraphs, the document is segmented into two sentence sets by the jieba tool, with the sentences in one paragraph as one sentence set. The BERT model is used to obtain the sentence vectors of each sentence, the sentence similarity between each pair of sentences is calculated, and a heterogeneous graph network is constructed. Among them, each sentence corresponds to a node, and the nodes corresponding to two sentences with a sentence similarity greater than the similarity threshold and in the same paragraph are connected. The GraphSAGE algorithm is used to analyze the heterogeneous graph network, and the target node is output. The sentence corresponding to the target node is the sentence in the document that serves as the abstract.

[0071] The output of the target node is to output the information of the target node, that is, the sentence vector. According to the sentence vector, the corresponding sentence is retrieved from all sentences to determine the target sentence.

[0072] Spectral-based GCN expresses each node by training the embedding of each node through convolution, while the GCN based on the GraphSAGE algorithm expresses each node by sampling and aggregating the neighboring nodes of that node. Therefore, the GCN based on the GraphSAGE algorithm solves the problems that traditional GCN cannot estimate new nodes and must train the entire network.

[0073] The two most important stages in GraphSAGE are sampling and aggregation. Sampling is for the neighboring nodes of the node corresponding to sentence v. Through a fixed number of sampling methods, the embedding of the node corresponding to sentence v is aggregated to express the node corresponding to sentence v. Among them, set the number S of neighboring nodes required, that is, the sampling number. If the number of neighboring nodes of a node is less than S, then the sampling method with replacement is used until S neighboring nodes are sampled. If the number of neighboring nodes of a node is greater than S, then sampling without replacement is used to directly sample S neighboring nodes. The aggregation function uses the average aggregation function, concatenates the vectors of the k−1 layer of a certain node and its neighboring nodes, then performs an averaging operation on each dimension of the vector, and performs a non-linear transformation on the obtained result to generate the representation vector of the k layer of that node.

[0074] For example, a 2-layer (k = 2) graph neural network is constructed, the data of the neighboring nodes is set to 25, and the labeled document is used for iterative training to maximize the F1-score of its classification. After the training is completed, the parameter data of the graph neural network is obtained. Among them, the annotation method of the document is: if the sentence is the abstract of the document, it is annotated as 1, and if the sentence is not the abstract of the document, it is annotated as 0.

[0075] Optionally, determining the target sentence as the abstract sentence of the document to be extracted includes:

[0076] Check whether the number of sentences in the target sentence is greater than a preset value;

[0077] If it is detected that the number of sentences in the target sentence is greater than this value, then sort the importance levels of each sentence in the target sentence from largest to smallest, and determine the sentences ranked in the top N as the summary sentences of the document to be extracted, where N is an integer greater than zero.

[0078] Among them, according to requirements, the number of sentences in the summary may not exceed a certain quantity (i.e., the preset value). Therefore, when the number of sentences in the target sentence is large, it is necessary to select some sentences from them as summary sentences. After analyzing the importance levels of the nodes in this application, the importance levels of the nodes are also sorted, that is, the importance levels of the sentences are sorted, and the top N sentences (i.e., topN) with higher importance levels are taken as the summary sentences. In one implementation, the preset value is equal to N.

[0079] This application can be applied to the summary extraction of medical documents. By using sentence vectors and the position information of sentences in medical documents, a heterogeneous graph network of medical documents is constructed, and then the sentences in the medical document are predicted to obtain the key sentences of the medical document. Taking its top N sentences as the summary sentences of the medical document, the way of graph-structuring the document can well learn the information before and after the sentences and the meaning information of the hidden layer of the sentences, which is of great significance for the summary extraction of medical documents.

[0080] According to the sentence vectors and position information of each sentence in the document to be extracted in the embodiments of this application, the sentence similarity between corresponding sentences is obtained, and the sentences are used as nodes. According to the sentence similarity and position information, the nodes are connected to obtain a heterogeneous graph network. The importance levels of the information and connection relationships of each node in the heterogeneous graph network are analyzed, and the sentences corresponding to the nodes with importance levels greater than the threshold are determined as the summary sentences of the document to be extracted, realizing summary extraction. The heterogeneous graph network constructed by combining the sentence similarity and position information of sentences can better reflect the long-distance dependence relationship of sentences, thereby improving the accuracy of summary extraction.

[0081] See Figure 2 which is a schematic flowchart of a summary extraction method based on a heterogeneous graph network provided in the second embodiment of this application. As Figure 2 shown, this summary extraction method may include the following steps:

[0082] Step S201, obtain the sentence vectors and position information of each sentence in the document to be extracted.

[0083] Among them, the content of step S201 is the same as that of the above step S101, and reference can be made to the description of step S101, which will not be elaborated here.

[0084] Step S202: Obtain the vector similarity between corresponding sentences according to the sentence vectors of any two sentences.

[0085] Among them, referring to the content of the above step S102, the cosine distance between the sentence vectors of two sentences can be calculated through the TF calculation method. This cosine distance is the vector similarity between the sentences, and this vector similarity is used to characterize the similarity of the sentence vector features between the sentences.

[0086] Step S203: Obtain the word similarity between corresponding sentences according to the words in any two of the said sentences.

[0087] Among them, the similarity of sentences includes similarities in multiple dimensions such as the similarity of sentence vectors, the similarity of words, and the similarity of grammar. In order to obtain a better semantic relationship in this application, the similarity of sentence vectors and the similarity of words are fused to obtain the sentence similarity.

[0088] In the above step S202, any two sentences are taken to calculate the vector similarity between the sentences. In step S203, the objects to be processed need to be these two sentences, so as to obtain the word similarity between the sentences.

[0089] Optionally, obtaining the word similarity between corresponding sentences according to the words in any two of the said sentences includes:

[0090] Obtain the number of identical words in the two sentences and the number of words in each sentence of the two sentences;

[0091] According to the number of identical words in the two sentences and the number of words in each sentence of the two sentences, obtain the word similarity between the two sentences.

[0092] In this application, the word similarity can be calculated based on the word sequence between two sentences through the TextRank algorithm. Among them, the formula for calculating the word similarity by the above TextRank algorithm is as follows:

[0093]

[0094] In the formula, represents the i th sentence, represents the j th sentence, represents a word in any sentence. Among them, the meaning of the numerator part is the number of identical words that appear in both sentences at the same time, and the denominator is the sum of the logarithms of the number of words in the sentences. At this time, the denominator can suppress the advantage of longer sentences in similarity calculation.

[0095] Step S204, calculate the weighted average of the vector similarity and the word similarity between sentences, and determine the weighted average as the sentence similarity between the corresponding sentences.

[0096] Among them, when calculating the weighted average, the weights of the vector similarity and the word similarity can be set. For example, the weight of the vector similarity is 0.5. Therefore, the weight of the word similarity is 0.5, which is equivalent to adding the vector similarity and the word similarity and then taking the average. If you want to highlight the importance of the vector similarity, you can set the weight of the vector similarity to be greater than 0.5. If you want to highlight the importance of the word similarity, you can set the weight of the word similarity to be greater than 0.5.

[0097] Step S205, take the sentences as nodes, and connect the nodes according to the sentence similarity between sentences and the position information of each sentence to obtain a heterogeneous graph network.

[0098] Step S206, analyze the importance of the information and connection relationships of each node in the heterogeneous graph network, and determine the target sentence as the abstract sentence of the document to be extracted.

[0099] Among them, the contents of Step S205 and Step S206 are the same as those of Step S103 and Step S104 above respectively. For the descriptions of Step S103 and Step S104, please refer to them and will not be elaborated here.

[0100] For example, for a document, the document includes two paragraphs. The document is segmented into two sentence sets by the jieba tool, with the sentences in one paragraph as one sentence set. The sentence vectors of each sentence are obtained using the BERT model, and the vector similarity between sentences is calculated using the TF calculation method. The word similarity between sentences is calculated using the TextRank algorithm. The sentence similarity between sentences is calculated by combining the vector similarity and the word similarity. Then, according to the sentence similarity and the position information, a heterogeneous graph network is constructed. Among them, each sentence corresponds to a node, and the nodes corresponding to two sentences with a sentence similarity greater than the similarity threshold and in the same paragraph are connected. The GraphSAGE algorithm is used to analyze the heterogeneous graph network, and the target node is output. The sentence corresponding to the target node is the sentence used as the abstract in the document.

[0101] In the embodiment of the present application, the weighted average of the vector similarity and the word similarity between sentences is calculated, and the weighted average is determined as the sentence similarity between the corresponding sentences. By jointly considering the vector similarity and the word similarity between sentences, the sentence similarity between sentences can be more accurately characterized, which helps to improve the extraction of the semantic relationships of long sentences, making the subsequent construction of the heterogeneous graph network more accurate and the extracted abstract sentences more accurate.

[0102] Corresponding to the abstract extraction method in the above embodiment Figure 3The block diagram of the abstract extraction device based on the heterogeneous graph network provided in the third embodiment of the present application is shown. The above-mentioned abstract extraction device is applied to a terminal device, and a trained text classification model and a relationship extraction model are configured on the terminal device. The terminal device can be connected to a corresponding server, a conversation collector, etc. to obtain data such as dialogue sentences to be analyzed. For the convenience of description, only the parts related to the embodiments of the present application are shown.

[0103] See Figure 3 , the abstract extraction device includes:

[0104] An information acquisition module 31, configured to acquire the sentence vector and position information of each sentence in the document to be extracted;

[0105] A sentence similarity determination module 32, configured to obtain the sentence similarity between corresponding sentences according to the sentence vectors of any two sentences;

[0106] A heterogeneous graph construction module 33, configured to use sentences as nodes, and connect the nodes according to the sentence similarity between sentences and the position information of each sentence to obtain a heterogeneous graph network, where the information of the nodes is the sentence vector of the corresponding sentence;

[0107] An abstract extraction module 34, configured to analyze the importance degree of the information and connection relationship of each node in the heterogeneous graph network, and determine that the target sentence is the abstract sentence of the document to be extracted, where the target sentence is the sentence corresponding to the node with an importance degree greater than the threshold.

[0108] Optionally, the above-mentioned heterogeneous graph construction module 33 includes:

[0109] A first connection unit, configured to connect the nodes corresponding to two sentences with a vector similarity greater than the similarity threshold;

[0110] A second connection unit, configured to connect the nodes corresponding to any two sentences whose position information indicates that they are in the same paragraph.

[0111] Optionally, the above-mentioned abstract extraction device further includes:

[0112] A text segmentation module, configured to perform text segmentation on the document to be extracted before acquiring the sentence vector and position information of each sentence in the document to be extracted, obtain each sentence in the document to be extracted and the paragraph where the sentence is located, and use the paragraph where the sentence is located as the position information of the sentence;

[0113] A sentence vector determination module, configured to extract the feature vector of each sentence to obtain the sentence vector of the corresponding sentence.

[0114] Optionally, the above-mentioned sentence similarity determination module 32 includes:

[0115] A vector similarity determination unit, configured to obtain the vector similarity between corresponding sentences according to the sentence vectors of any two sentences;

[0116] A word similarity determination unit, configured to obtain the word similarity between corresponding sentences according to the words in any two sentences described above;

[0117] A sentence similarity determination unit, configured to calculate the weighted average of the vector similarity and the word similarity between sentences, and determine the weighted average as the sentence similarity between corresponding sentences.

[0118] Optionally, the above-mentioned word similarity determination unit includes:

[0119] An acquisition subunit, configured to acquire the number of identical words in two sentences and the number of words in each of the two sentences;

[0120] A similarity determination subunit, configured to obtain the word similarity between two sentences according to the number of identical words in two sentences and the number of words in each of the two sentences.

[0121] Optionally, the above-mentioned abstract extraction module 34 includes:

[0122] A detection unit, configured to detect whether the number of sentences in the target sentence is greater than a preset value;

[0123] A sorting unit, configured to, if it is detected that the number of sentences in the target sentence is greater than the preset value, sort the importance levels of each sentence in the target sentence from largest to smallest, and determine the sentences ranked in the top N as the abstract sentences of the document to be extracted, where N is an integer greater than zero.

[0124] Optionally, the above-mentioned abstract extraction module 34 includes:

[0125] A node determination unit, configured to use the GraphSAGE algorithm to analyze the importance levels of the information and connection relationships of each node in the heterogeneous graph network, and output target nodes with importance levels greater than a threshold, where the sentences corresponding to the target nodes are the target sentences;

[0126] A sentence determination unit, configured to determine the target sentence as the abstract sentence of the document to be extracted.

[0127] It should be noted that, for the information interaction, execution process, etc. between the above-mentioned modules, since they are based on the same concept as the method embodiment of the present application, their specific functions and the technical effects brought thereby can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0128] Figure 4 This is a schematic structural diagram of a terminal device provided in Embodiment 4 of the present application. As Figure 4 shown, the terminal device 4 in this embodiment includes: at least one processor 40 (Figure 4 only one is shown), a memory 41, and a computer program 42 stored in the memory 41 and executable on at least one processor 40. When the processor 40 executes the computer program 42, the steps in any of the above-described embodiments of the abstract extraction method are implemented.

[0129] The terminal device 4 may include, but is not limited to, a processor 40 and a memory 41. Those skilled in the art can understand that Figure 4 merely an example of the terminal device 4, which does not constitute a limitation on the terminal device 4. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0130] The so-called processor 40 may be a CPU, and the processor 40 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0131] The memory 41 may be an internal storage unit of the terminal device 4 in some embodiments, such as the hard disk or memory of the terminal device 4. The memory 41 may also be an external storage device of the terminal device 4 in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 4. Further, the memory 41 may also include both the internal storage unit and the external storage device of the terminal device 4. The memory 41 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 41 may also be used to temporarily store data that has been output or will be output.

[0132] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above-mentioned device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0133] To implement all or part of the processes in the above-mentioned method embodiments of this application, it can also be completed by a computer program product. When the computer program product runs on a terminal device, the terminal device can be enabled to execute the steps in the above-mentioned method embodiments.

[0134] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0135] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0136] In the embodiments provided in this application, it should be understood that the disclosed apparatus / terminal device and method can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the apparatus or unit can be in electrical, mechanical or other forms.

[0137] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0138] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. A method for abstract extraction based on a heterogeneous graph network, characterized in that, The abstract extraction method includes: Obtaining the sentence vectors and position information of each sentence in the document to be extracted; Obtaining the sentence similarity between corresponding sentences according to the sentence vectors of any two sentences; Taking sentences as nodes, and connecting the nodes according to the sentence similarity between sentences and the position information of each sentence to obtain a heterogeneous graph network, where the information of the nodes is the sentence vectors of the corresponding sentences; Performing an importance analysis on the information and connection relationships of each node in the heterogeneous graph network to determine that the target sentence is the abstract sentence of the document to be extracted, where the target sentence is the sentence corresponding to the node with an importance greater than the threshold, and the importance refers to the degree of participation of the node in the heterogeneous graph network. The importance analysis rule is that the more edges formed according to the sentence similarity on a certain node, the more important the node is, and the fewer edges formed according to the position information, the more important the node is; Among them, the connecting of nodes according to the sentence similarity between sentences and the position information of each sentence includes: Connecting the nodes corresponding to the two sentences with a sentence similarity greater than the similarity threshold; Connecting the nodes corresponding to any two sentences whose position information indicates being in the same paragraph.

2. The abstract extraction method according to claim 1, characterized in that, Before obtaining the sentence vectors and position information of each sentence in the document to be extracted, it further includes: Performing text segmentation on the document to be extracted to obtain each sentence in the document to be extracted and the paragraph where the sentence is located, and taking the paragraph where the sentence is located as the position information of the sentence; Extracting the feature vectors of each sentence to obtain the sentence vectors of the corresponding sentences.

3. The abstract extraction method according to claim 1, characterized in that, The obtaining of the sentence similarity between corresponding sentences according to the sentence vectors of any two sentences includes: Obtaining the vector similarity between corresponding sentences according to the sentence vectors of any two sentences; Obtaining the word similarity between the corresponding sentences according to the words in any two of the sentences; Calculating the weighted average of the vector similarity and the word similarity between sentences, and determining the weighted average as the sentence similarity between the corresponding sentences.

4. The abstract extraction method according to claim 3, characterized in that, The obtaining of the word similarity between the corresponding sentences according to the words in any two of the sentences includes: Obtaining the number of identical words in the two sentences and the number of words in each of the two sentences; Obtaining the word similarity between the two sentences according to the number of identical words in the two sentences and the number of words in each of the two sentences.

5. The abstract extraction method according to claim 1, characterized in that, The determining that the target sentence is the abstract sentence of the document to be extracted includes: Detecting whether the number of sentences in the target sentence is greater than a preset value; If it is detected that the number of sentences in the target sentence is greater than the preset value, then sorting the importance of each sentence in the target sentence from large to small, and determining the top N sentences as the abstract sentences of the document to be extracted, where N is an integer greater than zero.

6. The abstract extraction method according to any one of claims 1 to 5, characterized in that, The performing of an importance analysis on the information and connection relationships of each node in the heterogeneous graph network to determine that the target sentence is the abstract sentence of the document to be extracted, where the target sentence is the sentence corresponding to the node with an importance greater than the threshold includes: Use the GraphSAGE algorithm to analyze the importance of the information and connection relationships of each node in the heterogeneous graph network, and output target nodes with importance greater than a threshold. The sentences corresponding to the target nodes are target sentences; Determine that the target sentence is the abstract sentence of the document to be extracted.

7. A device for abstract extraction based on a heterogeneous graph network, characterized in that, The abstract extraction device includes: An information acquisition module for acquiring the sentence vectors and position information of each sentence in the document to be extracted; A sentence similarity determination module for obtaining the sentence similarity between corresponding sentences according to the sentence vectors of any two sentences; A heterogeneous graph construction module for taking sentences as nodes and connecting the nodes according to the sentence similarity between sentences and the position information of each sentence to obtain a heterogeneous graph network, where the information of the nodes is the sentence vectors corresponding to the sentences; An abstract extraction module for analyzing the importance of the information and connection relationships of each node in the heterogeneous graph network, and determining that the target sentence is the abstract sentence of the document to be extracted. The target sentence is the sentence corresponding to the node with importance greater than the threshold. The importance refers to the degree of participation of the node in the heterogeneous graph network. The importance analysis rule is that the more edges formed according to the sentence similarity on a certain node, the more important the node is, and the fewer edges formed according to the position information, the more important the node is; Among them, the heterogeneous graph construction module includes: a first connection unit for connecting the nodes corresponding to two sentences with a sentence similarity greater than the similarity threshold; A second connection unit for connecting the nodes corresponding to any two sentences whose position information indicates being in the same paragraph.

8. A terminal device, characterized in that, The terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the abstract extraction method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the abstract extraction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Text summarization method and device based on heterogeneous graph, storage medium and terminal

    CN113127632A

  • Text abstract generation method and device, computer equipment and storage medium

    CN113254593A