A method and system for generating text summaries based on NLP
By merging redundant clauses in text and building an approximate neighborhood graph, and determining important clauses using the NLP prediction model, the redundancy and computational complexity problems in text summary generation are solved, and more efficient text summary generation is achieved.
Patent Information
- Application Number
- CN202310731001.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-06-20
AI Technical Summary
In the prior art, the text summary generation method has problems of redundant information and high computational complexity, resulting in inaccurate generation and increased computational cost.
By dividing the text into multiple clauses, using the neighbor clause of the neighbor clause as the neighbor clause for merge, an approximate neighborhood graph is constructed, and the importance is determined based on the NLP prediction model to generate a text summary.
Improves the accuracy of text summary and reduces the complexity and cost of model training.
Smart Images

Figure CN117271760B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text summarization, and in particular to a method and system for generating text summarization based on NLP. Background Art
[0002] The generation of text summaries mainly includes extractive and generative methods. The generative method mainly uses deep learning technologies such as neural networks to directly generate corresponding text summaries. This type of technology is relatively complex, so extractive text summary generation is often used. The extractive solution is mainly based on NLP (Natural Language Processing) processing technology to select important content such as sentences from text documents as text summaries.
[0003] At present, punctuation marks are mainly used to divide clauses, and then the importance of clauses is evaluated. Finally, certain clauses or important contents are selected as text summaries based on their importance. However, for text, there is often a certain correlation between the sentences before and after the punctuation marks, and there may be repeated and redundant information between the two sentences. This not only leads to inaccurate text summary generation, but also increases the complexity of calculation, which indirectly increases the calculation cost. Summary of the invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a method and system for generating text summaries based on NLP, which can effectively improve the accuracy of text summaries and reduce the training cost of the model.
[0005] According to the first aspect of the present invention, the NLP-based text summary generation method includes:
[0006] Get object text;
[0007] Dividing the object text into a plurality of first clauses;
[0008] Determine an adjacent clause corresponding to each of the first clauses, calculate a first similarity between each of the first clauses and the corresponding adjacent clause, and merge the first clauses whose first similarity is less than a merge similarity threshold with the corresponding adjacent clauses to obtain a plurality of second clauses; wherein the adjacent clauses of the first clause are neighbor clauses adjacent to the first clause and / or neighbor clauses of neighbor clauses adjacent to the first clause in the plurality of first clauses;
[0009] Determine the adjacent clause corresponding to each of the second clauses, calculate the second similarity between each of the second clauses and the corresponding adjacent clause, and obtain an approximate neighborhood graph of the multiple second clauses according to the second similarity; the adjacent clause of the second clause is the neighbor clause adjacent to the second clause among the multiple second clauses;
[0010] Input the approximate neighborhood graph into a prediction model based on NLP to obtain the importance degree of each of the second clauses output by the prediction model;
[0011] Select several of the second clauses from the multiple second clauses to construct a text summary according to the importance degree;
[0012] The NLP-based text summary generation method according to the embodiments of the present invention has at least the following beneficial effects:
[0013] This method first performs a merging process on the multiple first clauses divided. Through the idea that the neighbor clause of the neighbor clause of a clause may be the neighbor clause of this clause, taking similarity as a measurement condition, the neighbor clause of a clause or the neighbor clause of the neighbor clause is used as the adjacent clause of this sentence, and the clauses that meet the requirements and their adjacents are merged into one clause. In this way, the adjacent sentences containing redundant and repeated information are merged, reducing the complexity of subsequent model training, and thus improving the accuracy of subsequent summary formation; then, find the neighbor clause corresponding to each merged clause, construct an approximate domain graph based on similarity, and finally, input the approximate domain graph into a prediction model for training to obtain the importance degree of the clause, and obtain the summary result based on the importance degree. This method effectively improves the accuracy of text summary generation and can also reduce the training cost of the model.
[0014] According to some embodiments of the present invention, the adjacent clause corresponding to the first clause is determined in the following manner:
[0015] Select the first neighbor clause before and after the first clause and the second neighbor clause before or after the first neighbor clause;
[0016] Calculate the similarity between the first clause and the first neighbor clause, and calculate the similarity between the first clause and the second neighbor clause;
[0017] When the similarity between the first clause and the first neighbor clause is less than the similarity between the first clause and the second neighbor clause, take the second neighbor clause as the adjacent clause of the first clause; otherwise, take the first neighbor clause as the adjacent clause of the first clause.
[0018] According to some embodiments of the present invention, the first clause is a clause in the object text whose effective information amount is greater than an information amount threshold, where the effective information amount is determined by information features included in the clause of the object text.
[0019] According to some embodiments of the present invention, the information features include: part-of-speech features in the clause, entity features in the clause, and constituent features in the clause.
[0020] According to some embodiments of the present invention, the effective information amount of the clause is calculated in the following manner:
[0021] N = αN1 + βN2 + γN3
[0022] Where N is the effective information amount, N1 represents the number of part-of-speech features in the clause, N2 represents the number of entity features in the clause, N3 represents the number of constituent features in the clause, and α, β, γ represent weight values.
[0023] According to some embodiments of the present invention, the prediction model is a model obtained based on the TextRank algorithm.
[0024] According to some embodiments of the present invention, the selecting several second clauses from the multiple second clauses according to the importance degree to construct a text summary includes:
[0025] Sorting the multiple second clauses according to the importance degree;
[0026] Selecting the second clauses with the top importance degrees from the multiple second clauses to construct a text summary.
[0027] According to the second aspect embodiments of the present invention, a text summary generation system based on NLP, the text summary generation system based on NLP includes:
[0028] An object text acquisition unit, configured to acquire an object text;
[0029] A first clause division unit, configured to divide the object text into a plurality of first clauses;
[0030] A first clause merging unit, configured to determine adjacent clauses corresponding to each first clause, calculate a first similarity between each first clause and the corresponding adjacent clause, and respectively merge the first clauses with the corresponding adjacent clauses where the first similarity is less than a merging similarity threshold to obtain a plurality of second clauses; wherein, the adjacent clause of the first clause is a neighbor clause adjacent to the first clause among the plurality of first clauses and / or a neighbor clause of the neighbor clause adjacent to the first clause;
[0031] A neighborhood graph determination unit, configured to determine an adjacent clause corresponding to each of the second clauses, calculate a second similarity between each of the second clauses and the corresponding adjacent clause, and obtain an approximate neighborhood graph of the multiple second clauses according to the second similarity; the adjacent clause of the second clause is a neighbor clause adjacent to the second clause among the multiple second clauses;
[0032] An importance determination unit, configured to input the approximate neighborhood graph into a prediction model based on NLP, and obtain the importance of each of the second clauses output by the prediction model;
[0033] A text summary formation unit, configured to select several of the second clauses from the multiple second clauses to construct a text summary according to the importance;
[0034] An electronic device according to an embodiment of the third aspect of the present invention includes at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the above-mentioned NLP-based text summary generation method.
[0035] A computer-readable storage medium according to an embodiment of the fourth aspect of the present invention stores computer-executable instructions for causing a computer to execute the above-mentioned NLP-based text summary generation method.
[0036] Other features and advantages of the present invention will be described in the following specification, and will be partially apparent from the specification, or understood by implementing the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:
[0038] Figure 1 is a flowchart of a method for generating a text summary based on NLP provided by an embodiment of the present invention;
[0039] Figure 2 is Figure 1 a flowchart of step S102 in;
[0040] Figure 3 is Figure 1 a flowchart of step S103 in;
[0041] Figure 4 is Figure 1 a flowchart of step S106 in;
[0042] Figure 5 It is a schematic structural diagram of a text summarization generation system based on NLP provided by an embodiment of the present invention;
[0043] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0044] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0045] In the description of the present invention, if the first, second, etc. are described only for the purpose of distinguishing technical features, they should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.
[0046] In the description of the present invention, it should be understood that for the orientation description, such as up, down, etc., the indicated orientation or positional relationship is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the indicated device or element must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present invention.
[0047] In the description of the present invention, it should be noted that unless otherwise clearly defined, words such as setting, installing, connecting, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.
[0048] Natural Language Processing (NLP) is a discipline that takes language as the object and uses computer technology to analyze, understand, and process natural language, that is, uses the computer as a powerful tool for language research, conducts quantitative research on language information with the support of the computer, and provides a language description that can be commonly used between humans and computers. TextRank is a keyword extraction and summarization algorithm in NLP.
[0049] A neighbor clause refers to another clause adjacent to a clause. For example: “…, it doesn't mean you have to maximize the value of every moment. Due to the influence of the original family, combined with your own weaknesses and life experiences, …” Among them, “it doesn't mean you have to maximize the value of every moment” is the neighbor clause of “due to the influence of the original family”.
[0050] An adjacent clause is a neighbor clause of a clause, or a neighbor clause of a neighbor clause of a clause.
[0051] The generation of text summaries mainly includes extractive and generative methods. The generative method mainly uses deep learning technologies such as neural networks to directly generate corresponding text summaries. This type of technology is relatively complex, so extractive text summaries are often used. The extractive method mainly selects important content such as sentences from text documents as text summaries based on NLP processing technology. At present, punctuation marks are mainly used to divide clauses, and then the importance of clauses is evaluated. Finally, certain clauses or important content are selected as text summaries based on the importance. However, for text, there is often a certain correlation between the sentences before and after punctuation marks, and there may be repeated and redundant information between the two sentences. This not only leads to inaccurate text summary generation, but also increases the complexity of calculations, which indirectly increases the calculation cost.
[0052] For example, one of the passages includes "The experiences in life and the heights of life you eventually reach are the meaning of your life on earth. The meaning of time is to let you realize the full meaning of your life on earth. Therefore, treat every minute and every second of your life with your heart and face it well. It is not necessary for you to maximize the value of every moment. Due to the influence of your original family, your own weaknesses and life experiences, you are bound to take some tortuous roads, and time can allow you to finally get out of those tortuous forks and return to the road of realizing your life ideals, until Until you reach a distant place with birds singing and flowers blooming, this is the magic of time, because only time can make people forget those things that they thought were important at the time, but in hindsight seem insignificant; only time can make people think clearly or see clearly the things that they thought were so at the time, but in fact are very different and the real trend of the times, so they will eventually be relieved and no longer worry about it, choose the things they are interested in and follow the direction of historical development; and for those things that you feel sad, you will eventually understand that even if you start over, it is still inevitable and you are powerless, so you will be more wise and free to continue moving forward, and finally be satisfied. "
[0053] For example, the clauses divided by punctuation marks, "The experiences in life and the heights of life you eventually reach are the meaning of your life on earth" and "The meaning of time is to allow you to realize the full meaning of your life on earth", contain repeated information, "The meaning of your life on earth". For example, the clauses divided by punctuation marks, "You are bound to walk on some tortuous roads" and "Time can allow you to finally walk out of those tortuous forks", contain redundant information.
[0054] The technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described below are only some embodiments of the present invention, not all embodiments.
[0055] See Figure 1 For an embodiment of the present invention, a method for generating a text summary based on NLP is provided. The text summary method based on NLP includes the following steps S101 to S106:
[0056] Step S101, obtain the object text. In this embodiment, the object text is the text for which a summary needs to be extracted. There is no limitation here.
[0057] Step S102, divide the object text into multiple first clauses.
[0058] Normally, the text is divided into clauses by punctuation marks. However, different from the conventional means, see Figure 2 , in some embodiments, the multiple first clauses are not all the clauses of the object text, but are selected based on specific conditions, specifically including:
[0059] Step S1021, determine each clause in the object text. Here, the division is performed by punctuation marks.
[0060] Step S1022, extract the part-of-speech features, entity features, and constituent features in each clause.
[0061] In this embodiment, the part-of-speech features include but are not limited to nouns, verbs, prepositions, etc., the entity features include but are not limited to place names, personal names, organization names, time, etc., and the constituent features include but are not limited to subjects, predicates, objects, etc. These types of features contain a lot of effective information in the sentence, and then these information are used to select high-quality sentences.
[0062] Step S1023, calculate the number of part-of-speech features, the number of entity features, and the number of constituent features in each clause.
[0063] Step S1024, calculate the effective information amount of the sentence according to the number of part-of-speech features, the number of entity features, and the number of constituent features in each clause:
[0064] N = αN1 + βN2 + βN3
[0065] Wherein, N is the effective information amount, N1 represents the number of part-of-speech features, N2 represents the number of entity features, N3 represents the number of constituent features, and α, β, γ represent weight values.
[0066] Step S1025, select the clauses with an effective information amount greater than a certain value as the first clauses. The certain value here can be set in advance, and the value thereof is not limited here.
[0067] Steps S1021 to S1025 use the effective information features in the sentence as important indicators for selecting summary sentences, selectively deleting low-quality sentences (with relatively low effective information volume), improving the effect of summary formation, and reducing the cost of training the obtained summary.
[0068] Step S103: Determine the adjacent clauses corresponding to each first clause, calculate the first similarity between each first clause and its corresponding adjacent clause, and merge the first clause with its corresponding adjacent clause where the first similarity is less than the merging similarity threshold respectively to obtain multiple second clauses; wherein, the adjacent clause of a first clause is the neighbor clause adjacent to the first clause and / or the neighbor clause of the neighbor clause adjacent to the first clause among multiple first clauses.
[0069] In the existing scenario, there may be a certain degree of repetition and redundancy between two adjacent sentences divided by punctuation marks. For example: the first sentence is "Mom asked Xiaoming to buy painkillers at the third window of the First People's Hospital", and the second sentence is "Xiaoming came to the third window of the First People's Hospital to buy painkillers according to Mom's request". There is a certain amount of repetitive and redundant information between the two sentences, and the two sentences can be completely merged into one sentence, thereby reducing the repetitive and redundant information, saving the cost of training the obtained summary subsequently, and also improving the accuracy of summary formation. Therefore, in step S103, multiple first clauses are screened. If there is a large amount of similar information between a first clause and its neighboring first clause, then the two first clauses are merged into a new clause, which is called the second clause here.
[0070] In this step, first determine the adjacent clause corresponding to each first clause. If the similarity between the first clause and the adjacent clause corresponding to the first clause is less than a preset threshold, then the two clauses are merged to obtain the second clause. Among them, the cosine similarity is used to measure the degree of similar information between the two, which will not be elaborated here.
[0071] It should be noted that in the above steps, this method does not directly consider the neighbor clause of a first clause as the adjacent clause of the first clause, but explores the idea that "the neighbor clause of the neighbor clause of a first clause may be the neighbor clause of the first clause". According to the similarity between a clause and its neighbor and the similarity between the clause and the neighbor of its neighbor, the adjacent clauses of each clause (clauses that are likely to have similar information to the clause) are selected, so as to calculate the approximate neighborhood graph of each clause subsequently.
[0072] The purpose of this is: to construct an approximate neighborhood structure of clauses by exploring the similarity between the neighbors of the neighbors of clauses, avoiding the similarity calculation between a large number of clauses that are repetitive and redundant with each other, and obtaining as high accuracy as possible while improving the calculation efficiency.
[0073] See Figure 3 , in some embodiments, the above step S103 determines the adjacent clause corresponding to the first clause in the following manner:
[0074] Step S1031: Select the first neighbor clauses located before and after the first clause and the second neighbor clauses located before or after the first neighbor clauses.
[0075] Step S1032: Calculate the similarity between the first clause and the first neighbor clauses, and calculate the similarity between the first clause and the second neighbor clauses.
[0076] Step S1033: When the similarity between the first clause and the first neighbor clauses is less than the similarity between the first clause and the second neighbor clauses, take the second neighbor clause as the adjacent clause of the first clause; otherwise, take the first neighbor clause as the adjacent clause of the first clause.
[0077] Among them, step S103 first determines the adjacent clause of each first clause. There are two possibilities for the adjacent clause here. The first possibility is the neighbor clauses located before and after the first clause, and the second possibility is the neighbor clause before the neighbor clause before the first clause and the neighbor clause after the neighbor clause after the first clause.
[0078] For example: The multiple first clauses include: …, the 11th first clause, the 12th first clause, the 13th first clause, the 14th first clause, the 15th first clause, …, where the adjacent clauses of the 13th first clause include the 11th first clause, the 12th first clause, the 14th first clause or the 15th first clause. Calculate the similarity degrees between the 13th first clause and the 11th first clause, the 12th first clause, the 14th first clause and the 15th first clause respectively. Then, based on the similarity degrees, judge the adjacent clause of the 13th first clause. For example:
[0079] The similarity degrees between the 13th first clause and the 11th first clause, the 12th first clause, the 14th first clause and the 15th first clause are K11, K12, K14, K15 respectively. Among them, K14 > K11 > K12 > K15, then the adjacent clause of the 13th first clause is the 14th first clause;
[0080] Step S104: Determine the adjacent clause corresponding to each second clause, calculate the second similarity between each second clause and the corresponding adjacent clause, and obtain the approximate neighborhood graph of the multiple second clauses according to the second similarity; the adjacent clause of the second clause is the neighbor clause adjacent to the second clause among the multiple second clauses.
[0081] In this step, first, the adjacent clauses corresponding to the second clauses are determined. Different from the above, the adjacent clauses here are the neighbor clauses before and the next neighbor clause after the second clause selected from multiple second clauses. Then, the similarity between the second clause and its adjacent clauses is calculated, and thus an approximate neighborhood graph of multiple second clauses can be constructed. It should be noted that constructing an approximate neighborhood graph based on the similarity of data points is common general knowledge in the field and will not be elaborated here.
[0082] Step S105: Input the approximate neighborhood graph into the prediction model based on NLP to obtain the importance degree of each second clause output by the prediction model.
[0083] In some embodiments of the present application, the prediction model is a model obtained based on the TextRank algorithm. In the art, TextRank is a classical method that scores each sentence in the compressed text, regards the score obtained by the sentence as the weight of the sentence, and finally selects a specified number of sentences with the top-ranked weights to form the final text summary.
[0084] Step S106: Select several second clauses from multiple second clauses to construct a text summary according to the importance degree.
[0085] See Figure 4 , in some embodiments, selecting several second clauses from multiple second clauses to construct a text summary includes:
[0086] Step S1061: Sort multiple second clauses according to the importance degree.
[0087] Step S1062: Select the second clauses with the top-ranked importance degrees from multiple second clauses to construct a text summary.
[0088] The present application has the following beneficial effects:
[0089] This method first performs a merging process on the multiple first clauses divided. Through the idea that the neighbor clause of the neighbor clause of a clause may be the neighbor clause of this clause, using similarity as the measurement condition, the neighbor clause or the neighbor clause of the neighbor clause of a clause is used as the adjacent clause of this sentence, and the clause that meets the requirements and its adjacent are merged into a new clause. In this way, adjacent sentences containing redundant and repetitive information are merged, reducing the complexity of subsequent model training, and thus improving the accuracy of subsequent summary formation. Then, the neighbor clauses of each merged second clause are found to construct an approximate neighborhood graph based on similarity. Finally, the approximate neighborhood graph is input into the prediction model for training to obtain the importance degree of the clauses, and the summary result is obtained based on the importance degree. This method effectively improves the accuracy of text summary generation and can also reduce the training cost of the model.
[0090] Compared with the method of dividing the first clause based on symbols, this method measures the effective information volume according to the part-of-speech features, entity features, and constituent features in the clause, and then selects a certain number of first clauses from the clauses divided based on punctuation marks based on the effective information volume. This reduces the interference of useless information and improves the accuracy of abstract formation.
[0091] For the sake of understanding, a set of embodiments are provided below, including a method for generating a text summary based on NLP. This method includes the following steps:
[0092] Step S201: Obtain the object text to be processed.
[0093] Step S202: Divide the object text into N clauses according to the punctuation marks for sentence segmentation.
[0094] Step S203: Judge the number of part-of-speech features, entity features, and constituent features of the N clauses, and calculate the effective information volume of each clause in the N clauses based on the following formula:
[0095] N = αN1 + βN2 + γN3
[0096] Where N is the effective information volume, N1 represents the number of part-of-speech features, N2 represents the number of entity features, N3 represents the number of constituent features, and α, β, γ represent weight values.
[0097] Step S204: Sort according to the effective information volume, set a threshold, and select M clauses with an effective information volume greater than the threshold from the N clauses.
[0098] Step S205: Calculate the similarity between each clause in the M clauses and its neighbor clauses, and calculate the similarity between each clause and the neighbor clauses of its neighbor clauses. Judge the size of the similarity, and take the neighbor clause with a greater similarity as the adjacent clause of the clause, that is, the adjacent clause is the neighbor clause of the clause or the neighbor clause of the neighbor clause.
[0099] Step S206: Set a threshold, calculate the similarity between the M clauses and their corresponding adjacent clauses. If the similarity is less than the threshold, fuse the clause and its corresponding adjacent clause into a new clause, and finally obtain L clauses.
[0100] Step S207: Determine the similarity between each clause in the L clauses and its neighbor clauses, and obtain an approximate neighborhood graph of the L clauses according to the similarity.
[0101] Step S208: Input the approximate neighborhood graph into the TextRank model to obtain the importance degree of each clause in the L clauses output by the TextRank model.
[0102] Step S209: Select the clauses with the top-ranked importance degrees from the L clauses to construct a text summary.
[0103] This method first performs a merging process on multiple first clauses obtained by partitioning. Based on the idea that the neighbor's neighbor of a clause may be the neighbor of that clause, using similarity as a metric condition, the neighbor clause or the neighbor's neighbor clause of a clause is regarded as the adjacent clause of that sentence. The clauses that meet the requirements and their adjacents are merged into a new clause. In this way, the sentences containing redundant and repetitive information are merged, reducing the complexity of subsequent model training and thus improving the accuracy of subsequent summary formation. Then, the neighbor clauses of each merged second clause are found to construct an approximate neighborhood graph based on similarity. Finally, the approximate neighborhood graph is input into a prediction model for training to obtain the importance degree of the clauses, and the summary result is obtained based on the importance degree. This method effectively improves the accuracy of text summary generation and also reduces the training cost of the model.
[0104] Compared with the method of partitioning the first clauses based on symbols, this method uses the part-of-speech features, entity features, and constituent features in the clauses partitioned based on punctuation marks as a measure of effective information volume, and then selects a certain number of first clauses from the clauses partitioned based on punctuation marks based on the effective information volume. This reduces the interference of useless information and improves the accuracy of summary formation.
[0105] See Figure 5 , an embodiment of the present application provides a text summary generation system based on NLP. The text summary generation system based on NLP includes an object text acquisition unit 1100, a first clause partitioning unit 1200, a first clause merging unit 1300, a neighborhood graph determination unit 1400, an importance degree determination unit 1500, and a text summary formation unit 1600, as follows:
[0106] The object text acquisition unit 1100 is used to acquire the object text.
[0107] The first clause partitioning unit 1200 is used to partition the object text into multiple first clauses.
[0108] The first clause merging unit 1300 is used to determine the adjacent clause corresponding to each first clause, calculate the first similarity between each first clause and the corresponding adjacent clause, and merge the first clause with the corresponding adjacent clause whose first similarity is less than the merging similarity threshold to obtain multiple second clauses; wherein, the adjacent clause of the first clause is the neighbor clause adjacent to the first clause and / or the neighbor's neighbor clause adjacent to the first clause among the multiple first clauses.
[0109] The neighborhood graph determination unit 1400 is used to determine the adjacent clauses corresponding to each second clause, calculate the second similarity between each second clause and the corresponding adjacent clause, and obtain the approximate neighborhood graph of multiple second clauses according to the second similarity; the adjacent clause of a second clause is the neighbor clause adjacent to the second clause among multiple second clauses.
[0110] The importance determination unit 1500 is used to input the approximate neighborhood graph into the prediction model based on NLP, and obtain the importance of each second clause output by the prediction model.
[0111] The text summary formation unit 1600 is used to select several second clauses from multiple second clauses according to the importance to construct a text summary.
[0112] It should be noted that the embodiments of this system and the above method embodiments are based on the same inventive concept. Therefore, the relevant content of the above method embodiments also applies to the embodiments of this system.
[0113] This system first performs a merging process on the divided multiple first clauses. Through the idea that the neighbor clause of the neighbor clause of a clause may be the neighbor clause of this clause, and using the similarity as the measurement condition, the neighbor clause or the neighbor clause of the neighbor clause of a clause is used as the adjacent clause of this clause, and the clauses that meet the requirements and their adjacents are merged into a new clause. In this way, the adjacent sentences containing redundant and repeated information are merged, reducing the complexity of subsequent model training and thus improving the accuracy of subsequent summary formation; then, the neighbor clauses of each merged second clause are found to construct an approximate neighborhood graph based on similarity. Finally, the approximate neighborhood graph is input into the prediction model for training to obtain the importance of the clauses, and the summary result is obtained based on the importance. The system can effectively improve the accuracy of text summary generation and also reduce the training cost of the model.
[0114] See Figure 6 , the embodiments of this application also provide an electronic device, which includes:
[0115] At least one memory;
[0116] At least one processor;
[0117] At least one program;
[0118] The program is stored in the memory, and the processor executes at least one program to implement the above-mentioned NLP-based text summary generation method of this disclosure.
[0119] This electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0120] The electronic device according to the embodiments of the present application will be described in detail below.
[0121] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure;
[0122] The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700 and are called by the processor 1600 to execute the NLP-based text summary generation method according to the embodiments of the present disclosure. This method first performs a merging process on a plurality of first sub-clauses divided. Through the idea that the neighbor clause of the neighbor clause of a clause may be the neighbor clause of this clause, using similarity as a measurement condition, the neighbor clause or the neighbor clause of the neighbor clause of a clause is used as the adjacent clause of this sentence, and the clauses that meet the requirements and their adjacents are merged into a new clause. In this way, adjacent sentences containing redundant and repetitive information are merged, reducing the complexity of subsequent model training and thus improving the accuracy of subsequent summary formation; then, the neighbor clauses of each second clause after merging are found, and an approximate domain graph based on similarity is constructed. Finally, the approximate domain graph is input into a prediction model for training to obtain the importance degree of the clauses, and the summary result is obtained based on the importance degree. This method effectively improves the accuracy of text summary generation and can also reduce the training cost of the model.
[0123] The input / output interface 1800 is used to implement information input and output;
[0124] The communication interface 1900 is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);
[0125] The bus 2000 transmits information between the various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900);
[0126] Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.
[0127] The embodiments of the present disclosure also provide a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described NLP-based text summary generation method. The method first performs a merging process on a plurality of divided first clauses. Based on the idea that the neighbor clause of the neighbor clause of a clause may be the neighbor clause of this clause, using similarity as a measurement condition, the neighbor clause or the neighbor clause of the neighbor clause of a clause is used as the adjacent clause of this sentence, and the clause that meets the requirements and its adjacent clause are merged into a new clause. In this way, adjacent sentences containing redundant and repetitive information are merged, reducing the complexity of subsequent model training and thus improving the accuracy of subsequent summary formation. Then, the neighbor clauses of each merged second clause are found, and an approximate neighborhood graph based on similarity is constructed. Finally, the approximate neighborhood graph is input into a prediction model for training to obtain the importance degree of the clauses, and the summary result is obtained based on the importance degree. This method effectively improves the accuracy of text summary generation and also reduces the training cost of the model.
[0128] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0129] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.
[0130] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0131] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0132] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0133] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0134] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a, b, and c", where a, b, c can be single or multiple.
[0135] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0136] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0137] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0138] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0139] The above has described the embodiments of the present invention in detail with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made without departing from the spirit of the present invention within the knowledge scope of those of ordinary skill in the art to which the present invention pertains.
Claims
1. A method for generating text summaries based on NLP, characterized in that, The NLP-based text summarization method includes: Obtain the object text; Divide the object text into multiple first clauses; Determine the adjacent clause corresponding to each first clause, calculate the first similarity between each first clause and the corresponding adjacent clause, and merge the first clause and the corresponding adjacent clause where the first similarity is less than the merging similarity threshold to obtain multiple second clauses; wherein, the adjacent clause of the first clause is the neighbor clause adjacent to the first clause and / or the neighbor clause of the neighbor clause adjacent to the first clause among the multiple first clauses; determine the adjacent clause corresponding to the first clause in the following manner: Select the first neighbor clause before and after the first clause and the second neighbor clause before or after the first neighbor clause; Calculate the similarity between the first clause and the first neighbor clause, and calculate the similarity between the first clause and the second neighbor clause; When the similarity between the first clause and the first neighbor clause is less than the similarity between the first clause and the second neighbor clause, use the second neighbor clause as the adjacent clause of the first clause; otherwise, use the first neighbor clause as the adjacent clause of the first clause; Determine the adjacent clause corresponding to each second clause, calculate the second similarity between each second clause and the corresponding adjacent clause, and obtain the approximate neighborhood graph of the multiple second clauses according to the second similarity; the adjacent clause of the second clause is the neighbor clause adjacent to the second clause among the multiple second clauses; Input the approximate neighborhood graph into the NLP-based prediction model to obtain the importance degree of each second clause output by the prediction model; Select several second clauses from the multiple second clauses according to the importance degree to construct a text summary.
2. The method for generating a text summary based on NLP according to claim 1, wherein The first clause is a clause in the object text whose effective information amount is greater than the information amount threshold, where the effective information amount is determined by the information features included in the clause of the object text.
3. The method for generating a text summary based on NLP according to claim 2, wherein The information features include: part-of-speech features in the clause, entity features in the clause, and constituent features in the clause.
4. The method for generating a text summary based on NLP according to claim 3, wherein, Calculate the effective information amount of the clause in the following manner: N = αN1 + βN2 + γN3 where N is the effective information amount, N1 represents the number of part-of-speech features in the clause, N2 represents the number of entity features in the clause, N3 represents the number of constituent features in the clause, and α, β, γ represent weight values.
5. The method for generating a text summary based on NLP according to any one of claims 1 to 4, characterized in that, The prediction model is a model obtained based on the TextRank algorithm.
6. The method for generating a text summary based on NLP according to claim 1, wherein The selecting several second clauses from the multiple second clauses according to the importance degree to construct a text summary includes: Sort the multiple second clauses according to the importance degree; Select the second clauses with the top importance degrees from the multiple second clauses to construct a text summary.
7. A text summarization generation system based on NLP, characterized in that, The NLP-based text summary generation system includes: An object text acquisition unit for acquiring an object text; A first clause division unit for dividing the object text into multiple first clauses; The first clause merging unit is configured to determine the adjacent clauses corresponding to each of the first clauses, calculate the first similarity between each of the first clauses and the corresponding adjacent clauses, and merge the first clauses with the corresponding adjacent clauses where the first similarity is less than the merging similarity threshold, to obtain a plurality of second clauses; wherein, the adjacent clause of a first clause is a neighbor clause adjacent to the first clause and / or a neighbor clause of a neighbor clause adjacent to the first clause among the plurality of first clauses; the adjacent clause corresponding to the first clause is determined in the following manner: Select a first neighbor clause before and after the first clause and a second neighbor clause before or after the first neighbor clause; Calculate the similarity between the first clause and the first neighbor clause, and calculate the similarity between the first clause and the second neighbor clause; When the similarity between the first clause and the first neighbor clause is less than the similarity between the first clause and the second neighbor clause, use the second neighbor clause as the adjacent clause of the first clause; otherwise, use the first neighbor clause as the adjacent clause of the first clause; The neighborhood graph determination unit is configured to determine the adjacent clauses corresponding to each of the second clauses, calculate the second similarity between each of the second clauses and the corresponding adjacent clauses, and obtain an approximate neighborhood graph of the plurality of second clauses according to the second similarity; the adjacent clause of a second clause is a neighbor clause adjacent to the second clause among the plurality of second clauses; The importance determination unit is configured to input the approximate neighborhood graph into a prediction model based on NLP, to obtain the importance of each of the second clauses output by the prediction model; The text summary forming unit is configured to select several of the second clauses from the plurality of second clauses to construct a text summary according to the importance; 8. An electronic device, characterized in that: Comprising at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor, so that the at least one control processor can execute the NLP-based text summary generation method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the NLP-based text summary generation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-document abstract sentence generating method
CN104778157A
Text abstract generation method and apparatus, and computer device and storage medium
WO2022262266A1