Sentence compression method, device, equipment and computer-readable storage medium

By dividing the long statement to be compressed into phrases, filtering key statements, generating syntax trees and calculating information density, the problem of low information retention rate in statement compression is solved, and higher compression accuracy and information retention rate are achieved.

CN111444703BActive Publication Date: 2025-05-09CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010142436.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-03-04
Publication Date
2025-05-09
Estimated Expiration
2040-03-04

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively retain important information of the long statement to be compressed during the statement compression process, resulting in a low retention rate of compressed statement information.

Method used

The long statement to be compressed is divided into phrases through the separator, the key statements are filtered out, the syntax tree is generated, the information density of the candidate compression statement is calculated, and the candidate compression statement with the largest information density is finally selected as the compression statement.

Benefits of technology

Improve the accuracy of statement compression, ensuring that more important information is retained during the compression process, thereby improving information retention rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111444703B_ABST
    Figure CN111444703B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and discloses a sentence compression method, comprising the following steps: dividing a long sentence to be compressed into one or more short sentences by a separator; screening one or more key sentences from the one or more short sentences by a preset text classification model, and splicing the one or more key sentences into a test sentence; if the byte length of the test sentence is greater than the minimum compression byte length, segmenting the test sentence to obtain a segmented sentence; generating a syntactic tree corresponding to the segmented sentence by a preset strategy algorithm; calculating the information density of each candidate compressed sentence according to the importance index of proper nouns and keywords in the sentence; and outputting the candidate compressed sentence with the largest information density as the final compressed sentence. The present invention also discloses a sentence compression device, equipment, and computer-readable storage medium. The sentence compression method provided by the present invention improves the efficiency of sentence compression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a sentence compression method, device, equipment and computer-readable storage medium. Background Art

[0002] At present, sentence compression technology is one of the important research directions in the field of natural language processing. It can effectively trim redundant information from the original sentence, retain the main idea, and facilitate readers and machine recognition. It is currently widely used in the fields of automatic extraction of topics, automatic generation of summaries, and video and speech processing. Machines cannot parse and understand long and difficult sentences expressed by customers well, so they give wrong or useless responses, which causes a very poor experience for users. How to compress the long sentences to be compressed while retaining the important information of the original long sentences to be compressed as much as possible is a technical problem that needs to be solved urgently in this field. Summary of the invention

[0003] The main purpose of the present invention is to provide a sentence compression method, device, equipment and computer-readable storage medium, aiming to solve the technical problem of low retention rate of important information in sentences caused by sentence compression.

[0004] To achieve the above object, the present invention provides a statement compression method, which comprises the following steps:

[0005] Split the long sentence to be compressed into at least two short sentences using a separator;

[0006] Filtering at least one key sentence from the at least two short sentences by using a preset text classification model, and splicing the at least one key sentence into a test sentence;

[0007] Determining whether the byte length of the test statement is greater than or equal to the minimum compressed byte length;

[0008] If the byte length of the test sentence is greater than or equal to the minimum compressed byte length, segmenting the test sentence to obtain a segmented sentence;

[0009] Generate a syntax tree corresponding to the word segmentation sentence by a preset strategy algorithm, wherein the syntax tree includes at least one candidate compression sentence;

[0010] Calculating the information volume of the candidate compressed sentences;

[0011] determining importance indexes of proper nouns and keywords in the sentence based on the information amount;

[0012] Based on the importance index of the proper noun and the keyword in the sentence, the information density of each candidate compressed sentence is calculated:

[0013] The candidate compressed sentence with the largest information density is taken as the final compressed sentence.

[0014] Optionally, the generating a syntax tree from the segmented sentence by a preset strategy algorithm, wherein the syntax tree includes at least one candidate compression sentence including:

[0015] Calling a library function to perform dependency syntactic analysis and semantic dependency analysis on the segmented sentence to obtain an analysis result;

[0016] At least one syntax tree is constructed according to the analysis result, wherein the syntax tree includes at least one candidate compression sentence.

[0017] Optionally, the calling library function performs dependency syntactic analysis and semantic dependency analysis on the word segmentation sentence, and the analysis results obtained include:

[0018] Calling a library function to annotate the structure and semantics of the segmented sentence respectively to obtain an annotation result;

[0019] Traversing a preset annotation set to determine whether the annotation result exists in the annotation set;

[0020] If the annotation result exists in the annotation set, the analysis result is obtained.

[0021] Optionally, constructing at least one syntax tree according to the analysis result, wherein the syntax tree includes at least one candidate compression sentence including:

[0022] Segmenting the analysis results by using a segmentation tool HanLP to obtain candidate compressed sentences;

[0023] The candidate compressed sentence is cached into a preset initial syntax tree to obtain at least one syntax tree, wherein the syntax tree includes at least one candidate compressed sentence.

[0024] Optionally, calculating the information density of each candidate compressed sentence includes:

[0025] The information content of the candidate compressed sentence is calculated by the following formula;

[0026]

[0027] Among them, W i For the words in the text, tf ij W i Frequency in the text, idf ij W i The inverse document frequency of , w is the additional weight of proper nouns or keywords.

[0028] Optionally, the calculating the information density of each candidate compressed sentence based on the importance index of the proper noun and the keyword in the sentence includes:

[0029] Based on the importance index of the proper noun and the keyword in the sentence, the information density of each candidate compressed sentence is calculated by the following formula:

[0030]

[0031] Among them, D(s k ) is the information density, I(w i ) is the word w i The amount of information, P is the sentence S K List of all words in L(S K ) is a sentence S K The sentence length, S K For sentences.

[0032] Optionally, after calculating the information amount of the candidate compressed sentence by the following formula, the method further includes:

[0033] Determine whether there is a request to obtain proper nouns and keywords;

[0034] If there is a request to obtain the proper noun and the keyword, the current usage frequencies of the proper noun and the keyword are updated.

[0035] Furthermore, to achieve the above object, the present invention also provides a statement compression device, which includes the following modules:

[0036] A segmentation module, used for segmenting the long sentence to be compressed into at least two short sentences by using a separator;

[0037] A screening module, configured to screen out at least one key sentence from the at least two short sentences by using a preset text classification model, and to splice the at least one key sentence into a test sentence;

[0038] A length determination module, used to determine whether the byte length of the test statement is greater than or equal to the minimum compressed byte length;

[0039] A word segmentation module, configured to segment the test sentence to obtain a segmented sentence if the byte length of the test sentence is greater than or equal to the minimum compressed byte length;

[0040] A construction module, used to generate a syntax tree corresponding to the segmented sentence by a preset strategy algorithm, wherein the syntax tree includes at least one candidate compressed sentence;

[0041] An information volume calculation module, used for calculating the information volume of the candidate compressed sentence;

[0042] An importance index determination module, used to determine the importance index of proper nouns and keywords in the sentence based on the information amount;

[0043] A calculation module is used to calculate the information density of each candidate compressed sentence based on the importance index of the proper noun and the keyword in the sentence:

[0044] The compressed sentence output module is used to take the candidate compressed sentence with the largest information density as the final compressed sentence.

[0045] Optionally, the building block includes the following units:

[0046] An analysis unit, used for calling a library function to perform dependency syntactic analysis and semantic dependency analysis on the segmentation sentence to obtain an analysis result;

[0047] A construction unit is used to construct at least one syntax tree according to the analysis result, wherein the syntax tree includes at least one candidate compression sentence.

[0048] Optionally, the analysis unit is used to:

[0049] Calling a library function to annotate the structure and semantics of the segmented sentence respectively to obtain an annotation result;

[0050] Traversing a preset annotation set to determine whether the annotation result exists in the annotation set;

[0051] If the annotation result exists in the annotation set, the analysis result is obtained.

[0052] Optionally, the building block is used to:

[0053] Segmenting the analysis results by using a segmentation tool HanLP to obtain candidate compressed sentences;

[0054] The candidate compressed sentence is cached into a preset initial syntax tree to obtain at least one syntax tree, wherein the syntax tree includes at least one candidate compressed sentence.

[0055] Optionally, the information volume calculation module includes:

[0056] An information volume calculation unit, used to calculate the information volume of the candidate compressed sentence by the following formula;

[0057]

[0058] Among them, W i For the words in the text, tf ij W i Frequency in the text, idfij W i The inverse document frequency of , w is the additional weight of proper nouns or keywords.

[0059] Optionally, the calculation module includes:

[0060] The information density calculation unit is used to calculate the information density of each candidate compressed sentence based on the importance index of the proper noun and the keyword in the sentence by the following formula:

[0061]

[0062] Among them, D(s k ) is the information density, I(w i ) is the word w i The amount of information, P is the sentence S K List of all words in L(S K ) is a sentence S K The sentence length, S K Optionally, the sentence compression device further comprises:

[0063] A request judgment module is used to judge whether there is a request to obtain proper nouns and keywords;

[0064] An updating module is used to update the current usage frequency of the proper noun and the keyword if there is a request to obtain the proper noun and the keyword.

[0065] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a statement compression device, which includes a memory, a processor, and a statement compression program stored in the memory and executable on the processor, and the statement compression program, when executed by the processor, implements the steps of the statement compression method as described in any one of the above-mentioned items.

[0066] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, on which a statement compression program is stored, and when the statement compression program is executed by a processor, the steps of the statement compression method as described in any one of the above items are implemented.

[0067] In the embodiment of the present invention, emphasis is placed on first classifying sentences into long and short sentences and assembling them into test sentences. If the byte length of the test sentence is greater than or equal to the minimum compression byte length, the test sentence is segmented. A syntax tree is used in the compression process. In order to ensure short sentences that retain more important information, the information density of each candidate compressed sentence is innovatively calculated through an information density formula. The information density is used as the basis for obtaining the final compressed short sentence, which can improve the accuracy of compression. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a structural schematic diagram of the operating environment of the statement compression device involved in the embodiment of the present invention;

[0069] Figure 2 It is a flowchart of the first embodiment of the sentence compression method of the present invention;

[0070] Figure 3 for Figure 2 A detailed flow chart of an embodiment of step S50; Figure 4 The figure is a functional module diagram of an embodiment of a sentence compression device of the present invention.

[0071] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0072] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0073] The sentence compression method involved in the embodiment of the present invention is mainly applied to a sentence compression device, which may be a device with display and processing functions such as a PC, a portable computer, or a mobile terminal.

[0074] Reference Figure 1 , Figure 1 Schematic diagram of the hardware structure of the statement compression device involved in the embodiment of the present invention. In the embodiment of the present invention, the statement compression device may include a processor 1001 (such as a CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components; the user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard); the network interface 1004 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface); the memory 1005 may be a high-speed RAM memory, or a stable memory (non-volatile memory), such as a disk memory, and the memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.

[0075] Those skilled in the art will understand that Figure 1 The hardware structure shown in the figure does not constitute a limitation on the statement compression device, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0076] Continue to refer to Figure 1 , Figure 1The memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, and a statement compression program.

[0077] exist Figure 1 In the embodiment, the network communication module is mainly used to connect to the server and perform data communication with the server; and the processor 1001 can call the statement compression program stored in the memory 1005 and execute the statement compression method provided by the embodiment of the present invention.

[0078] An embodiment of the present invention provides a statement compression method.

[0079] Reference Figure 2 , Figure 2 1 is a flow chart of a first embodiment of the sentence compression method of the present invention. In this embodiment, the sentence compression method includes the following steps:

[0080] Step S10, dividing the long sentence to be compressed into at least two short sentences by using a separator;

[0081] In this embodiment, the long sentence to be compressed can be divided into at least two short sentences by separators such as commas, question marks, periods, exclamation marks, spaces, etc. The separators are set in advance in the program. Therefore, if there are separators in the sentence, the long sentence to be compressed can be divided into at least two short sentences.

[0082] Step S20, selecting at least one key sentence from at least two short sentences by using a preset text classification model, and splicing the at least one key sentence into a test sentence;

[0083] In this embodiment, the preset text classification model here refers to a pre-trained text classification model. The pre-trained text classification model can be used to classify key sentences from at least two short sentences. After the key sentences are classified, the key sentences can be spliced ​​by string splicing.

[0084] For example, "I want to buy e-life insurance", the preset text classification model can be used to determine whether the short sentence is a key sentence. If the short sentence is a key sentence, the key sentence is retained; if the short sentence is a non-key sentence, the key sentence is deleted.

[0085] Step S30, determining whether the byte length of the test statement is greater than or equal to the minimum compressed byte length;

[0086] In this embodiment, in order to obtain a test statement that meets the preset byte length, the minimum compressed byte length is pre-defined in this embodiment, so it is only necessary to determine whether the byte length of the current test statement is greater than or equal to the minimum compressed byte length.

[0087] Step S40, if the byte length of the test sentence is greater than or equal to the minimum compressed byte length, segment the test sentence to obtain a segmented sentence;

[0088] In this embodiment, a word segmentation technique may be used to segment the sentence, such as the jieba word segmentation technique.

[0089] Step S50, generating a syntax tree corresponding to the word segmentation sentence by a preset strategy algorithm, wherein the syntax tree includes at least one candidate compression sentence;

[0090] In this embodiment, the word segmentation sentence is generated into a syntactic tree through a preset strategy algorithm. The syntactic tree is composed of a root node and left and right nodes, and includes at least two layers. Except for the first time, which is a root node, each other layer is composed of left and right nodes. In order to obtain a syntactic tree that meets the compression requirements, dependency syntactic analysis and semantic dependency analysis can be pre-set in each node so that different words of the word segmentation sentence can be cached in different nodes. Each layer is a compression method.

[0091] Step S60, calculating the information volume of the candidate compressed sentence;

[0092] In this embodiment, the calculation can be performed using a preset formula for calculating the amount of information, for example,

[0093]

[0094] Among them, W i For the words in the text, tf ij W i Frequency in the text, idf ij W i The inverse document frequency of , w is the additional weight of proper nouns or keywords.

[0095] Step S70, determining the importance index of proper nouns and keywords in the sentence based on the information amount;

[0096] Step S80, calculating the information density of each candidate compressed sentence based on the importance index of the proper noun and the keyword in the sentence;

[0097] In this embodiment, an importance index is set for each keyword and proper noun to facilitate calculation of the information density of each candidate compressed sentence. The larger the importance index, the greater the amount of information contained in the word. For example, the importance index of "Xiangshan" is greater than that of "Beijing". Assuming that the nth layer of the syntax tree is "I love Beijing Xiangshan", the n+1th layer is "I love Beijing", and the n+2th layer is "I love Xiangshan", then according to the importance index, the formula is:

[0098]

[0099] Among them, D(s k ) is the information density, I(w i ) is the word w i The amount of information, P is the sentence S K List of all words in L(S K ) is a sentence S K The sentence length, S K for sentences;

[0100] The information density of each candidate compressed sentence is calculated, wherein the size of the information density is the basis for obtaining the final compressed sentence.

[0101] Step S90: taking the candidate compressed sentence with the largest information density as the final compressed sentence.

[0102] The purpose of retaining the candidate compressed sentences with the highest information density is to achieve compression of the sentences while ensuring the important information of the data to the greatest extent.

[0103] First, the sentences are classified into long and short categories and assembled into test sentences. If the test sentence is greater than or equal to the minimum compression byte length, the test sentence is segmented. The syntax tree is used in the compression process. In order to ensure that short sentences with more important information are retained, the information density of each candidate compression sentence is innovatively calculated through the information density formula. The information density is used as the basis for obtaining the final compressed short sentence, which can improve the accuracy of compression.

[0104] Reference Figure 3 , Figure 3 for Figure 2 A detailed flow chart of an embodiment of step S50 in FIG. 1 is a flow chart of a detailed flow chart of an embodiment of step S50 in FIG. 1. In this embodiment, step S50 includes the following steps:

[0105] Step S501, calling a library function to perform dependency syntactic analysis and semantic dependency analysis on the word segmentation sentence to obtain an analysis result;

[0106] In this embodiment, semantic dependency parsing (SDP) analyzes the semantic associations between the various language units of a sentence and presents the semantic associations in a dependency structure. The advantage of using semantic dependency to depict the semantics of a sentence is that there is no need to abstract the vocabulary itself, but to describe the vocabulary through the semantic framework that the vocabulary bears, and the number of arguments is always much less than that of the vocabulary. The goal of semantic dependency analysis is to transcend the constraints of the surface syntactic structure of the sentence and directly obtain deep semantic information. For example, the same semantic information is expressed in different ways, that is, Zhang San performs an action of eating, and the action of eating is performed on apples. Dependency syntactic parsing is one of the key technologies in natural language processing. It is a processing process for analyzing the input text sentence to obtain the syntactic structure of the sentence. Analyzing the syntactic structure is, on the one hand, a demand of language understanding itself, and syntactic analysis is an important part of language understanding, and on the other hand, it also provides support for other natural language processing tasks. For example, syntax-driven statistical machine translation requires syntactic analysis of the source language or the target language (or both languages ​​at the same time); semantic analysis usually uses the output of syntactic analysis as input to obtain more indicative information. Dependency grammar analysis reveals the syntactic structure by analyzing the dependency relationship between components within a language unit, such as the "subject-verb-object" structure; semantic dependency analysis analyzes the semantic associations between the various language units in a sentence and presents the semantic associations in a dependency structure, such as "the cat ate the fish", "the fish was eaten by the cat", and "the cat ate the fish".

[0107] Step S502: construct at least one syntax tree according to the analysis result, wherein the syntax tree includes at least one candidate compressed sentence.

[0108] In this embodiment, the syntax tree has a multi-layer structure, and each layer has a corresponding candidate compressed sentence.

[0109] The following is a detailed process of calling the library function to perform dependency syntactic analysis and semantic dependency analysis on the word segmentation sentence to obtain the analysis result:

[0110] Calling a library function to annotate the structure and semantics of the segmented sentence respectively to obtain an annotation result;

[0111] Traversing a preset annotation set to determine whether the annotation result exists in the annotation set;

[0112] If the annotation result exists in the annotation set, the analysis result is obtained.

[0113] In this embodiment, the structure and semantics of the participle sentence need to be annotated respectively. The structure of the sentence can be subject, predicate, and object. The participle sentence is traversed.

[0114] The following is a refinement process of constructing at least one syntax tree according to the analysis result, wherein the syntax tree includes at least one candidate compressed sentence:

[0115] Segmenting the analysis results by using a segmentation tool HanLP to obtain candidate compressed sentences;

[0116] The candidate compressed sentence is cached into a preset initial syntax tree to obtain at least one syntax tree, wherein the syntax tree includes at least one candidate compressed sentence.

[0117] In this embodiment, after the segmented sentences are marked, the segmentation tool HanLP needs to be used to perform segmentation according to the marked results, for example, the subject and the object are segmented and stored in the branch specified by the initial syntax tree for compression.

[0118] The following is a detailed process of calling the library function to perform dependency syntactic analysis and semantic dependency analysis on the word segmentation sentence to obtain the analysis result:

[0119] Calling a library function to annotate the structure and semantics of the segmented sentence respectively to obtain an annotation result;

[0120] Traversing a preset annotation set to determine whether the annotation result exists in the annotation set;

[0121] If the annotation result exists in the annotation set, the analysis result is obtained.

[0122] In this embodiment, the structure and semantics of the segmentation sentence can be annotated by calling a library function. After the annotation is completed, a pre-set annotation set is traversed to determine whether the annotation result exists in the annotation set. The annotation set includes annotation information of multiple pre-set segmentation sentences.

[0123] The following is a detailed process for calculating the information amount of the candidate compressed sentence:

[0124] The information content of the candidate compressed sentence is calculated by the following formula;

[0125]

[0126] Among them, W i For the words in the text, tf ij W i Frequency in the text, idf ij W i The inverse document frequency of , w is the additional weight of proper nouns or keywords,

[0127] In this embodiment, the information volume of the candidate compressed sentences is calculated. The following is a detailed process of calculating the information density of each candidate compressed sentence based on the importance index of the proper noun and the keyword in the sentence:

[0128] Based on the importance index of the proper noun and the keyword in the sentence, the information density of each candidate compressed sentence is calculated by the following formula:

[0129]

[0130] Among them, D(s k ) is the information density, I(w i ) is the word w i The amount of information, P is the sentence S K List of all words in L(S K ) is a sentence S K The sentence length, S K For sentences.

[0131] The following is a second embodiment, after calculating the information amount of the candidate compressed sentence by the following formula, further comprising:

[0132] Determine whether there is a request to obtain proper nouns and keywords;

[0133] If there is a request to obtain the proper noun and the keyword, the current usage frequencies of the proper noun and the keyword are updated.

[0134] In this embodiment, the higher the number of times a proper noun and a keyword are obtained, the more frequently the proper noun and the keyword are used, and the greater the amount of information contained in the sentence containing the proper noun and the keyword.

[0135] Reference Figure 4 , Figure 4 This is a functional module diagram of an embodiment of a sentence compression device of the present invention. In this embodiment, the sentence compression device includes:

[0136] A segmentation module 10, used to segment the long sentence to be compressed into at least two short sentences by using a separator;

[0137] A screening module 20, configured to screen out at least one key sentence from the at least two short sentences by using a preset text classification model, and to splice the at least one key sentence into a test sentence;

[0138] A length determination module 30, used to determine whether the byte length of the test statement is greater than or equal to the minimum compressed byte length;

[0139] A word segmentation module 40, configured to segment the test sentence to obtain a segmented sentence if the byte length of the test sentence is greater than or equal to the minimum compressed byte length;

[0140] A construction module 50 is used to generate a syntax tree corresponding to the segmented sentence by a preset strategy algorithm, wherein the syntax tree includes at least one candidate compressed sentence;

[0141] An information volume calculation module 60, used to calculate the information volume of the candidate compressed sentence;

[0142] An importance index determination module 70, for determining the importance index of proper nouns and keywords in the sentence based on the information amount;

[0143] The calculation module 80 is used to calculate the information density of each candidate compressed sentence based on the importance index of the proper noun and the keyword in the sentence:

[0144] The compressed sentence output module 90 is used to take the candidate compressed sentence with the largest information density as the final compressed sentence.

[0145] In this embodiment, through the module of this device, the sentences can be first classified into long and short sentences and assembled into test sentences. If the test sentence is greater than or equal to the minimum compression byte length, the test sentence is segmented. A syntax tree is used in the compression process. In order to ensure that short sentences with more important information are retained, the information density of each candidate compressed sentence is innovatively calculated through the information density formula. The information density is used as the basis for obtaining the final compressed short sentence, which can improve the accuracy of compression.

[0146] The present invention also provides a computer-readable storage medium.

[0147] In this embodiment, a statement compression program is stored on the computer-readable storage medium, and when the statement compression program is executed by the processor, the steps of the statement compression method described in any of the above embodiments are implemented.

[0148] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM) and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0149] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims. All equivalent structures or equivalent process changes made using the contents of the specification and drawings of the present invention, or directly or indirectly used in other related technical fields, are protected by the present invention.

Claims

1. A sentence compression method, characterized in that: The statement compression method comprises: Split the long sentence to be compressed into at least two short sentences using a separator; Filtering at least one key sentence from the at least two short sentences by using a preset text classification model, and splicing the at least one key sentence into a test sentence; Determining whether the byte length of the test statement is greater than or equal to the minimum compressed byte length; If the byte length of the test sentence is greater than or equal to the minimum compressed byte length, segmenting the test sentence to obtain a segmented sentence; Generate a syntax tree corresponding to the word segmentation sentence by a preset strategy algorithm, wherein the syntax tree includes at least one candidate compression sentence; Calculating the information volume of the candidate compressed sentences; determining importance indexes of proper nouns and keywords in the sentence based on the information amount; Calculating the information density of each candidate compressed sentence based on the importance index of the proper noun and the keyword in the sentence; Taking the candidate compressed sentence with the largest information density as the final compressed sentence; The calculating the information amount of the candidate compressed sentence comprises: The information content of the candidate compressed sentence is calculated by the following formula; in, For the words of the text, for The frequency in the text, for The inverse document frequency of Additional weight for proper nouns or keywords; The calculating the information density of each candidate compressed sentence based on the importance index of the proper noun and the keyword in the sentence comprises: Based on the importance index of the proper noun and the keyword in the sentence, the information density of each candidate compressed sentence is calculated by the following formula: in, is the information density, For words The amount of information, P is the sentence List of all words in For Sentences The length of the sentence, For sentences.

2. The sentence compression method according to claim 1, characterized in that: The word segmentation sentence is generated into a syntax tree by a preset strategy algorithm, wherein the syntax tree includes at least one candidate compression sentence including: Calling a library function to perform dependency syntactic analysis and semantic dependency analysis on the segmented sentence to obtain an analysis result; At least one syntax tree is constructed according to the analysis result, wherein the syntax tree includes at least one candidate compression sentence.

3. The sentence compression method according to claim 2, characterized in that: The calling library function performs dependency syntactic analysis and semantic dependency analysis on the word segmentation sentence, and the analysis results obtained include: Calling a library function to annotate the structure and semantics of the segmented sentence respectively to obtain an annotation result; Traversing a preset annotation set to determine whether the annotation result exists in the annotation set; If the annotation result exists in the annotation set, the analysis result is obtained.

4. The sentence compression method according to claim 2, characterized in that: The step of constructing at least one syntax tree according to the analysis result, wherein the syntax tree includes at least one candidate compression sentence including: Segmenting the analysis results by using a segmentation tool HanLP to obtain candidate compressed sentences; The candidate compressed sentence is cached into a preset initial syntax tree to obtain at least one syntax tree, wherein the syntax tree includes at least one candidate compressed sentence.

5. The sentence compression method according to claim 1, characterized in that: After calculating the information amount of the candidate compressed sentence, the method further includes: Determine whether there is a request to obtain proper nouns and keywords; If there is a request to obtain the proper noun and the keyword, the current usage frequencies of the proper noun and the keyword are updated.

6. A sentence compression device, characterized in that: The sentence compression device comprises the following modules: A segmentation module, used for segmenting the long sentence to be compressed into at least two short sentences by using a separator; A screening module, configured to screen out at least one key sentence from the at least two short sentences by using a preset text classification model, and to splice the at least one key sentence into a test sentence; A length determination module, used to determine whether the byte length of the test statement is greater than or equal to the minimum compressed byte length; A word segmentation module, configured to segment the test sentence to obtain a segmented sentence if the byte length of the test sentence is greater than or equal to the minimum compressed byte length; A construction module, used to generate a syntax tree corresponding to the segmented sentence by a preset strategy algorithm, wherein the syntax tree includes at least one candidate compressed sentence; An information volume calculation module, used for calculating the information volume of the candidate compressed sentence; An importance index determination module, used to determine the importance index of proper nouns and keywords in the sentence based on the information amount; A calculation module is used to calculate the information density of each candidate compressed sentence based on the importance index of the proper noun and the keyword in the sentence: A compressed sentence output module, used for taking the candidate compressed sentence with the largest information density as the final compressed sentence; The calculating the information amount of the candidate compressed sentence comprises: The information content of the candidate compressed sentence is calculated by the following formula; in, For the words of the text, for The frequency in the text, for The inverse document frequency of Additional weight for proper nouns or keywords; The calculating the information density of each candidate compressed sentence based on the importance index of the proper noun and the keyword in the sentence comprises: Based on the importance index of the proper noun and the keyword in the sentence, the information density of each candidate compressed sentence is calculated by the following formula: in, is the information density, For words The amount of information, P is the sentence List of all words in For Sentences The length of the sentence, For sentences.

7. A sentence compression device, characterized in that: The statement compression device includes a memory, a processor, and a statement compression program stored in the memory and executable on the processor. When the statement compression program is executed by the processor, the steps of the statement compression method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a statement compression program, which, when executed by a processor, implements the steps of the statement compression method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-document abstract generation method and device, and terminal

    CN108959312A

  • A method and a device for extracting key information in corpus

    CN109582968A