Document compression method and device, electronic equipment and computer readable storage medium
By extracting keywords and building scoring candidate paragraphs in the information retrieval system, the redundant content problem is solved, efficient compression and accurate output of documents are achieved, and the document processing capability of the information retrieval system is improved.
Patent Information
- Application Number
- CN202510466682.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing information retrieval and document processing systems, it is difficult for users to obtain core content from massive data, and the existing document compression technology has redundant content and structural defects, resulting in poor information overload and compression effects.
By extracting keywords from the initial document, a candidate paragraph containing keywords is constructed, and the compressed document is determined based on paragraph scores, removing redundant content with low correlation with core content.
Improves the accuracy and simplicity of documents, enhances the document processing and compression quality of the search system, and ensures that the output document is logically coherent and relevant to query requirements.
Smart Images

Figure CN120373265A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information processing, and in particular, to a document compression method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] In current information retrieval and document processing systems, users usually need to obtain key content from a vast amount of data. However, existing retrieval techniques often return a large number of redundant documents, which may contain duplicate, irrelevant, or low-quality content. This information redundancy not only causes information overload for users but also affects the subsequent document processing and content compression effects. Summary of the Invention
[0003] In view of this, the purpose of the embodiments of the present application is to provide a document compression method, apparatus, electronic device, and computer-readable storage medium, which can improve the quality of document processing and compression.
[0004] In a first aspect, the embodiments of the present application provide a document compression method, including: extracting keywords from an initial document; constructing candidate paragraphs containing the keywords according to the keywords; and determining one or more paragraphs in the candidate paragraphs as the compressed document according to the paragraph scores of each paragraph in the candidate paragraphs.
[0005] In the above implementation process, after determining the initial document, by extracting the keywords in the initial document, constructing candidate paragraphs where the keywords are located, performing paragraph scoring on the candidate paragraphs, and determining the candidate paragraphs for constructing the compressed document based on the paragraph scoring results. Since the candidate paragraphs for constructing the compressed document are constructed based on the keywords and determined according to the paragraph scores. Therefore, the candidate paragraphs for constructing the compressed document can accurately express the core content of the initial document, improving the accuracy of the document. In addition, some redundant content with low relevance to the core content is removed from the compressed document, improving the conciseness of the document, and thus improving the quality of document processing and compression in the retrieval system.
[0006] In an embodiment, the constructing candidate paragraphs containing the keywords according to the keywords includes: splitting the initial document to obtain one or more sub-documents; where the sub-documents include: keyword sub-documents containing keywords and non-keyword sub-documents not containing keywords; generating a plurality of document combinations containing the keyword sub-documents based on a preset rule; and each document combination is a candidate paragraph.
[0007] In the above implementation process, by splitting the initial document into one or more sub-documents, the processing difficulty of the initial document can be simplified, and the processing and compression efficiency of the initial document can be improved. In addition, by setting that each candidate paragraph includes keywords, the accuracy and effectiveness of the content in the candidate paragraph can be improved.
[0008] In one embodiment, generating a plurality of document combinations including the keyword sub-documents based on a preset rule includes: for each of the keyword sub-documents, starting from the keyword sub-document, determining one or more of the sub-documents located in front of and / or behind the keyword sub-document as one or more target sub-documents; wherein, the target sub-documents include other keyword sub-documents and / or the non-keyword sub-documents; and for each time one or more of the obtained target sub-documents, according to the order of the target sub-documents and the keyword sub-documents in the initial document, sequentially combining one or more of the target sub-documents and the keyword sub-documents to obtain a plurality of document combinations including the keyword sub-documents.
[0009] In the above implementation process, when determining the document combination, determining that each document combination includes keywords can improve the accuracy and effectiveness of the content in the document combination. In addition, when generating the document combination, sequentially combining according to the order of the keyword sub-document and the target sub-document in the initial document can improve the continuity of the content and logic in the document combination.
[0010] In one embodiment, the sub-documents in each of the document combinations are consecutive sub-documents in the initial document.
[0011] In the above implementation process, by setting that the sub-documents in the document combination are consecutive sub-documents in the initial document, discontinuous sub-documents in the document combination can be avoided, thereby affecting the logical coherence, and the coherence of the content and logic of the combined document can be improved.
[0012] In one embodiment, determining one or more paragraphs in the candidate paragraphs as the compressed document according to the paragraph scores of the respective paragraphs in the candidate paragraphs includes: calculating the paragraph score of each of the candidate paragraphs; determining the candidate paragraph corresponding to the highest paragraph score as the compressed document; wherein, in the case where the paragraph scores corresponding to a plurality of the candidate paragraphs are all the highest paragraph scores, determining the candidate paragraph corresponding to the highest paragraph score with a shorter length as the compressed document.
[0013] In the above implementation process, the candidate paragraph corresponding to the highest paragraph score is usually the paragraph with relatively high comprehensive performance such as logical coherence, accuracy, and conciseness. It can accurately and concisely express the content in the initial document. Therefore, determining the candidate paragraph corresponding to the highest paragraph score as the compressed document can improve the accuracy and conciseness of the compressed document.
[0014] In one embodiment, calculating the paragraph score of each candidate paragraph includes: converting the candidate paragraph and the corresponding query requirement into a set format and inputting them into a scoring model; wherein the set format includes: a prompt word for judging the relationship between the candidate paragraph and the query requirement; determining the probability corresponding to the set prompt word output by the scoring model as the paragraph score of the candidate paragraph.
[0015] In the above implementation process, by converting the candidate paragraph and the corresponding query requirement into a set format and inputting them into the scoring model, and calculating the paragraph score of the candidate paragraph through the scoring model, it can be automatically calculated based on the actual content of the candidate paragraph and the set algorithm, which can quantify the relationship between the candidate paragraph and the initial document and improve the accuracy of candidate paragraph selection. In addition, when the output is a set prompt word, determining the probability corresponding to the set prompt word as the paragraph score of the candidate paragraph can obtain the candidate paragraph corresponding to the query requirement and improve the relevance between the compressed document and the query requirement.
[0016] In one embodiment, after constructing the candidate paragraphs containing the keywords according to the keywords, the method further includes: checking whether the candidate paragraphs contain the keywords; and / or removing duplicate sub-documents in the candidate paragraphs.
[0017] In the above implementation process, after constructing the candidate paragraphs, further detecting whether the candidate paragraphs contain keywords can improve the accuracy of the content in the candidate paragraphs. In addition, removing duplicate sub-documents in the candidate paragraphs can reduce the redundant information in the candidate paragraphs and improve the conciseness of the candidate paragraphs.
[0018] In a second aspect, an embodiment of the present application further provides a document compression device, including: an extraction module for extracting keywords from an initial document; a construction module for constructing candidate paragraphs containing the keywords according to the keywords; and a determination module for determining one or more paragraphs in the candidate paragraphs as the compressed document according to the paragraph scores of each paragraph in the candidate paragraphs.
[0019] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor and a memory, the memory stores machine-readable instructions executable by the processor, and when the electronic device runs, the machine-readable instructions are executed by the processor to execute the method steps in the above first aspect or any possible implementation manner of the first aspect.
[0020] Fourthly, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the document compression method in the first aspect or any possible implementation manner of the first aspect.
[0021] To make the above objects, features, and advantages of the present application more obvious and understandable, specific embodiments are hereinafter given, and detailed descriptions are made in conjunction with the accompanying drawings as follows. Description of the Drawings
[0022] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a block diagram of the electronic device provided by the embodiment of the present application;
[0024] Figure 2 It is a flowchart of the document compression method provided by the embodiment of the present application;
[0025] Figure 3 It is a schematic diagram of the functional modules of the document compression device provided by the embodiment of the present application. Detailed Embodiments
[0026] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0027] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0028] In a retrieval system (such as, a RAG (Retrieval-Augmented Generation) system), there are dual challenges of core information extraction and logic preservation when generating documents. On the one hand, the documents or text fragments retrieved by the retrieval system often contain a large amount of redundant content, such as background descriptions, repetitive arguments, or irrelevant details. These information noises significantly increase the difficulty of generating concise answers, making it difficult to accurately locate the core semantic units, resulting in missing or redundant residues in the generated results.
[0029] On the other hand, existing document compression technologies have structural defects. The mainstream methods perform sentence filtering through single-sentence level scoring (such as TF-IDF-based, sentence embedding similarity, or pre-trained model scoring). This local optimization strategy ignores cross-sentence logical associations. The compressed text often suffers from paragraph topic drift and broken argument chains due to the loss of key connecting sentences, and ultimately the generated answers are affected in both factual correctness and expression fluency.
[0030] In view of this, the present application proposes a document compression method. After determining the initial document, by extracting keywords in the initial document, constructing candidate paragraphs where the keywords are located, performing paragraph scoring on the candidate paragraphs, and determining the candidate paragraphs for constructing the compressed document based on the paragraph scoring results. Since the candidate paragraphs for constructing the compressed document are constructed based on keywords and determined according to paragraph scoring. Therefore, the candidate paragraphs for constructing the compressed document can accurately express the core content of the initial document, improving the accuracy of the document. In addition, some redundant content with low relevance to the core content is removed from the compressed document, improving the conciseness of the document, and thus improving the document processing and compression quality of the retrieval system.
[0031] To facilitate the understanding of this embodiment, first, the electronic device for executing the document compression method disclosed in the embodiments of the present application will be introduced in detail.
[0032] As Figure 1 shown, it is a block diagram of the electronic device. The electronic device 100 may include a memory 111 and a processor 113. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the electronic device 100. For example, the electronic device 100 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0033] The above-mentioned memory 111 and processor 113 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components may be electrically connected to each other through one or more communication buses or signal lines. The above-mentioned processor 113 is used to execute the executable module stored in the memory.
[0034] Among them, the memory 111 can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electric Erasable Programmable Read-Only Memory (EEPROM), etc. Among them, the memory 111 is used to store a program. After receiving an execution instruction, the processor 113 executes the program. The method executed by the electronic device 100 defined by the process disclosed in any embodiment of the embodiments of the present application can be applied to or implemented by the processor 113.
[0035] The above-mentioned processor 113 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 113 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a digital signal processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0036] The electronic device 100 in this embodiment can be used to execute each step in the various methods provided by the embodiments of the present application. The implementation process of the document compression method will be described in detail through several embodiments below.
[0037] Please refer to Figure 2 , which is a flowchart of the document compression method provided by the embodiments of the present application. The following will elaborate on the specific process shown in Figure 2 in detail.
[0038] Step 201, extract keywords from the initial document.
[0039] Among them, the initial document is a document initially generated by the retrieval system according to external access information.
[0040] A retrieval system is a technical system specifically designed to search for, filter, and return relevant information from a large-scale data collection. Its core objective is to quickly and accurately find the most relevant information according to the user's query requirements. After the retrieval system obtains the user's query requirements, it searches a large amount of data and generates corresponding initial documents based on the collected data. However, these initial documents usually contain lengthy information and it is difficult to achieve a concise and accurate answer.
[0041] To improve the accuracy and conciseness of the documents output by the retrieval system, after generating the initial documents, the initial documents can be compressed and then output to the user.
[0042] The keyword here refers to the words or phrases that express the core content, have a high degree of generalization and representativeness in the initial document, and are used to describe the core content and important information of the initial document. Among them, the initial document can include one or more keywords.
[0043] In one embodiment, the keyword can be extracted through a pre-trained BERT (Bidirectional Encoder Representations from Transformers) question-answering model. Among them, the BERT question-answering model is a deep bidirectional language model based on the Transformer architecture.
[0044] For ease of understanding, taking the keyword can be extracted through a pre-trained BERT question-answering model as an example, the specific implementation process of step 201 is shown as follows:
[0045] Convert the query requirements and the initial document into an input format and input them into the BERT question-answering model. The BERT question-answering model calculates the start probability and end probability of each character in the initial document and determines the keyword according to the start probability and end probability.
[0046] The input format here refers to the format that can be recognized by the BERT question-answering model. For example, [CLS] query requirements [SEP] text block [SEP].
[0047] The calculation formula for the above start probability is:
[0048] P start (i) = softmax(W s H i + b s );
[0049] The calculation formula for the end probability is:
[0050] P end (j) = softmax(We H j +b e );
[0051] Among them, P start (i) is the starting probability, H i is the hidden layer output of the i-th character, W s is the weight of the starting probability in the Q&A model, b s is the bias term of the starting probability in the Q&A model, P end (j) is the ending probability, H j is the hidden layer output of the j-th character, W e is the weight of the ending probability in the Q&A model, b e is the bias term of the ending probability in the Q&A model.
[0052] After calculating the two probabilities of the starting probability and the ending probability of each character, determine the position with the highest starting probability and the position with the highest ending probability behind the position with the highest starting probability. Among them, each character includes two probability values of the starting probability and the ending probability.
[0053] Determine the character at the position with the highest starting probability as the starting word, and the character at the position with the highest ending probability behind the position with the highest starting probability as the ending word. Starting from the starting word and ending with the ending word, determine that the phrase composed of the characters between the starting word and the ending word is the keyword.
[0054] Exemplarily, if the query requirement input to the retrieval system is: The third-quarter report in 2020 pointed out how China Unicom responded to the improvement of the broadband networking and speed increase query requirements caused by personnel management?
[0055] The keyword extracted from the initial document can be: Adhere to rational and standardized competition.
[0056] Step 202, according to the keyword, construct a candidate paragraph containing the keyword.
[0057] It should be understood that since the keyword is a word or phrase used to describe the core content and important information of the initial document, it is usually difficult to represent the complete overall logic of the initial document. Therefore, after extracting the keyword, one or more sentences or paragraphs where the keyword is located can be determined according to the keyword to form a candidate paragraph.
[0058] The candidate paragraph here is the candidate paragraph to be output to the customer. In order to further improve the accuracy and conciseness of the output document, the candidate paragraph can also be further screened to reduce the output of redundant content.
[0059] Optionally, one keyword can construct one or more candidate paragraphs.
[0060] In one embodiment, candidate paragraphs are formed based on keywords and the phrases or sentences before and after the keywords.
[0061] Step 203, based on the paragraph scores of each paragraph in the candidate paragraphs, determine one or more paragraphs in the candidate paragraphs as the compressed document.
[0062] The paragraph scores here can be scored from multiple dimensions such as relevance, accuracy, integrity, and logical coherence. Among them, the candidate paragraphs can determine the paragraph scores by calling a re-ranking callback function.
[0063] Optionally, the candidate paragraphs used to form the compressed document can be determined according to actual query requirements. For example, determine according to the high and low of the paragraph scores, determine according to the relationship between the paragraph scores and the score threshold, etc.
[0064] Exemplarily, if determined according to the high and low of the paragraph scores, after calculating the paragraph scores of each candidate paragraph, sort the paragraph scores, and determine the candidate paragraph with the highest paragraph score as the compressed document.
[0065] If determined according to the high and low of the paragraph scores, after calculating the paragraph scores of each candidate paragraph, sort the paragraph scores, and determine the multiple candidate paragraphs with the top-ranked paragraph scores as the compressed document.
[0066] If determined according to the score threshold, after calculating the paragraph scores of each candidate paragraph, match the paragraph scores that meet the score threshold, and determine the candidate paragraphs with the paragraph scores that meet the score threshold as the compressed document.
[0067] The above-mentioned determination method of the compressed document is only exemplary, and the determination method of the compressed document can be adjusted according to the actual situation.
[0068] In the above implementation process, after determining the initial document, by extracting the keywords in the initial document, constructing the candidate paragraphs where the keywords are located, scoring the candidate paragraphs, and determining the candidate paragraphs for constructing the compressed document based on the paragraph score results. Since the candidate paragraphs for constructing the compressed document are constructed based on keywords and determined according to the paragraph scores. Therefore, the candidate paragraphs for constructing the compressed document can accurately express the core content of the initial document, improving the accuracy of the document. In addition, the compressed document removes some redundant content with low relevance to the core content, improving the conciseness of the document, and thus improving the document processing and compression quality of the retrieval system.
[0069] In a possible implementation manner, step 202 includes: splitting the initial document to obtain one or more sub-documents; generating multiple document combinations including the sub-documents with keywords based on preset rules.
[0070] Among them, the sub-documents include: keyword sub-documents containing keywords and non-keyword sub-documents not containing keywords.
[0071] Since the initial document is a long document generated by the retrieval system according to the user's query requirements, in order to reduce the processing difficulty of the initial document, the initial document can be first segmented into multiple smaller sub-documents. Among them, the sub-documents can be a sentence or a paragraph.
[0072] The initial document here can be segmented by a preset delimiter. The delimiter can be a punctuation mark, such as one or more of a full stop, a question mark, an exclamation mark, etc.
[0073] Exemplarily, if the delimiter is a "full stop", the initial document is: Thanks to the active and effective adjustment of the mobile business development strategy, the quality of the company's mobile business development has been gradually improved, and the mobile main business income has further increased by 2.5% year-on-year in the third quarter. In terms of the fixed-line broadband business, the personnel management has significantly driven the broadband networking and speed-up query requirements. The company adheres to rational and standardized competition, actively gives play to the comprehensive advantages of high broadband speed, rich content, and excellent service, accelerates the promotion of the smart home series products, and promotes the common growth of the broadband access business and other related businesses. In the first three quarters of 2020, the net increase in fixed-line broadband users was 3.08 million, reaching 86.56 million; the fixed-line broadband access revenue reached 32.096 billion yuan, a year-on-year increase of 3.7%.
[0074] After segmenting the initial document, the obtained sub-documents may include:
[0075] Sub-document one: Thanks to the active and effective adjustment of the mobile business development strategy, the quality of the company's mobile business development has been gradually improved, and the mobile main business income has further increased by 2.5% year-on-year in the third quarter.
[0076] Sub-document two: In terms of the fixed-line broadband business, the personnel management has significantly driven the broadband networking and speed-up query requirements.
[0077] Sub-document three: The company adheres to rational and standardized competition, actively gives play to the comprehensive advantages of high broadband speed, rich content, and excellent service, accelerates the promotion of the smart home series products, and promotes the common growth of the broadband access business and other related businesses.
[0078] Sub-document four: In the first three quarters of 2020, the net increase in fixed-line broadband users was 3.08 million, reaching 86.56 million; the fixed-line broadband access revenue reached 32.096 billion yuan, a year-on-year increase of 3.7%.
[0079] It should be understood that since the keywords in the initial document have been extracted in step 201, after the initial document is divided into multiple sub-documents, it is possible to classify the sub-documents by determining whether the keywords exist in each sub-document. Among them, the sub-documents containing keywords are keyword sub-documents, and the sub-documents not containing keywords are non-keyword sub-documents.
[0080] Continuing with the above example, if the keyword extracted from the initial document is "rational and regulated competition", then sub-document three is a keyword sub-document, and sub-documents one, two, and four are non-keyword sub-documents.
[0081] It should be understood that since the initial document is divided into multiple sub-documents, if the keyword sub-documents are directly output, it may lead to illogical coherence and be difficult for users to understand. Especially in the case of fine segmentation, in extreme cases, the keyword sub-document may be just a single sentence, which is difficult to convey specific meanings. If the keyword sub-documents are directly output to users, users may have difficulty understanding.
[0082] Based on this, after determining the keyword sub-documents, it is also possible to generate a continuous document combination according to a preset rule and based on the keyword sub-documents, combined with other sub-documents around the keyword sub-documents, so as to make the output document logically coherent and easy to understand.
[0083] The preset rule here refers to the rule set in advance for generating a document combination containing keyword sub-documents. For example, the preset rule is: generate a document combination according to 2 sub-documents around the keyword sub-document and the keyword sub-document. Another example is that the preset rule is: starting from the keyword sub-document, generate a document combination with 1 sub-document in front of the keyword sub-document and the keyword sub-document. Another example is that the preset rule is: starting from the keyword sub-document, generate a document combination with 1 sub-document in front of the keyword sub-document and the keyword sub-document; and generate a document combination according to 2 sub-documents around the keyword sub-document, etc. The preset rule can be adjusted according to the actual situation.
[0084] Among them, each keyword sub-document can form one or more document combinations, and each document combination is a candidate paragraph.
[0085] In the above implementation process, by dividing the initial document into one or more sub-documents, the processing difficulty of the initial document can be simplified, and the processing and compression efficiency of the initial document can be improved. In addition, by setting that each candidate paragraph includes keywords, the accuracy and effectiveness of the content in the candidate paragraph can be improved.
[0086] In a possible implementation, based on a preset rule, multiple document combinations including keyword sub-documents are generated, including: for each keyword sub-document, starting from the keyword sub-document, determining one or more sub-documents located before and / or after the keyword sub-document as one or more target sub-documents; and for each time one or more target sub-documents are obtained, according to the order of the target sub-documents and the keyword sub-document in the initial document, combining one or more target sub-documents and the keyword sub-document in sequence to obtain multiple document combinations including the keyword sub-document.
[0087] Among them, the target sub-documents include other keyword sub-documents and / or non-keyword sub-documents.
[0088] The determination of the target sub-documents here includes the following methods:
[0089] Method 1: Starting from the keyword sub-document, determining one or more sub-documents located before the keyword sub-document as one or more target sub-documents.
[0090] Method 2: Starting from the keyword sub-document, determining one or more sub-documents located after the keyword sub-document as one or more target sub-documents.
[0091] Method 3: Starting from the keyword sub-document, determining one or more sub-documents located before and one or more sub-documents located after the keyword sub-document as one or more target sub-documents.
[0092] For each keyword sub-document, one or more target sub-documents can be determined according to one or more of the above Method 1, Method 2, and Method 3.
[0093] Optionally, after determining one or more target sub-documents, the target sub-documents and the keyword sub-document can be arranged and combined to obtain one or more document combinations. Among them, when arranging and combining the target sub-documents and the keyword sub-document, they are combined according to the order of the target sub-documents and the keyword sub-document in the initial document.
[0094] Continuing with the example in the above embodiment, in the case where sub-document three is the keyword sub-document, the determination method of the target sub-documents of sub-document three:
[0095] If the preset rule is: starting from the keyword sub-document, determining one sub-document located before the keyword sub-document and one sub-document located after the keyword sub-document as the target sub-documents. Then the target sub-documents can include: sub-document two and sub-document four.
[0096] Since the order of Sub-document Two, Sub-document Four, and Sub-document Three in the initial document is: Document Two, Sub-document Three, and Sub-document Four, the document combinations generated according to the target sub-documents include: [Document Two, Sub-document Three, Sub-document Four], [Document Two, Sub-document Three], [Sub-document Three, Sub-document Four].
[0097] Continuing with the example in the above embodiment, in the case where Sub-document Three is the keyword sub-document, the method for determining the target sub-document of Sub-document Three:
[0098] If the preset rule is: starting from the keyword sub-document, determine the two sub-documents located before the keyword sub-document and the one sub-document located after the keyword sub-document as the target sub-documents. Then the target sub-documents may include: Sub-document One, Sub-document Two, and Sub-document Four.
[0099] Since the order of Sub-document One, Sub-document Two, Sub-document Four, and Sub-document Three in the initial document is: Sub-document One, Document Two, Sub-document Three, and Sub-document Four, the document combinations generated according to the target sub-documents include: [Sub-document One, Document Two, Sub-document Three, Sub-document Four], [Sub-document One, Document Two, Sub-document Three], [Sub-document Two, Sub-document Three, Sub-document Four].
[0100] The above methods for determining the target sub-documents and document combinations are only exemplary, and the method for determining the target sub-documents can be selected according to the actual situation.
[0101] It should be understood that in the case where multiple keywords are extracted from the initial document, for each keyword, the corresponding document combination can be obtained in the above manner.
[0102] In one embodiment, in order to improve the overall coherence and logic of the compressed document. In the case where the initial document includes multiple keywords, after respectively determining the document combinations corresponding to each keyword sub-document, it is further possible to combine the document combinations corresponding to two adjacent keyword sub-documents to form an overall document combination. Among them, when combining the document combinations corresponding to two adjacent keyword sub-documents, all the sub-documents between these two keyword sub-documents can be regarded as sub-documents.
[0103] Exemplarily, if after the initial document is segmented, Sub-document A, Sub-document B, Sub-document C, Sub-document D, Sub-document E, Sub-document F, Sub-document G, and Sub-document H are obtained. Among them, Sub-document C and Sub-document G are keyword sub-documents, and the other sub-documents are non-keyword sub-documents.
[0104] If the document combination corresponding to Sub-document C includes: [Sub-document B, Sub-document C, Sub-document D], [Sub-document C, Sub-document D], [Sub-document B, Sub-document C].
[0105] The document combinations corresponding to sub-document G include: [sub-document F, sub-document G, sub-document H], [sub-document F, sub-document G], [sub-document G, sub-document H].
[0106] When further combining the corresponding document combinations of sub-document C and sub-document G, sub-document D, sub-document E, and sub-document F can all be regarded as target sub-documents. The overall document combinations formed include: [sub-document B, sub-document C, sub-document D, sub-document E, sub-document F, sub-document G, sub-document H], [sub-document B, sub-document C, sub-document D, sub-document E, sub-document F, sub-document G], [sub-document C, sub-document D, sub-document E, sub-document F, sub-document G].
[0107] In the above implementation process, when determining the document combinations, determining that each document combination includes keywords can improve the accuracy and effectiveness of the content in the document combinations. Additionally, when generating the document combinations, combining them in the order of the keyword sub-documents and target sub-documents in the initial document can improve the continuity of the content and logic in the document combinations.
[0108] In a possible implementation manner, among them, the sub-documents in each document combination are consecutive sub-documents in the initial document.
[0109] The consecutive sub-documents here refer to the sub-documents that are consecutive in the initial document.
[0110] For example, after the initial document is segmented, we get: sub-document A, sub-document B, sub-document C, sub-document D, sub-document E. Sub-document B is the keyword sub-document, and the order of these sub-documents in the initial document is: sub-document A, sub-document B, sub-document C, sub-document D, sub-document E. Then the document combinations can be: [sub-document A, sub-document B, sub-document C, sub-document D, sub-document E], [sub-document A, sub-document B, sub-document C, sub-document D], [sub-document B, sub-document C, sub-document D, sub-document E], [sub-document B, sub-document C, sub-document D], [sub-document B, sub-document C].
[0111] Among them, combinations such as [sub-document B, sub-document D, sub-document E], [sub-document B, sub-document C, sub-document E], [sub-document B, sub-document E] cannot be used as the document combinations corresponding to sub-document B because they have discontinuous paragraphs.
[0112] In the above implementation process, by setting the sub-documents in the document combinations to be consecutive sub-documents in the initial document, it is possible to avoid discontinuous sub-documents in the document combinations, thereby affecting logical coherence and improving the coherence of the content and logic of the combined documents.
[0113] In a possible implementation, step 203 includes: calculating the passage score for each candidate passage; determining the candidate passage corresponding to the highest passage score as the compressed document.
[0114] The passage score here can be determined by means such as a scoring model, a scoring calculation formula, etc. For each candidate passage, its corresponding passage score is calculated.
[0115] The passage score here can be used to reflect the logical coherence, accuracy, conciseness, etc. of the candidate passage.
[0116] After calculating the passage score for each candidate passage, the candidate passage corresponding to the highest passage score can be determined as the compressed document. Among them, in the case where the passage scores corresponding to multiple candidate passages are all the highest passage score, the candidate passage corresponding to the highest passage score with a shorter length is determined as the compressed document.
[0117] In the above implementation process, the candidate passage corresponding to the highest passage score is usually a passage with relatively high comprehensive performance such as logical coherence, accuracy, and conciseness. It can accurately and concisely express the content in the initial document. Therefore, determining the candidate passage corresponding to the highest passage score as the compressed document can improve the accuracy and conciseness of the compressed document.
[0118] In a possible implementation, calculating the passage score for each candidate passage includes: converting the candidate passage and the corresponding query requirement into a set format and inputting it into the scoring model; determining the probability corresponding to the set prompt word output by the scoring model as the passage score of the candidate passage.
[0119] Among them, the set format includes: a prompt word for judging the relationship between the candidate passage and the query requirement. The set format can also include the user's query requirement, the initial document, etc. For example, the set format is: "A:[question]\nB:[passage]\n[prompt]".
[0120] The prompt word here is an instruction for judging whether the candidate passage is an answer to the user's query requirement. The set prompt word can be "Given a query requirement A and a passage B, determine whether the passage contains the answer to the query and provide a prediction of 'Yes' or 'No'", or the set prompt word can be "Given a query A and a passage B, determine whether the passage contains an answer to the query by providing a prediction of either 'Yes' or 'No'", etc. The set prompt word can be set according to the actual situation.
[0121] After the candidate paragraph and the query requirement are input into the scoring model, the scoring model evaluates the candidate paragraph and the query requirement, and uses the output of the last layer in the scoring model as the basis for scoring. The probability that the result corresponding to the generated prompt word of the output is "yes" or "Yes" is the paragraph score.
[0122] In one embodiment, when calculating the paragraph score, if there is a corresponding document combination, the paragraph scores of the same document combination will be cached to avoid repeated calculation. New documents are calculated preferentially, and the corresponding paragraph scores of the cached documents can be directly obtained.
[0123] In the above implementation process, by converting the candidate paragraph and the corresponding query requirement into a set format and inputting them into the scoring model, and calculating the paragraph score of the candidate paragraph through the scoring model, it can be automatically calculated based on the actual content of the candidate paragraph and the set algorithm, which can quantify the relationship between the candidate paragraph and the initial document and improve the accuracy of candidate paragraph selection. In addition, when the output is a set prompt word, determining the probability corresponding to the set prompt word as the paragraph score of the candidate paragraph can obtain the candidate paragraph corresponding to the query requirement and improve the relevance between the compressed document and the query requirement.
[0124] In a possible implementation manner, according to the keyword, after step 202, the method further includes: checking whether the candidate paragraph contains the keyword; and / or removing the duplicate sub-documents in the candidate paragraph.
[0125] It should be understood that during the task execution process, due to the occurrence of errors such as human error or algorithm error, some candidate paragraphs may not contain keywords. Therefore, after determining the candidate paragraphs, the candidate paragraphs can be further checked to determine whether the candidate paragraphs contain keywords. If the candidate paragraph contains the keyword, the candidate paragraph is retained. If the candidate paragraph does not contain the keyword, the candidate paragraph is removed.
[0126] In addition, when generating the combined document according to the target sub-document and the keyword sub-document, some duplicate combined documents may be formed. Therefore, after determining the candidate paragraphs, the candidate paragraphs can be filtered to remove the duplicate sub-documents in the candidate paragraphs and improve the conciseness of the candidate paragraphs.
[0127] Next, through a specific example, the document obtained by the document compression method in the embodiments of the present application is intuitively shown:
[0128] Exemplarily, if the query requirement input by the user is: The third quarter report in 2020 pointed out how China Unicom responded to the improvement of the broadband networking and speed increase query requirement caused by personnel management?
[0129] The initial retrieved document is as follows: Thanks to the positive and effective adjustment of the mobile business development strategy, the quality of the company's mobile business development has been gradually improved, and the main mobile business revenue has further increased by 2.5% year-on-year in the third quarter. In terms of fixed-line broadband services, the demand for broadband networking and speed-up query driven by personnel management has increased significantly. The company adheres to rational and standardized competition, gives full play to the comprehensive advantages of high broadband speed, rich content and excellent service, accelerates the promotion of smart home series products, and promotes the common growth of broadband access services and other related services. In the first three quarters of 2020, the net increase in fixed-line broadband users was 3.08 million, reaching 86.56 million; the fixed-line broadband access revenue reached 32.096 billion yuan, a year-on-year increase of 3.7%.
[0130] The extracted keywords are: adhere to rational and standardized competition.
[0131] The compressed document is as follows: In terms of fixed-line broadband services, the demand for broadband networking and speed-up query driven by personnel management has increased significantly. The company adheres to rational and standardized competition, gives full play to the comprehensive advantages of high broadband speed, rich content and excellent service, accelerates the promotion of smart home series products, and promotes the common growth of broadband access services and other related services.
[0132] From the above comparison, it can be seen that compared with the initial document, redundant information is removed and only the content related to the query demand is retained. In addition, since the compressed document contains keywords, the loss of key information can be avoided.
[0133] In the above implementation process, after constructing the candidate paragraph, further detecting whether the candidate paragraph contains keywords can improve the accuracy of the content in the candidate paragraph. In addition, the repeated sub-documents in the candidate paragraph are removed, which can reduce the redundant information in the candidate paragraph and improve the conciseness of the candidate paragraph.
[0134] Based on the same application concept, a document compression device corresponding to the document compression method is further provided in the embodiments of the present application. Since the principle of solving problems by the device in the embodiments of the present application is similar to that of the foregoing document compression method embodiments, the embodiments of the device in the present embodiment can refer to the description in the embodiments of the above method, and the repeated parts will not be elaborated.
[0135] Please refer to Figure 3 , which is a schematic diagram of the functional modules of the document compression device provided in the embodiments of the present application. Each module in the document compression device in the present embodiment is used to execute each step in the above method embodiment. The document compression device includes an extraction module 301, a construction module 302, and a determination module 303;
[0136] Among them,
[0137] The extraction module 301 is used to extract keywords from the initial document.
[0138] The building module 302 is configured to build candidate paragraphs containing the keyword according to the keyword.
[0139] The determining module 303 is configured to determine one or more paragraphs in the candidate paragraphs as the compressed document according to the paragraph scores of the paragraphs in the candidate paragraphs.
[0140] In a possible implementation, the building module 302 is further configured to: split the initial document to obtain one or more sub-documents; where the sub-documents include: keyword sub-documents containing keywords and non-keyword sub-documents not containing keywords; generate multiple document combinations containing the keyword sub-documents based on a preset rule; where each document combination is a candidate paragraph.
[0141] In a possible implementation, the building module 302 is specifically configured to: for each keyword sub-document, starting from the keyword sub-document, determine one or more of the sub-documents located in front of and / or behind the keyword sub-document as one or more target sub-documents; where the target sub-documents include other keyword sub-documents and / or the non-keyword sub-documents; and for each time one or more of the obtained target sub-documents are obtained, combine one or more of the target sub-documents and the keyword sub-documents in sequence according to the order of the target sub-documents and the keyword sub-documents in the initial document to obtain multiple document combinations containing the keyword sub-documents.
[0142] In a possible implementation, the determining module 303 is further configured to: calculate the paragraph score of each candidate paragraph; determine the candidate paragraph corresponding to the highest paragraph score as the compressed document; where in the case where the paragraph scores corresponding to multiple candidate paragraphs are all the highest paragraph score, determine the candidate paragraph corresponding to the highest paragraph score with a shorter length as the compressed document.
[0143] In a possible implementation, the determining module 303 is specifically configured to: convert the candidate paragraph and the corresponding query requirement into a set format and input them into a scoring model; where the set format includes: a prompt word for judging the relationship between the candidate paragraph and the query requirement; determine the probability corresponding to the set prompt word output by the scoring model as the paragraph score of the candidate paragraph.
[0144] In a possible implementation, the document compression device further includes a reprocessing module, configured to check whether the candidate paragraph contains the keyword; and / or remove duplicate sub-documents in the candidate paragraph.
[0145] In addition, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the document compression method described in the above method embodiment.
[0146] A computer program product of the document compression method provided by an embodiment of the present application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the steps of the document compression method described in the above method embodiment. For details, please refer to the above method embodiment and will not be elaborated here.
[0147] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0148] In addition, in each embodiment of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0149] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the said elements.
[0150] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0151] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or replacements, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A document compression method, characterized in that Comprising: Extracting keywords from an initial document; Constructing candidate paragraphs containing the keywords according to the keywords; Determining one or more paragraphs in the candidate paragraphs as the compressed document according to the paragraph scores of each paragraph in the candidate paragraphs.
2. The method according to claim 1, wherein The constructing candidate paragraphs containing the keywords according to the keywords includes: Segmenting the initial document to obtain one or more sub-documents; wherein, the sub-documents include: keyword sub-documents containing keywords and non-keyword sub-documents not containing keywords; Generating multiple document combinations containing the keyword sub-documents based on preset rules; Wherein, each document combination is a candidate paragraph.
3. The method according to claim 2, wherein The generating multiple document combinations containing the keyword sub-documents based on preset rules includes: For each keyword sub-document, starting from the keyword sub-document, determining one or more of the sub-documents located in front of and / or behind the keyword sub-document as one or more target sub-documents; wherein, the target sub-documents include other keyword sub-documents and / or the non-keyword sub-documents; and For each time one or more of the obtained target sub-documents are obtained, combining one or more of the target sub-documents and the keyword sub-documents in sequence according to the order of the target sub-documents and the keyword sub-documents in the initial document to obtain multiple document combinations containing the keyword sub-documents.
4. The method according to claim 3, characterized in that, Wherein, The sub-documents in each document combination are consecutive sub-documents in the initial document.
5. The method according to claim 1, characterized in that, Determining one or more paragraphs in the candidate paragraphs as the compressed document according to the paragraph scores of each paragraph in the candidate paragraphs includes: Calculating the paragraph score of each candidate paragraph; Determining the candidate paragraph corresponding to the highest paragraph score as the compressed document; Wherein, in the case where the paragraph scores corresponding to multiple candidate paragraphs are all the highest paragraph scores, determining the candidate paragraph corresponding to the highest paragraph score with a shorter length as the compressed document.
6. The method according to claim 5, characterized in that, The calculating the paragraph score of each candidate paragraph includes: Converting the candidate paragraph and the corresponding query requirement into a set format and inputting them into a scoring model; wherein, the set format includes: a prompt word for judging the relationship between the candidate paragraph and the query requirement; Determining the probability corresponding to the set prompt word output by the scoring model as the paragraph score of the candidate paragraph.
7. The method according to any one of claims 1-6, characterized in that, After the constructing candidate paragraphs containing the keywords according to the keywords, the method further includes: Checking whether the candidate paragraph contains the keywords; and / or Removing duplicate sub-documents in the candidate paragraph.
8. A document compression device, characterized in that, Comprising: An extraction module for extracting keywords from an initial document; A construction module for constructing candidate paragraphs containing the keywords according to the keywords; A determination module for determining one or more paragraphs in the candidate paragraphs as the compressed document according to the paragraph scores of each paragraph in the candidate paragraphs.
9. An electronic device, characterized in that, Comprising: A processor and a memory, the memory storing machine-readable instructions executable by the processor, and when the electronic device runs, the machine-readable instructions, when executed by the processor, perform the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program runs on the processor, it performs the steps of the method according to any one of claims 1 to 7.