Document summary extraction method, system, electronic device and storage medium

By performing optical character recognition and image type recognition of enterprise documents, and combining industry tag encoding, input into the pre-constructed abstract extraction model, the problem of difficulty in extracting key image data and industry background information in the existing technology is solved. The generated graphic and text abstracts have high targetedness and generalization capabilities, improving the efficiency of application scenarios.

CN119131829BActive Publication Date: 2025-05-16南京中孚信息技术有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411629311.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-05-16
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

Existing document summary technology is difficult to effectively extract key image data and industry background information from corporate documents, resulting in the lack of targeted and generalized capabilities of generated abstracts.

Method used

By parsing the document to be extracted into fragments, optical character recognition and image type recognition, preliminary semantic text and structural description text are generated, and combined with industry tag encoding, input into the pre-constructed abstract extraction model to generate a graphic and text document summary.

Benefits of technology

The key semantics and key illustrations in corporate documents are extracted. The generated abstract has an industry background and the effect of "one picture to offset a thousand words", which improves the targetedness and generalization capabilities of the abstract, and improves the efficiency of application scenarios such as machine-assisted data sensitivity analysis and judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119131829B_ABST
    Figure CN119131829B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a document summary extraction method, system, electronic device and storage medium, which belongs to the field of navigation technology. The method includes: parsing the document to be summarized into fragments to generate a fragment set; performing optical character recognition to form a preliminary semantic text, and determining a first token sequence; performing image type recognition on the illustrations in the fragment set to form a structural description text, and determining a second token sequence; identifying the industry or field label of the document to be summarized, and determining the corresponding code; inputting the first and second token sequences and the codes corresponding to the labels into a pre-built summary extraction model to obtain a summary text. The graphic document summary generation scheme based on a recurrent neural network utilizes key semantic extraction, document structure extraction, and document industry or field identification to extract key semantics and key illustrations from corporate documents to form a graphic document summary, which is targeted and has strong summarization capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of navigation technology, and in particular to a document summary extraction method, system, electronic equipment and storage medium. Background Art

[0002] Document summarization is an important technical field of data processing. In enterprise-level applications, it is often used in application scenarios such as document classification and grading, document asset sorting, data leakage analysis, and machine-assisted data sensitivity assessment within large enterprises or government regulatory departments.

[0003] The mainstream document summarization technology is based on various neural network language models. Although the language models based on various neural networks have achieved good results in plain text summarization and short text summarization in the public domain, the content summarization of long documents within the enterprise has not yet reached a satisfactory level:

[0004] (1) Enterprise documents are usually medium-length documents containing rich graphic and text data. A lot of high-value information is hidden in the pictures in the documents, such as key drawings, architecture diagrams, flow charts, key component assembly diagrams, etc. in technical documents. They are highly generalized and serve as a summary of the document, with the effect of "one picture is worth a thousand words". Existing document summarization technology focuses on key text extraction and content abstraction, lacks the extraction of key image data in the document, and the resulting document summary is in the form of pure text with a single expression form. It has the disadvantages of being unintuitive to read and some key information difficult to describe.

[0005] (2) Traditional document summarization technology lacks attention to the industry background of the document. Corporate documents are usually highly specialized documents in a specific field, and contain many professional terms and unknown words. Existing language models have good generalization capabilities for text summarization tasks, but they lack knowledge of the specific field of the enterprise, and thus lack the correct focus. The generated summaries often lack specificity, and the summary results lack focus and are general. Summaries without industry attributes often fail to achieve the semantic expression effect of "hitting the nail on the head" and "having substance".

[0006] (3) Traditional document summarization technology ignores the internal structural features of documents. Corporate documents usually have document titles and rich title hierarchies. This information is generally of high weight in the document and should have a significant impact on the document summary. Traditional document summarization technology only parses the document into plain text during the document preprocessing stage, which loses some of the key semantics contained in the title. Summary of the invention

[0007] The purpose of the embodiments of the present invention is to provide a document summary extraction method, system, electronic device and storage medium, which are used to fully or partially solve the technical problems existing in the above-mentioned prior art.

[0008] In order to achieve the above object, an embodiment of the present invention provides a method for extracting a document summary, comprising:

[0009] Parse the document to be summarized into fragments and generate a fragment set;

[0010] Performing optical character recognition on the illustrations in the fragment set to form a preliminary semantic text, and determining a first token sequence based on the preliminary semantic text;

[0011] Performing image type recognition on the illustrations in the fragment set to form a structure description text, and determining a second token sequence based on the structure description text;

[0012] Identify the industry or field labels of the documents to be summarized and determine the codes corresponding to the labels;

[0013] The first token sequence, the second token sequence, and the codes corresponding to the tags are input into a pre-built summary extraction model to obtain a summary text.

[0014] Optionally, each fragment in the fragment set has attributes of fragment type, document ID, fragment sequence number, title content, title level, text content, and image content, wherein the fragment type includes paragraph, illustration, and title types.

[0015] Optionally, performing optical character recognition on the illustrations in the fragment set to form a preliminary semantic text, and determining a first token sequence based on the preliminary semantic text includes:

[0016] Perform optical character recognition on the image content in the clip set, and wrap the optical character recognition results with preset keywords;

[0017] According to the fragment sequence number, the title, paragraph content and image optical character recognition results are merged to form a preliminary semantic text;

[0018] Using the TextRank algorithm, extract key sentences from the preliminary semantic text, wherein the key sentences include text wrapped with preset keywords;

[0019] The key sentences are segmented, word vectors are embedded, and truncated and padded to obtain the first token sequence.

[0020] Optionally, performing image type recognition on the illustrations in the fragment set to form a structure description text, and determining a second token sequence based on the structure description text includes:

[0021] Perform image type recognition on the illustrations in the clip set, and wrap them with preset keywords before and after recognition, and regard the wrapped image type as the structural description of the clip; image type recognition includes but is not limited to table, architecture diagram, flow chart, map, network topology diagram, wireframe diagram, and interface screenshot recognition;

[0022] The structural descriptions of all fragments are merged in the order of fragment numbers to form a structural description text, and the structural description text is segmented, converted into word embedding vectors, truncated and padded to obtain the second token sequence.

[0023] Optionally, the summary extraction model is designed based on the encoder-decoder framework of the Seq2Seq model, wherein a first, second and third encoders are used in the encoding stage, and a fourth decoder, a first attention reader unit, a second attention reader unit and a fully connected unit are used in the decoding stage.

[0024] Optionally, define the objective function of the summary extraction model as:

[0025] ;

[0026] In the formula, n represents the length of the sequence, X represents the first token sequence, D represents the second token sequence, T represents the encoding corresponding to the label, and log represents the logarithmic function. Indicates the summary content. Represents the summary content before time t , the first token sequence X, the second token sequence D, and the encoding T corresponding to the label, the next token is probability.

[0027] Optionally, inputting the first token sequence, the second token sequence, and the encoding corresponding to the label into a pre-built summary extraction model to obtain a summary text includes:

[0028] Converting the first token sequence into a first hidden representation using a first encoder via a bidirectional long short-term memory unit;

[0029] Using a second encoder to convert the second token sequence into a second hidden representation through a bidirectional long short-term memory unit;

[0030] The third encoder is used to convert the encoding corresponding to the label into a dense embedding vector;

[0031] Using the first attention reader unit, the state at time t-1 is used as the query vector, the first hidden representation is used as the key vector and the value vector for conversion, and the semantic context vector is output;

[0032] Using the second attention reading unit, the state at time t-1 is used as the query vector, and the second hidden representation is used as the key vector and value vector for conversion, and the structural context vector is output;

[0033] Use a fully connected unit to connect the semantic context vector, structural context vector, and dense embedding vector into one vector;

[0034] The fourth decoder is used to take the word vector, semantic context vector, structural context vector and label context vector output by the fully connected unit at time t-1 as input to update the state at time t.

[0035] On the other hand, the present invention also provides a document summary extraction system, comprising:

[0036] A document parsing unit, used for parsing the document to be extracted into fragments and generating a fragment set;

[0037] A semantic extraction unit, configured to perform optical character recognition on the illustrations in the fragment set to form a preliminary semantic text, and determine a first token sequence based on the preliminary semantic text;

[0038] A structure extraction unit, configured to perform image type recognition on the illustrations in the fragment set to form a structure description text, and determine a second token sequence based on the structure description text;

[0039] A label recognition unit, used to recognize the industry or field label of the document to be extracted and to determine the code corresponding to the label;

[0040] The summary inference unit is used to input the first token sequence, the second token sequence and the encoding corresponding to the label into a pre-built summary extraction model to obtain a summary text.

[0041] On the other hand, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the document summary extraction method described above when executing the program.

[0042] On the other hand, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the document summary extraction method described above are implemented.

[0043] Through the above technical solution, a graphic document summary generation solution based on recurrent neural network utilizes key semantic extraction, document structure extraction, document industry or field identification, and recurrent neural network technology with multiple encoders to extract key semantics and key illustrations from corporate documents to form a graphic document summary. The text generated by the summary has an industry or field background, and the illustrations can make the summary have the effect of "one picture is worth a thousand words". The generated summary is targeted and has strong summarization ability, which greatly improves the efficiency of application scenarios such as machine-assisted data sensitivity analysis.

[0044] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:

[0046] Figure 1 It is an implementation flow chart of a method for extracting a document summary provided by an embodiment of the present invention;

[0047] Figure 2 It is a structural schematic diagram of a document summary extraction model provided by an embodiment of the present invention;

[0048] Figure 3 It is a structural diagram of a document summary extraction system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The specific implementation of the embodiment of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the embodiment of the present invention, and is not used to limit the embodiment of the present invention.

[0050] See also Figure 1 As shown, it is a flowchart of an implementation method of extracting a document summary provided by an embodiment of the present invention, including the following execution steps:

[0051] Step 100: Parse the document to be summarized into segments and generate a segment set.

[0052] It should be noted that each segment in the segment set has segment type, document ID, segment sequence number, title content, title level, text content, and image content attributes, wherein the segment type includes paragraph, illustration, and title types.

[0053] In some implementations, the fragment set is also cached, and a file-type relational database is used to store the fragment set, each file corresponds to a database, and the fragment content is stored in a "fragment table".

[0054] Step 101: perform optical character recognition on the illustrations in the fragment set to form a preliminary semantic text, and determine a first token sequence based on the preliminary semantic text.

[0055] Specifically, when executing step 101, the following steps may be specifically performed:

[0056] S1010: Perform optical character recognition on the image content in the segment set, and wrap the optical character recognition results with preset keywords.

[0057] S1011: According to the fragment sequence number, the title, paragraph content and image optical character recognition results are merged to form a preliminary semantic text.

[0058] S1012: extract key sentences from the preliminary semantic text using the TextRank algorithm, wherein the key sentences include text wrapped with preset keywords.

[0059] S1013: Perform word segmentation, word vector embedding conversion, truncation and padding processing on the key sentence to obtain the first token sequence.

[0060] Step 102: Perform image type recognition on the illustrations in the fragment set to form a structure description text, and determine a second token sequence based on the structure description text.

[0061] Specifically, each time a set of fragments of a document is read, OCR (optical character recognition) is performed on the "image content", and the "[illustration START]" and "[illustration END]" keywords are used to wrap the OCR results before and after, and the corresponding records in the file database are updated at the same time. Then, according to the fragment sequence number, the title, paragraph content, and image OCR results are merged to form a preliminary semantic text. Then, the TextRank algorithm is used to extract key sentences in the semantic text, among which the text wrapped by "[illustration START]" and "[illustration END]" has a higher weight and is retained first; the key text is segmented, word vector embedded conversion, truncated and padding are performed to obtain the first token sequence. .

[0062] Specifically, when executing step 102, the following steps may be specifically performed:

[0063] S1020: Perform image type recognition on the illustrations in the fragment set, and wrap them with preset keywords before and after recognition, and regard the wrapped image type as a structural description of the fragment; wherein image type recognition includes but is not limited to table, architecture diagram, flow chart, map, network topology diagram, wireframe diagram, and interface screenshot recognition.

[0064] S1021: Merge the structural descriptions of all the fragments in the order of the fragment numbers to form a structural description text, and perform word segmentation, word embedding vector conversion, truncation and padding processing on the structural description text to obtain a second token sequence.

[0065] Specifically, each time a set of fragments of a document is read, for the "illustration" type fragments, a pre-trained image classification model is used to identify the image type (tables, architecture diagrams, flow charts, maps, network topology diagrams, wireframes, interface screenshots, etc.), and the images are wrapped with the keywords "[illustration START]" and "[illustration END]". The wrapped image type is regarded as the structural description of the fragment, and the corresponding records in the file database are updated at the same time; for "title" type fragments, the title level and title text are regarded as the structural description of the fragment; for "paragraph" type fragments, continuous paragraphs are merged, and the keyword "text paragraph" is used as the structural description of the fragment; the structural descriptions of all fragments are merged in order of fragment numbers to form a structural description text. The structured description text is segmented, converted to a word embedding vector, truncated and padded to obtain the second token sequence. .

[0066] Step 103: Identify the industry or field label of the document to be extracted and determine the code corresponding to the label.

[0067] Specifically, each time a set of document fragments is read, industry or field labels are added to the documents based on keyword matching, pattern recognition and other technologies. One-hot encoding is used to obtain the label encoding T (T is a vector).

[0068] Step 104: Input the first token sequence, the second token sequence, and the codes corresponding to the tags into a pre-built summary extraction model to obtain a summary text.

[0069] Specifically, the summary extraction model is designed based on the encoder-decoder framework of the Seq2Seq model, wherein the first, second and third encoders are used in the encoding stage, and the fourth decoder, the first attention reader unit, the second attention reading unit and the fully connected unit are used in the decoding stage.

[0070] Specifically, the objective function of the summary extraction model is defined as:

[0071] ;

[0072] In the formula, n represents the length of the sequence, X represents the first token sequence, D represents the second token sequence, T represents the encoding corresponding to the label, and log represents the logarithmic function. Indicates the summary content. Represents the summary content before time t , the first token sequence X, the second token sequence D, and the encoding T corresponding to the label, the next token is probability.

[0073] Specifically, when executing step 104, the following steps may be specifically performed:

[0074] S1040: Using a first encoder to convert a first token sequence into a first hidden representation through a bidirectional long short-term memory unit.

[0075] S1041: Utilize a second encoder to convert a second token sequence into a second hidden representation through a bidirectional long short-term memory unit.

[0076] S1042: Use a third encoder to convert the code corresponding to the label into a dense embedding vector.

[0077] S1043: Using the first attention reader unit, taking the state at time t-1 as the query vector, and converting the first hidden representation into a key vector and a value vector, outputting a semantic context vector.

[0078] S1044: Using the second attention reading unit, the state at time t-1 is used as the query vector, and the second hidden representation is used as the key vector and value vector for conversion, and the structural context vector is output.

[0079] S1045: Use a fully connected unit to connect the semantic context vector, the structural context vector, and the dense embedding vector into one vector.

[0080] S1046: Using the fourth decoder, the word vector, semantic context vector, structural context vector and label context vector output by the fully connected unit at time t-1 are used as input to update the state at time t.

[0081] More specifically, each time the extraction results (X, D, T) of a document are read, use Figure 2The recurrent neural network shown in the figure performs reasoning, generates text, obtains a word sequence, and connects the elements in Y to obtain the summary text. The model is designed based on the encoder-decoder framework of the general Seq2Seq model. Three encoders are used in the encoding stage, including LSTMEncoder 1, LSTM Encoder 2, and Embedding 3; one decoder and two attention reader units are used in the decoding stage, including LSTM Decoder 4, Attention 6, Attention 7, and concat 5. The input of the model includes: , represents a token sequence consisting of key semantic text segmentation in a document; , represents the token sequence formed after the document structure summary text is segmented in a document. T is the one-hot encoding vector of the document label. The output of the model is , represents each word in the summary of a document output, and the final text summary is obtained by connecting the elements of Y.

[0082] The model maximizes the true summary content The log-likelihood is used for training, and the objective function is defined as:

[0083] ;

[0084] LSTM Encoder1, which transforms the sequence into Convert to hidden representation ; Any one of the elements The formula for generating is: LSTM Encoder2, which uses a bidirectional long short-term memory (LSTM) unit to encode the sequence Convert to hidden representation ; Any one of the elements The formula for generating is: . Embedding 3, convert the document tag one-hot encoding vector T into a dense embedding vector , can be regarded as feature embedding representing document tags; the conversion formula is , where e represents the word embedding conversion function. LSTM Decoder 4 is a decoder based on unidirectional LSTM. In step t, the decoder converts , the vector of the word output before , context vector ( ) as input to update the state at time t , the conversion formula is: , where [..;..] represents vector concatenation; finally, the decoder uses a fully connected layer to calculate the output probability and samples the current word to be output from the output probability distribution, the formula is: ,in is a weight matrix obtained through learning. concat 5, These four vectors are concatenated into one vector, where is the semantic context, representing the most relevant semantic information in the document at time t. is the structural context, which represents the information about the structure in the document at time t. is the tag context, representing the industry or field of the document. Attention 6 is the first attention reading unit. As Query, As KEY and Value, the conversion formula is: ; ; .in is the weight matrix obtained during training. Attention 7 is the second attention reading unit. As Query, As KEY and Value, the conversion formula is: ; ; .in is the weight matrix obtained during training.

[0085] In some embodiments, the text wrapped by "[illustration START]" and "[illustration END]" is parsed from the preliminary summary text, and compared with the OCR text and image type of the illustration fragment in the "fragment set cache" using a similarity comparison algorithm to obtain the most similar illustration, extract the image data, record the slot position corresponding to the illustration in the text summary, combine the image and text, generate a PDF file, and form a document summary that has both key text descriptions and key illustrations.

[0086] This application is a graphic document summary generation solution based on a recurrent neural network. It uses key semantic extraction, document structure extraction, document industry or field identification, and recurrent neural network technology with multiple encoders to extract key semantics and key illustrations from corporate documents to form a graphic document summary. The text generated by the summary has an industry or field background, and the illustrations can make the summary have the effect of "one picture is worth a thousand words". The generated summary is targeted and has strong summarization ability, which greatly improves the efficiency of application scenarios such as machine-assisted data sensitivity analysis.

[0087] See also Figure 3FIG. 1 is a schematic diagram of a structure of a summary document extraction system provided by an embodiment of the present invention, including:

[0088] The document parsing unit 30 is used to parse the document to be extracted into fragments and generate a fragment set;

[0089] A semantic extraction unit 31, configured to perform optical character recognition on the illustrations in the fragment set to form a preliminary semantic text, and determine a first token sequence based on the preliminary semantic text;

[0090] A structure extraction unit 32, configured to perform image type recognition on the illustrations in the fragment set to form a structure description text, and determine a second token sequence based on the structure description text;

[0091] The tag identification unit 33 is used to identify the industry or field tag of the document to be extracted and determine the code corresponding to the tag;

[0092] The summary inference unit 34 is used to input the first token sequence, the second token sequence and the codes corresponding to the tags into a pre-built summary extraction model to obtain a summary text.

[0093] On the other hand, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the document summary extraction method described in any one of the above embodiments when executing the program.

[0094] On the other hand, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the document summary extraction method described in any of the above embodiments are implemented.

[0095] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0096] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0097] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0099] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0100] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0101] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0102] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0103] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the present application.

Claims

1. A method for extracting a document summary, characterized in that: include: Parse the document to be summarized into fragments and generate a fragment set; Performing optical character recognition on the illustrations in the fragment set to form a preliminary semantic text, and determining a first token sequence based on the preliminary semantic text; Performing image type recognition on the illustrations in the fragment set to form a structure description text, and determining a second token sequence based on the structure description text; Identify the industry or field labels of the documents to be summarized and determine the codes corresponding to the labels; Inputting the first token sequence, the second token sequence, and the codes corresponding to the tags into a pre-built summary extraction model to obtain a summary text; Among them, the objective function of the summary extraction model is defined as: ; In the formula, n represents the length of the sequence, X represents the first token sequence, D represents the second token sequence, T represents the encoding corresponding to the label, and log represents the logarithmic function. Indicates the content of the summary. Represents the summary content before time t , the first token sequence X, the second token sequence D, and the encoding T corresponding to the label, the next token is probability; The first token sequence, the second token sequence, and the code corresponding to the tag are input into a pre-built summary extraction model to obtain a summary text, including: Converting the first token sequence into a first hidden representation using a first encoder via a bidirectional long short-term memory unit; Using a second encoder to convert the second token sequence into a second hidden representation through a bidirectional long short-term memory unit; The third encoder is used to convert the encoding corresponding to the label into a dense embedding vector; Using the first attention reader unit, the state at time t-1 is used as the query vector, the first hidden representation is used as the key vector and the value vector for conversion, and the semantic context vector is output; Using the second attention reading unit, the state at time t-1 is used as the query vector, and the second hidden representation is used as the key vector and value vector for conversion, and the structural context vector is output; Use a fully connected unit to connect the semantic context vector, structural context vector, and dense embedding vector into one vector; The fourth decoder is used to take the word vector, semantic context vector, structural context vector and label context vector output by the fully connected unit at time t-1 as input to update the state at time t.

2. The method for extracting a document summary according to claim 1, characterized in that: Each segment in the segment set has segment type, document ID, segment sequence number, title content, title level, text content, and image content attributes, wherein the segment type includes paragraph, illustration, and title types.

3. The method for extracting a document summary according to claim 2, characterized in that: Performing optical character recognition on the illustrations in the fragment set to form a preliminary semantic text, and determining a first token sequence based on the preliminary semantic text, including: Perform optical character recognition on the image content in the clip set, and wrap the optical character recognition results with preset keywords; According to the fragment sequence number, the title, paragraph content and image optical character recognition results are merged to form a preliminary semantic text; Using the TextRank algorithm, extract key sentences from the preliminary semantic text, wherein the key sentences include text wrapped with preset keywords; The key sentences are segmented, word vectors are embedded, and truncated and padded to obtain the first token sequence.

4. The method for extracting a document summary according to claim 1, characterized in that: Performing image type recognition on the illustrations in the fragment set to form a structure description text, and determining a second token sequence based on the structure description text, including: Perform image type recognition on the illustrations in the clip set, and wrap them with preset keywords before and after recognition, and regard the wrapped image type as the structural description of the clip; image type recognition includes but is not limited to table, architecture diagram, flow chart, map, network topology diagram, wireframe diagram, and interface screenshot recognition; The structural descriptions of all fragments are merged in the order of fragment numbers to form a structural description text, and the structural description text is segmented, converted into word embedding vectors, truncated and padded to obtain the second token sequence.

5. The method for extracting a document summary according to claim 1, characterized in that: The summary extraction model is designed based on the encoder-decoder framework of the Seq2Seq model, in which the first, second and third encoders are used in the encoding stage, and the fourth decoder, the first attention reader unit, the second attention reader unit and the fully connected unit are used in the decoding stage.

6. A document summary extraction system applied to the document summary extraction method according to any one of claims 1 to 5, characterized in that: include: A document parsing unit, used for parsing the document to be extracted into fragments and generating a fragment set; A semantic extraction unit, configured to perform optical character recognition on the illustrations in the fragment set to form a preliminary semantic text, and determine a first token sequence based on the preliminary semantic text; A structure extraction unit, configured to perform image type recognition on the illustrations in the fragment set to form a structure description text, and determine a second token sequence based on the structure description text; A label recognition unit, used to recognize the industry or field label of the document to be extracted and to determine the code corresponding to the label; The summary inference unit is used to input the first token sequence, the second token sequence and the encoding corresponding to the label into a pre-built summary extraction model to obtain a summary text.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the document summary extraction method according to any one of claims 1 to 5 are implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the document summary extraction method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Abstract generation method and related device

    CN115034194A