Text abstract generation method and device, electronic equipment and medium

By extracting the head, tail and middle content of long text or ultra-long text, generating compressed text and entering a text summary model, the problem of obtaining main content in long text or ultra-long text is solved, and the efficiency and effect of text summary generation is improved.

CN120123503APending Publication Date: 2025-06-10SHANGHAI IQIYI NEW MEDIA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510191475.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the film and television industry, streaming media and user interaction platforms, long text or ultra-long text makes it difficult for users to quickly obtain main content, making it difficult to generate effective text summary.

Method used

Generate text summary by extracting the head, tail and intermediate text content from the pending text, generate the compressed pending text, and input it into the text summary generation model.

Benefits of technology

It effectively shortens the length of long text or ultra-long text, so that it can be directly processed by the text summary generation model, improves the generation effect of text summary, and reduces the difficulty and cost of generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123503A_ABST
    Figure CN120123503A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text abstract generation method and device, electronic equipment and a medium, and relates to the technical field of text processing. The method comprises the steps of obtaining a to-be-processed text and a total length value of the to-be-processed text, extracting the to-be-processed text according to a first preset length value to obtain a head text content, extracting the to-be-processed text according to a second preset length value to obtain a tail text content, and extracting the to-be-processed text according to the to-be-processed text, the head text content and the tail text content. Determining middle text content, extracting preset segmented words in the middle text content, generating a compressed to-be-processed text according to the head text content, the preset segmented words in the middle text content and the tail text content, and generating a text abstract of the to-be-processed text. By adopting the technical scheme, the long text is compressed, so that the compressed text length can be directly processed by the text abstract generation model, and the difficulty and the cost of generating the text abstract of the long text are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of text processing, and in particular, to a method, apparatus, electronic device, and medium for generating a text summary. Background Art

[0002] With the rapid development of the film and television industry, streaming media, and various user interaction platforms, a large number of long texts or ultra-long texts have emerged on the film and television industry, streaming media, and various user interaction platforms. For long texts or ultra-long texts, users often cannot quickly obtain the main content of the text.

[0003] Therefore, how to obtain a text summary from a long text or an ultra-long text has become a technical problem that needs to be solved urgently at present, so as to assist users in quickly obtaining the main content from a long text or an ultra-long text based on the text summary. Summary of the Invention

[0004] To solve the above technical problems, the present disclosure provides a method, apparatus, electronic device, and medium for generating a text summary.

[0005] In a first aspect, the present disclosure provides a method for generating a text summary, which is applied to an electronic device. The method includes:

[0006] Obtain the text to be processed and the total length value of the text to be processed;

[0007] Starting from the first word segmentation of the text to be processed, extract the text to be processed according to a first preset length value to obtain the head text content, and starting from the last word segmentation of the text to be processed, extract the text to be processed according to a second preset length value to obtain the tail text content; wherein, the sum of the first preset length value and the second preset length value is less than the total length value;

[0008] Determine the middle text content according to the text to be processed, the head text content, and the tail text content, and extract the preset word segmentations in the middle text content;

[0009] Generate the compressed text to be processed according to the head text content, the preset word segmentations in the middle text content, and the tail text content;

[0010] Input the compressed text to be processed and a guiding statement into a text summary generation model, and generate a text summary of the text to be processed based on the text summary generation model; wherein, the text summary is used to represent the content described by the text to be processed.

[0011] In one example, the determining the middle text content according to the text to be processed, the head text content, and the tail text content includes:

[0012] Delete the head text content and the tail text content from the text to be processed, obtain the remaining text content, and determine the remaining text content as the intermediate text content.

[0013] In one example, extracting the preset word segments in the intermediate text content includes:

[0014] Divide the intermediate text content according to a preset unit to obtain multiple segmented text contents;

[0015] Extract a preset number of word segments at a preset position in each of the segmented text contents as the preset word segments.

[0016] In one example, the dividing the intermediate text content according to a preset unit to obtain multiple segmented text contents includes:

[0017] Divide the intermediate text content with sentences as the preset unit to obtain multiple segmented text contents.

[0018] In one example, the extracting a preset number of word segments at a preset position in each of the segmented text contents as the preset word segments includes:

[0019] Extract two word segments starting from the first word segment in each of the segmented text contents as the preset word segments.

[0020] In one example, the guiding language includes a pre-instruction and a post-instruction. Inputting the compressed text to be processed and the guiding language into the text summary generation model includes:

[0021] Obtain the guiding language from a preset code file; wherein, the preset code file is a pre-configured file; wherein, the pre-instruction and the post-instruction are used to represent the processing tasks of the compressed text to be processed in the text summary generation model;

[0022] Concatenate the pre-instruction, the head text content, the preset word segments in the intermediate text content, the tail text content, and the post-instruction in sequence to generate a concatenated text;

[0023] Input the concatenated text into the text summary generation model.

[0024] In one example, the extracting the head text content by extracting the text to be processed from the first word segment of the text to be processed according to a first preset length value, and extracting the tail text content by extracting the text to be processed from the last word segment of the text to be processed according to a second preset length value includes:

[0025] Starting from the first word segment of the text to be processed, intercept the word segment text content with a quantity equal to the first preset length value in the text to be processed, and determine the intercepted word segment text content with the first preset length value as the head text content;

[0026] Starting from the last word segment of the text to be processed, intercept the word segment text content with a quantity equal to the second preset length value in the text to be processed, and determine the intercepted word segment text content with the second preset length value as the tail text content.

[0027] In a second aspect, the present disclosure provides a text summary generation device configured in an electronic device, and the device includes:

[0028] An acquisition module for acquiring the text to be processed and the total length value of the text to be processed;

[0029] A determination module for extracting the text to be processed from the first word segment of the text to be processed according to the first preset length value to obtain the head text content, and extracting the text to be processed from the last word segment of the text to be processed according to the second preset length value to obtain the tail text content; wherein, the sum of the first preset length value and the second preset length value is less than the total length value;

[0030] An extraction module for determining the intermediate text content according to the text to be processed, the head text content and the tail text content, and extracting the preset word segments in the intermediate text content;

[0031] A first generation module for generating the compressed text to be processed according to the head text content, the preset word segments in the intermediate text content and the tail text content;

[0032] A second generation module for inputting the compressed text to be processed and the guiding language into a text summary generation model, and generating a text summary of the text to be processed based on the text summary generation model; wherein, the text summary is used to characterize the content described by the text to be processed.

[0033] In a third aspect, an embodiment of the present disclosure further provides an electronic device, including:

[0034] A processor;

[0035] A memory for storing executable instructions;

[0036] Wherein, the processor is used to read the executable instructions from the memory and execute the executable instructions to implement the method provided in the first aspect.

[0037] Fourthly, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method provided in the first aspect is implemented.

[0038] An embodiment of the present disclosure provides a method, apparatus, electronic device and medium for generating a text summary. The method includes: obtaining a text to be processed and the total length value of the text to be processed, and then starting from the first word segmentation of the text to be processed, extracting the text to be processed according to a first preset length value to obtain the head text content, and starting from the last word segmentation of the text to be processed, extracting the text to be processed according to a second preset length value to obtain the tail text content. Then, according to the text to be processed, the head text content and the tail text content, the middle text content is determined, and the preset word segmentations in the middle text content are extracted. Then, according to the head text content, the preset word segmentations in the middle text content and the tail text content, the compressed text to be processed is generated. The compressed text to be processed and a guiding statement are input into a text summary generation model, and based on the text summary generation model, the text summary of the text to be processed is generated. In this way, in the face of long texts or ultra-long texts, without losing the description content that helps users understand the text, the long texts or ultra-long texts are compressed, so that the length of the compressed text can be directly processed by the text summary generation model, and the text summary generation model is used to accurately extract the text summary of the compressed long text or ultra-long text. At the same time, the difficulty and cost of generating the text summary of the long text or ultra-long text are reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is a schematic flowchart of a method for generating a text summary provided by an embodiment of the present disclosure;

[0042] Figure 2a It is a schematic flowchart of a method for generating a text summary provided by an embodiment of the present disclosure;

[0043] Figure 2b It is a schematic structural diagram of a compressed text to be processed provided by an embodiment of the present disclosure;

[0044] Figure 3 A structural schematic diagram of a text summary generation device provided by an embodiment of the present disclosure;

[0045] Figure 4 A structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0046] In order to more clearly understand the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0047] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.

[0048] With the rapid development of large language models (LLMs) and the emergence of a large number of open-source models, they have shown great potential in text summary generation tasks. For example, using large language models to generate summaries of novels. However, large language models perform well in processing short texts, but when processing long texts or extremely long texts, the text summary generation effect significantly decreases. In addition, due to the limited computing resources of large language models, it further affects the processing effect of large language models on long texts or extremely long texts.

[0049] Therefore, in the face of long texts or extremely long texts, a text summary generation method with high accuracy needs to be provided to improve the text summary generation effect.

[0050] The following combines Figure 1 To illustrate the text summary generation method provided by an embodiment of the present disclosure. In the embodiment of the present disclosure, the text summary generation method may be executed by an electronic device or a server. The electronic device may include devices with communication functions such as tablet computers, desktop computers, laptop computers, etc., and may also include devices simulated by virtual machines or emulators. The server may be a cloud server or a server cluster. Hereinafter, taking the electronic device as the execution subject, the text summary generation method will be specifically explained.

[0051] Figure 1 Shows a flowchart of a text summary generation method provided by an embodiment of the present disclosure.

[0052] As Figure 1 shown, the text summary generation method may include the following steps:

[0053] S101. Obtain the text to be processed and the total length value of the text to be processed.

[0054] In this embodiment, in the face of long texts or ultra-long texts in different application scenarios, the electronic device takes the long text or ultra-long text as the text to be processed to perform the operation of extracting the text summary from the long text or ultra-long text.

[0055] Optionally, the text to be processed may be a long script in a film and television scenario, a long novel in a streaming media scenario, a long paper in a user retrieval platform scenario, etc. Among them, the total length value of the text to be processed may be 32K - 256K word segments.

[0056] S102. Starting from the first word segment of the text to be processed, extract the text to be processed according to the first preset length value to obtain the head text content, and starting from the last word segment of the text to be processed, extract the text to be processed according to the second preset length value to obtain the tail text content; wherein, the sum value of the first preset length value and the second preset length value is less than the total length value.

[0057] In an example, the first preset length value may be thirty percent of the total length value, or may be ten percent. The second preset length value may be the same length value as the first preset length value, or may be a different length value, and the sum value of the first preset length value and the second preset length value is less than the total length value. For example, the first preset length value is ten percent of the total length value, and the second preset length value is ten percent of the total length value.

[0058] In an example, starting from the first word segment of the text to be processed, extracting the text to be processed according to the first preset length value to obtain the head text content, and starting from the last word segment of the text to be processed, extracting the text to be processed according to the second preset length value to obtain the tail text content includes:

[0059] Starting from the first word segment of the text to be processed, intercept the word segment text content with the quantity of the first preset length value in the text to be processed, and determine the intercepted word segment text content with the first preset length value as the head text content;

[0060] Starting from the last word segment of the text to be processed, intercept the word segment text content with the quantity of the second preset length value in the text to be processed, and determine the intercepted word segment text content with the second preset length value as the tail text content.

[0061] In one example, for instance, if the first preset length value is ten percent of the total length value, then starting from the first token of the text to be processed, intercept the token text content that is ten percent of the total length value, and use this part of the token text content as the header text content. For example, if the second preset length value is ten percent of the total length value, then starting from the last token of the text to be processed, intercept the token text content that is ten percent of the total length value, and use this part of the token text content as the tail text content.

[0062] S103. Determine the middle text content based on the text to be processed, the header text content, and the tail text content, and extract the preset tokens in the middle text content.

[0063] In one example, the middle text content is also a part of the token text content in the text to be processed. The header text content, the tail text content, and the middle text content form the text to be processed. After determining the middle text content, extract the preset tokens in the middle text content.

[0064] S104. Generate the compressed text to be processed based on the header text content, the preset tokens in the middle text content, and the tail text content.

[0065] In this embodiment, since the middle text content becomes the preset tokens in the middle text content, the total length value of the compressed text to be processed obtained from the header text content, the preset tokens in the middle text content, and the tail text content is smaller than the total length value of the text to be processed, thus achieving the compression of the text to be processed.

[0066] S105. Input the compressed text to be processed and the guiding language into the text summary generation model, and generate the text summary of the text to be processed based on the text summary generation model; wherein, the text summary is used to represent the content described by the text to be processed.

[0067] In one example, the guiding language is used to represent that the text summary generation model generates preset indication content according to the compressed text to be processed.

[0068] In one example, the text summary generation model is pre-trained to extract the text summary of the input text. Specifically, the training process of the text summary generation model is as follows: First, manually annotate the text summaries of multiple short texts. The specific content of the text summary includes character introduction, events, time, causes and consequences. Record the token length of each text summary, and then combine multiple short texts according to different strategies to obtain test long texts, and then input the test long texts into the original text summary generation model, where the test long texts carry the annotated text summaries.

[0069] Based on the output of the original text summarization model for the test long text, obtain the predicted text summary, determine the loss function between the predicted text summary and the annotated text summary, and until the loss function meets the preset conditions, then determine the text summarization model at this time as the text summarization model to be determined, and evaluate the text summarization model to be determined. If the evaluation passes, then determine the text summarization model to be determined as the trained text summarization model.

[0070] In this embodiment, the evaluation set for evaluating the text summarization model to be determined is multiple evaluation long texts determined from a preset database. Specifically, the number of word segments of the evaluation long texts is divided into 16K, 32K, 64K, and 128K. Then, each evaluation long text is used to generate multiple versions of the text summary by a closed-source model, and then the multiple versions of the text summary are compared and modified to obtain the evaluation text summary. Input the evaluation long text into the text summarization model to be determined to obtain the real evaluation text summary, and obtain the evaluation index value of the text summarization model to be determined through the difference value between the real evaluation text summary and the evaluation text summary. If the evaluation index value meets the threshold, the evaluation passes.

[0071] Optionally, the preset closed-source models include but are not limited to types such as doubao, qwen, moonshot, gemini, etc.

[0072] Among them, the evaluation index value includes but is not limited to the quality index value (Rouge), the semantic similarity index value (bert-score). In this way, through the semi-automated model testing method, the stability and reliability of the text summarization model are improved. At the same time, the cost of manually generating the text summary corresponding to the compressed test text is saved to achieve a rapid evaluation of the performance of this summary generation model.

[0073] An embodiment of the present disclosure provides a method for generating a text summary. The method includes: obtaining the text to be processed and the total length value of the text to be processed, then starting from the first word segment of the text to be processed, extracting the text to be processed according to a first preset length value to obtain the head text content, and starting from the last word segment of the text to be processed, extracting the text to be processed according to a second preset length value to obtain the tail text content. Then, according to the text to be processed, the head text content, and the tail text content, determine the middle text content, and extract the preset word segments in the middle text content. Then, according to the head text content, the preset word segments in the middle text content, and the tail text content, generate the compressed text to be processed. Input the compressed text to be processed and the guiding language into the text summary generation model, and based on the text summary generation model, generate the text summary of the text to be processed; wherein, the text summary is used to represent the content described by the text to be processed. In this way, in the face of long texts or extremely long texts, without losing the descriptive content that helps users understand the text, compress the long text or extremely long text, so that the length of the compressed text can be directly processed by the text summary generation model, and use the text summary generation model to accurately extract the text summary of the compressed long text or extremely long text. At the same time, the difficulty and cost of generating the text summary of the long text or extremely long text are reduced.

[0074] Figure 2a The flowchart of a method for generating a text summary provided by an embodiment of the present disclosure is shown.

[0075] As Figure 2a shown, the method for generating a text summary may include the following steps:

[0076] S201. Obtain the text to be processed and the total length value of the text to be processed.

[0077] In one example, this step may refer to the content of step S101.

[0078] S202. Starting from the first word segment of the text to be processed, extract the text to be processed according to a first preset length value to obtain the head text content, and starting from the last word segment of the text to be processed, extract the text to be processed according to a second preset length value to obtain the tail text content; wherein, the sum of the first preset length value and the second preset length value is less than the total length value.

[0079] In one example, this step may refer to the content of step S102.

[0080] S203. Delete the head text content and the tail text content from the text to be processed to obtain the remaining text content, and determine the remaining text content as the middle text content.

[0081] In one example, the text to be processed has its head text content removed and its tail text content removed, leaving only the middle part of the text content, and the middle part of the text content is regarded as the intermediate text content.

[0082] S204. Divide the intermediate text content according to a preset unit to obtain multiple segmented text contents.

[0083] In one example, the preset unit can be a chapter, a paragraph, or a sentence, and there is no limitation here. Since the intermediate text content is relatively long, it is necessary to process the intermediate text content to obtain multiple segmented text contents. Specifically, if the preset unit is a chapter, the multiple segmented text contents are the text contents of multiple chapters; if the preset unit is a paragraph, the multiple segmented text contents are the text contents of multiple paragraphs; if the preset unit is a sentence, the multiple segmented text contents are the text contents of multiple sentences.

[0084] In one example, dividing the intermediate text content according to a preset unit to obtain multiple segmented text contents includes:

[0085] Divide the intermediate text content with sentences as the preset unit to obtain multiple segmented text contents.

[0086] In one example, the intermediate text content can be divided with sentences as the preset unit by identifying punctuation marks to obtain multiple segmented text contents.

[0087] S205. At a preset position in each segmented text content, extract a preset number of word segments as preset word segments.

[0088] In one example, the preset position can be any position set in advance. For example, it can be the last position of each segmented text content or the middle position of each segmented text content. The preset number can be two, three, or four.

[0089] In one example, at a preset position in each segmented text content, extracting a preset number of word segments as preset word segments includes:

[0090] Extract two word segments starting from the first word segment in each segmented text content as the preset word segments.

[0091] In this embodiment, since the first two word segments starting from the first word segment in each segmented text content can represent the core meaning of the segmented text content, the first two word segments can be extracted as the preset word segments.

[0092] S206. Obtain the guiding statement from a preset code file; wherein, the preset code file is a pre-configured file; wherein, the pre-instruction and the post-instruction are used to characterize the processing tasks of the text to be processed after compression in the text summary generation model.

[0093] In one example, the pre-instruction and the post-instruction are used to indicate the processing tasks that the text summary generation model performs on the text to be processed. In a specific processing task, the pre-instruction and the post-instruction are both fixed and pre-stored in the preset code file. In this embodiment, the pre-instruction can be "What is input is a fragment of a novel. Please extract the text summary of the fragment", and the post-instruction can be "The above is the content of the novel fragment. The generated text summary is as follows:".

[0094] S207. Concatenate the pre-instruction, the head text content, the preset word segmentation in the middle text content, the tail text content, and the post-instruction in sequence to generate a concatenated text.

[0095] In one example, a concatenated text is generated by concatenating the pre-instruction, the head text content, the preset word segmentation in the middle text content, the tail text content, and the post-instruction in sequence. For a clearer illustration, reference can be made to Figure 2b the structural schematic diagram of a text to be processed after compression shown. Specifically, the pre-instruction can be "What is input is a fragment of a novel. Please extract the text summary of the fragment", and the post-instruction can be "The above is the content of the novel fragment. The generated text summary is as follows:". The head text content can be "I tightly held the yellowed letter in my hand. There were only a few words in the letter: 'Come back. Everything will be revealed.'", and the preset word segmentation in the middle text content can be "investigated person", "ten years ago", "dilapidated old house". The tail text content can be "Everything has settled down. The task in the letter has been completed".

[0096] S208. Input the concatenated text into the text summary generation model, and generate the text summary of the text to be processed based on the text summary generation model; wherein, the text summary is used to characterize the content described by the text to be processed.

[0097] In one example, the content of this step can refer to the content of step S105.

[0098] An embodiment of the present disclosure provides a method for generating a text summary. The method includes: deleting the head text content and the tail text content from the text to be processed to obtain the remaining text content, and determining the remaining text content as the intermediate text content; dividing the intermediate text content according to a preset unit to obtain a plurality of segmented text contents; at a preset position of each segmented text content, extracting a preset number of word segments as preset word segments. Obtain a guiding statement from a preset code file; wherein, the preset code file is a pre-configured file. Concatenate the pre-instruction, the head text content, the preset word segments in the intermediate text content, the tail text content, and the post-instruction in sequence to generate a concatenated text. Input the concatenated text into a text summary generation model, and generate a text summary of the text to be processed based on the text summary generation model. By adopting this technical solution, key information in the head, key information in the tail, and word segment information in the middle can be retained, enabling the model to learn the overall structure and key information of the long text, thereby reducing the length of the text input into the model, reducing the consumption of computing resources, and improving the execution efficiency.

[0099] Figure 3 The figure shows a schematic structural diagram of a text summary generation device provided by an embodiment of the present disclosure. Among them, as Figure 3 shown, the text summary generation device is configured in an electronic device or a server. The electronic device may include devices with communication functions such as a tablet computer, a desktop computer, a laptop computer, etc., and may also include a device simulated by a virtual machine or an emulator. The server may be a cloud server or a server cluster. Hereinafter, the text summary generation device is explained as an electronic device.

[0100] As Figure 3 shown, the text summary generation device 30 may include:

[0101] An acquisition module 301, configured to acquire the text to be processed and the total length value of the text to be processed.

[0102] A determination module 302, configured to extract the text to be processed from the first word segment of the text to be processed according to a first preset length value to obtain the head text content, and extract the text to be processed from the last word segment of the text to be processed according to a second preset length value to obtain the tail text content; wherein, the sum of the first preset length value and the second preset length value is less than the total length value;

[0103] An extraction module 303, configured to determine the intermediate text content according to the text to be processed, the head text content, and the tail text content, and extract the preset word segments in the intermediate text content;

[0104] A first generation module 304, configured to generate the compressed text to be processed according to the head text content, the preset word segments in the intermediate text content, and the tail text content;

[0105] The second generation module 305 is configured to input the compressed text to be processed and the guiding text into a text summary generation model, and generate a text summary of the text to be processed based on the text summary generation model; wherein, the text summary is used to represent the content described in the text to be processed.

[0106] In one example, the extraction module 303 is specifically configured to:

[0107] Delete the head text content and the tail text content of the text to be processed to obtain the remaining text content, and determine the remaining text content as the intermediate text content.

[0108] In one example, the extraction module 303 is specifically configured to: divide the intermediate text content according to a preset unit to obtain a plurality of segmented text contents;

[0109] Extract a preset number of word segments at a preset position of each segmented text content as preset word segments.

[0110] In one example, the extraction module 303 is specifically configured to: divide the intermediate text content by sentences as a preset unit to obtain a plurality of segmented text contents.

[0111] In one example, the extraction module 303 is specifically configured to: extract two word segments starting from the first word segment of each segmented text content as preset word segments.

[0112] The guiding text includes a pre-instruction and a post-instruction. The second generation module 305 is specifically configured to: obtain the guiding text from a preset code file; wherein, the preset code file is a pre-configured file; wherein, the pre-instruction and the post-instruction are used to represent the processing tasks of the compressed text to be processed in the text summary generation model.

[0113] Concatenate the pre-instruction, the head text content, the preset word segments in the intermediate text content, the tail text content, and the post-instruction in sequence to generate a concatenated text;

[0114] Input the concatenated text into the text summary generation model.

[0115] The determination module 302 is specifically configured to: starting from the first word segment of the text to be processed, intercept a word segment text content with a quantity of a first preset length value in the text to be processed, and determine the intercepted word segment text content with the first preset length value as the head text content;

[0116] Starting from the last word segment of the text to be processed, intercept a word segment text content with a quantity of a second preset length value in the text to be processed, and determine the intercepted word segment text content with the second preset length value as the tail text content.

[0117] It should be noted thatFigure 3 The text summary generation device 30 shown can execute Figure 1 each step in the method embodiment shown, and implement Figure 1 each process and effect in the method embodiment shown, which will not be elaborated here.

[0118] Figure 4 The structural schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown.

[0119] As Figure 4 shown, the electronic device may include a processor 401 and a memory 402 storing computer program instructions.

[0120] Specifically, the above-mentioned processor 401 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0121] The memory 402 may include a mass storage for information or instructions. By way of example and not limitation, the memory 402 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 402 may include removable or non-removable (or fixed) media. In a suitable case, the memory 402 may be internal or external to the integrated gateway device. In a specific embodiment, the memory 402 is a non-volatile solid state memory. In a specific embodiment, the memory 402 includes a read-only memory (ROM). In a suitable case, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0122] The processor 401 reads and executes the computer program instructions stored in the memory 402 to execute the steps of the text summary generation method provided by the embodiments of the present disclosure.

[0123] In one example, the electronic device may further include a transceiver 403 and a bus 404. Among them, as Figure 4 shown, the processor 401, the memory 402, and the transceiver 403 are connected through the bus 404 and complete communication with each other.

[0124] The bus 404 includes hardware, software, or both. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side BUS (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 404 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0125] The following are embodiments of a computer-readable storage medium provided by the embodiments of the present disclosure. The computer-readable storage medium and the text summary generation method of the above embodiments belong to the same inventive concept. Details not described in detail in the embodiments of the computer-readable storage medium may refer to the embodiments of the text summary generation method.

[0126] This embodiment provides a storage medium containing computer-executable instructions. When the computer-executable instructions are executed by a computer processor, they are used to execute a text summarization generation method, which is applied to an electronic device. The method includes: obtaining the text to be processed and the total length value of the text to be processed; starting from the first word segmentation of the text to be processed, extracting the text to be processed according to the first preset length value to obtain the head text content, and starting from the last word segmentation of the text to be processed, extracting the text to be processed according to the second preset length value to obtain the tail text content; wherein, the sum of the first preset length value and the second preset length value is less than the total length value; determining the middle text content according to the text to be processed, the head text content and the tail text content, and extracting the preset word segmentations in the middle text content; generating the compressed text to be processed according to the head text content, the preset word segmentations in the middle text content and the tail text content; inputting the compressed text to be processed and the guiding language into a text summarization generation model, and generating a text summary of the text to be processed based on the text summarization generation model; wherein, the text summary is used to represent the content described by the text to be processed.

[0127] Of course, the computer-executable instructions of a storage medium containing computer-executable instructions provided by the embodiments of the present disclosure are not limited to the above method operations, and can also execute the related operations in the text summarization generation method provided by any embodiment of the present disclosure.

[0128] From the above description of the embodiments, those skilled in the art can clearly understand that the present disclosure can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present disclosure, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk or an optical disc of a computer, etc., including several instructions for causing a computer cloud platform (which can be a personal computer, a server, or a network cloud platform, etc.) to execute the text summarization generation method provided by each embodiment of the present disclosure.

[0129] Note that the above is only a preferred embodiment of the present disclosure and the technical principles applied. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present disclosure. Therefore, although the present disclosure has been described in more detail through the above embodiments, the present disclosure is not limited to the above embodiments only. Without departing from the concept of the present disclosure, more other equivalent embodiments can be included, and the scope of the present disclosure is determined by the scope of the appended claims.

Claims

1. A text summary generation method, characterized in that: Applied to electronic equipment, the method comprises: Obtain the text to be processed and the total length value of the text to be processed; Starting from the first word segment of the text to be processed, extracting the text to be processed according to a first preset length value to obtain the head text content, and starting from the last word segment of the text to be processed, extracting the text to be processed according to a second preset length value to obtain the tail text content; wherein the sum of the first preset length value and the second preset length value is less than the total length value; Determine the intermediate text content according to the text to be processed, the header text content and the tail text content, and extract the preset segmentation words in the intermediate text content; Generate a compressed text to be processed according to the header text content, the preset segmented words in the middle text content and the tail text content; The compressed text to be processed and the guide words are input into a text summary generation model, and a text summary of the text to be processed is generated based on the text summary generation model; wherein the text summary is used to characterize the content described by the text to be processed.

2. The method according to claim 1, characterized in that The step of determining the intermediate text content according to the text to be processed, the header text content, and the tail text content includes: The header text content and the tail text content are deleted from the text to be processed to obtain remaining text content, and the remaining text content is determined as the intermediate text content.

3. The method according to claim 1, characterized in that The extracting the preset segmented words in the intermediate text content includes: Dividing the intermediate text content according to preset units to obtain a plurality of segmented text contents; At a preset position of each of the segmented text contents, a preset number of segmented words are extracted as the preset segmented words.

4. The method according to claim 3, characterized in that: The intermediate text content is divided according to preset units to obtain a plurality of segmented text contents, including: The intermediate text content is divided into sentences as the preset units to obtain a plurality of segmented text contents.

5. The method according to claim 3, characterized in that: The step of extracting a preset number of segmented words at a preset position of each segmented text content as the preset segmented words comprises: Starting from the first segmentation in each of the segmented text contents, two segmentations are extracted as the preset segmentations.

6. The method according to claim 1, characterized in that The guide words include a pre-instruction and a post-instruction, and the step of inputting the compressed text to be processed and the guide words into a text summary generation model includes: Obtaining the guide words from a preset code file; wherein the preset code file is a pre-configured file; wherein the pre-instructions and the post-instructions are used to characterize the processing task of the compressed text to be processed in the text summary generation model; The pre-instruction, the header text content, the preset segmentation words in the middle text content, the tail text content and the post-instruction are sequentially spliced ​​to generate a spliced ​​text; The concatenated text is input into the text summary generation model.

7. The method according to claim 1, characterized in that The method starts from the first word segment of the text to be processed, extracts the text to be processed according to a first preset length value to obtain the head text content, and starts from the last word segment of the text to be processed, extracts the text to be processed according to a second preset length value to obtain the tail text content, including: Starting from the first segmented word of the text to be processed, segmented text contents of a quantity equal to the first preset length value are intercepted from the text to be processed, and the segmented text contents of the first preset length value that are intercepted are determined as the header text contents; Starting from the last word segmentation of the text to be processed, word segmentation text contents of the second preset length value are intercepted from the text to be processed, and the intercepted word segmentation text contents of the second preset length value are determined as the tail text contents.

8. A text summary generation device, characterized in that: Applied to electronic equipment, the device comprises: An acquisition module, used for acquiring the text to be processed and the total length value of the text to be processed; A determination module is used to extract the text to be processed according to a first preset length value starting from the first participle of the text to be processed to obtain the head text content, and to extract the text to be processed according to a second preset length value starting from the last participle of the text to be processed to obtain the tail text content; wherein the sum of the first preset length value and the second preset length value is less than the total length value; An extraction module, used to determine the intermediate text content according to the text to be processed, the header text content and the tail text content, and extract the preset segmentation words in the intermediate text content; A first generating module, used to generate a compressed text to be processed according to the header text content, the preset segmentation words in the middle text content and the tail text content; The second generation module is used to input the compressed text to be processed and the guide words into a text summary generation model, and generate a text summary of the text to be processed based on the text summary generation model; wherein the text summary is used to characterize the content described by the text to be processed.

9. An electronic device, characterized in that: include: processor; A memory for storing executable instructions; The processor is used to read the executable instructions from the memory and execute the executable instructions to implement the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the method according to any one of claims 1 to 7.