Abstract generation method and apparatus, intelligent terminal, and computer-readable storage medium

By allowing users to select text from a list of texts to generate a summary text, this technology solves the problem of inconvenient summary text acquisition in existing technologies and enables convenient summary text retrieval.

CN116719927BActive Publication Date: 2026-04-14IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2023-04-28
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, users need to manually review the text data to obtain the summary text, which makes obtaining the summary text inconvenient.

Method used

A summary generation method is provided, which generates and displays summary text based on the selected text in the text to be processed by the user.

Benefits of technology

It improves the ease with which users can obtain summary text, allowing users to obtain summary text with a high degree of matching with the selected text simply by selecting a portion of the text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116719927B_ABST
    Figure CN116719927B_ABST
Patent Text Reader

Abstract

The application discloses a summary generation method and device, an intelligent terminal and a computer readable storage medium. The method comprises the following steps: obtaining selected Chinese text selected by a user in a text to be processed; obtaining a summary text based on at least the selected Chinese text; and displaying the summary text. The above scheme can improve the convenience of obtaining the summary text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, smart terminal, and computer-readable storage medium for generating abstracts. Background Technology

[0002] With the advent of the data era, users encounter vast amounts of text data in various aspects of life, work, and study. Therefore, users increasingly desire a convenient and reliable way to extract summaries from this text data. Current technologies typically only present the text data to users, while the process of extracting summaries relies on users reviewing the text data and extracting the desired summaries, which is extremely inconvenient for them. Therefore, improving the ease of extracting summaries has become a pressing issue. Summary of the Invention

[0003] The main technical problem addressed by this application is to provide a method, apparatus, smart terminal, and computer-readable storage medium for generating abstracts, which can improve the convenience of obtaining abstract text.

[0004] To address the aforementioned technical problems, the first aspect of this application provides a summary generation method, comprising: obtaining selected text selected by a user in a text to be processed; obtaining summary text based at least on the selected text; and displaying the summary text.

[0005] To address the aforementioned technical problems, a second aspect of this application provides a summary generation apparatus, comprising: an acquisition module, a generation module, and a display module, wherein the acquisition module is used to acquire selected text selected by a user in a text to be processed; the generation module is used to obtain summary text based at least on the selected text; and the display module is used to display the summary text.

[0006] To address the aforementioned technical problems, a third aspect of this application provides a smart terminal, including a display screen, a memory, and a processor, wherein the display screen and the memory are respectively coupled to the processor, the display screen is used at least to display content to a user and to allow the user to select content, the memory stores program instructions, and the processor is used to execute the program instructions to implement the summary generation method described in the first aspect.

[0007] To address the aforementioned technical problems, a fourth aspect of this application provides a computer-readable storage medium storing program instructions thereon, which, when executed by a processor, implement the digest generation method described in the first aspect.

[0008] The above solution involves obtaining the selected text from the text to be processed after the user selects at least a portion of the text. Based on this selected text, a summary text is generated, ensuring that the summary text is at least related to the selected text. In other words, the reference range for generating the summary text includes at least the selected text. After obtaining the summary text, it is displayed. Therefore, the user only needs to view the text and select a portion, and the process of selecting the selected text satisfies the user's need for free choice. Once the user selects at least a portion of the text from the text to be processed, a summary text with a high degree of matching to the selected text is obtained, improving the convenience for the user to obtain the summary text. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0010] Figure 1 This is a flowchart illustrating one embodiment of the abstract generation method of this application;

[0011] Figure 2 This is a flowchart illustrating another embodiment of the abstract generation method of this application;

[0012] Figure 3 yes Figure 2 A schematic diagram of an application scenario corresponding to step S201 in the middle;

[0013] Figure 4 yes Figure 2 A schematic diagram of an application scenario corresponding to another implementation method of step S201;

[0014] Figure 5 yes Figure 2 A schematic diagram of an application scenario corresponding to another implementation method in step S201;

[0015] Figure 6 This is a flowchart illustrating another implementation of the abstract generation method of this application;

[0016] Figure 7 yes Figure 6 A schematic diagram of an application scenario corresponding to step S602 in the middle;

[0017] Figure 8 This is a flowchart illustrating another implementation of the abstract generation method of this application;

[0018] Figure 9 yes Figure 8 A schematic diagram of an application scenario corresponding to step S802 in the middle step;

[0019] Figure 10 yes Figure 8 A schematic diagram of an application scenario corresponding to step S806 in the middle step;

[0020] Figure 11 yes Figure 8 A schematic diagram of an application scenario corresponding to another implementation method of step S806;

[0021] Figure 12 This is a schematic diagram of one embodiment of the abstract generation apparatus for this application.

[0022] Figure 13 This is a schematic diagram of the structure of one embodiment of the smart terminal of this application;

[0023] Figure 14 This is a schematic diagram of one embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0025] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper means two or more.

[0026] The abstract generation method provided in this application relies on an application on a smart terminal or at least a smart terminal with integrated text processing functions. The execution entity corresponding to the abstract generation method provided in this application is the processor of the smart terminal.

[0027] Please see Figure 1 , Figure 1 This is a flowchart illustrating one embodiment of the abstract generation method of this application, which includes:

[0028] S101: Get the selected text selected by the user in the text to be processed.

[0029] Specifically, after a user selects at least part of the text in the text to be processed, the text selected by the user in the text to be processed is obtained as the selected text.

[0030] Furthermore, users can select text from the text to be processed using preset selection methods. After a user selects at least a portion of the text using preset selection methods, the selected text is retrieved. The preset selection methods include at least one of drawing, checking, and dragging operations.

[0031] In one application method, in response to detecting a user's drawing operation, the text selected by the user in the text to be processed through the drawing operation is obtained as the selected text. The drawing operation includes drawing a closed shape or drawing a preset symbol; this application does not impose specific limitations on closed shapes and preset symbols.

[0032] In one application scenario, a selection guide for drawing operations is displayed to inform users that by drawing a closed shape (such as an ellipse or polygon), the text contained within the shape can be selected as the selected text. In response to the user selecting a portion of the text to be processed using the drawing operation, the text selected by the user in the text to be processed through the drawing operation is obtained as the selected text.

[0033] In another application scenario, the drawing operation is displayed with a selection description so that the user can know that by drawing a preset symbol (such as parentheses or a straight line), the text corresponding to the preset symbol can be selected as the selected text. For example, the text inside the parentheses is selected as the selected text, or the text corresponding to the straight line is selected as the selected text. In response to the user selecting part of the text in the text to be processed by using the drawing operation, the text selected by the user in the text to be processed by the drawing operation is obtained as the selected text.

[0034] In another application, in response to detecting a user's checkmark operation, the text selected by the user in the text to be processed through the checkmark operation is obtained as the selected text. The text to be processed includes multiple first candidate options, and the checkmark operation includes at least one of single-click, double-click, and sliding selection. This application does not impose specific limitations on the shape of the first candidate options or the triggering form of the checkmark operation.

[0035] In one application scenario, a first candidate option is displayed before each natural paragraph in the text to be processed using a rectangle. In response to the user checking at least some of the first candidate options and confirming, the text selected by the user in the text to be processed is obtained as the selected text.

[0036] In another application scenario, at least some of the first candidate options are displayed in the text to be processed using circles. The first candidate options are related to the number of characters in a natural paragraph. The number of characters following each first candidate option exceeds a character count threshold. In response to the user checking at least some of the first candidate options and confirming, the text selected by the user in the text to be processed through the check operation is obtained as the selected text.

[0037] In another application, in response to detecting a user's drag-and-select operation, the text selected by the user in the text to be processed through the drag-and-select operation is obtained as the selected text. The drag-and-select operation includes continuous movement operations following a click or touch; this application does not impose specific limitations on the method of drag-and-select operation.

[0038] In one application scenario, the drag-and-drop operation is explained to inform the user that by dragging the cursor, the selected text can be designated as the selected text. In response to the user selecting a portion of the text to be processed using the drag-and-drop operation, the text selected by the user in the text to be processed is obtained as the selected text.

[0039] In another application scenario, the drag-and-drop operation is explained to inform users that by dragging and selecting text, the selected text can be designated as the selected text. In response to the user selecting a portion of the text to be processed using the drag-and-drop operation, the text selected by the user in the text to be processed is obtained as the selected text.

[0040] In a specific application scenario, the text to be processed is displayed on the touch screen of a smart terminal (such as a mobile phone or tablet). The system detects the user's actions of drawing an ellipse or parentheses on the touch screen, and selects the text enclosed in the ellipse or the text between the parentheses as the selected text.

[0041] In another specific application scenario, the text to be processed is displayed on the screen of a smart terminal (such as a laptop or desktop computer), and the text to be processed has a checkbox and multiple radio buttons. When the checkbox is checked, all the radio buttons are selected. The system detects the user's operation of checking the checkbox or at least some of the radio buttons, and the text corresponding to the selected radio buttons is taken as the selected text.

[0042] In another specific application scenario, the text to be processed is displayed on the screen of a smart terminal (such as a laptop or desktop computer). The text selected by the user by dragging the cursor in the text to be processed is detected, and the text selected by the cursor is taken as the selected text.

[0043] S102: Obtain the summary text based at least on the selected text.

[0044] Specifically, extract key content from the selected text to obtain summary text, or extract key content from at least a portion of the selected and unselected text to obtain summary text, wherein the unselected text includes text other than the selected text in the text to be processed.

[0045] In one application method, the first keyword and the first key sentence are extracted from the selected text. The first keyword and the first key sentence are then integrated to obtain the summary text corresponding to the selected text, so as to match the user's WYSIWYG requirement and improve the user's convenience in obtaining the summary text.

[0046] In another application, the first keyword and the first key sentence are extracted from the selected text. Based on the first keyword and the first key sentence, reference text matching the first keyword and the first key sentence is obtained from the unselected text. The second keyword and the second key sentence are extracted from the reference text. The first keyword and the first key sentence, as well as the second keyword and the second key sentence, are integrated to obtain the summary text corresponding to the selected text. This makes the summary text have a higher degree of key content relevance with the selected text, improving the convenience for users to obtain the summary text and the accuracy of the obtained summary text.

[0047] In another application, the semantic information corresponding to the selected text is obtained, and text that matches the semantic information of the selected text is extracted from the unselected text as reference text. The first keyword and the first key sentence in the selected text, as well as the second keyword and the second key sentence in the reference text, are extracted. The first keyword and the first key sentence, as well as the second keyword and the second key sentence, are integrated to obtain the summary text corresponding to the selected text. This makes the summary text have a higher semantic fit with the selected text, improves the convenience for users to obtain the summary text, and increases the accuracy of the obtained summary text.

[0048] In one application scenario, the first keyword and the first key sentence in the selected text are extracted based on the semantics of the selected text. The first keyword and the first key sentence are then integrated to obtain the summary text corresponding to the selected text.

[0049] In another application scenario, the first keyword and the first key sentence are extracted from the selected text. Based on the frequency of the first keyword and the first key sentence appearing in the unselected text, reference text is extracted from the unselected text. The frequency of the first keyword or the first key sentence in the reference text is greater than the frequency threshold, and the frequency threshold can be any positive integer (such as 2, 3 or 4). The second keyword and the second key sentence are extracted from the reference text. The first keyword and the first key sentence, as well as the second keyword and the second key sentence, are integrated to obtain the summary text corresponding to the selected text.

[0050] In another application scenario, the selected and unselected texts are fed into a pre-trained semantic extraction model to obtain the output of the semantic extraction model. The semantic information corresponding to the selected text and the unselected text is obtained. Content that matches the semantic information of the selected text is extracted from the unselected text to obtain the reference text. The first keyword and the first key sentence in the selected text, as well as the second keyword and the second key sentence in the reference text, are extracted. The first keyword and the first key sentence, as well as the second keyword and the second key sentence, are integrated to obtain the summary text corresponding to the selected text.

[0051] S103: Display the summary text.

[0052] Specifically, after obtaining the summary text, the summary text is displayed. Therefore, users only need to select a portion of the text they notice to obtain summary text that closely matches the selected text, thus improving the ease with which users can obtain summary text.

[0053] In one application mode, a display area matching the summary text is obtained, and the summary text is displayed in the display area. When the user selects the editing option for the display area, the display area is adjusted to be an editable area.

[0054] In another application, an editable box is used to display the summary text until the user selects the confirmation option, at which point the editable box is adjusted to become the summary display area. Within this editable box, the user can add, delete, and modify the summary text.

[0055] In one application scenario, a display area matching the summary text is obtained in the display interface of the text to be processed. The summary text is displayed after the text to be processed. When the user selects the editing option of the display area, the display area is adjusted to an editable area. When the user selects the confirmation option, the summary text in the display area is moved to another interface independent of the display interface and saved.

[0056] In another application scenario, other interfaces that do not overlap with the display interface of the text to be processed are obtained. The display area that matches the summary text is obtained in the other interface. The summary text is displayed using an editable box so that the user can add, delete, or modify it within the editable box until the user selects the confirmation option of the editable box. The editable box is then adjusted to the summary display area, and the summary text within the summary display area is saved in the corresponding interface.

[0057] The above solution involves obtaining the selected text from the text to be processed after the user selects at least a portion of the text. Based on this selected text, a summary text is generated, ensuring that the summary text is at least related to the selected text. In other words, the reference range for generating the summary text includes at least the selected text. After obtaining the summary text, it is displayed. Therefore, the user only needs to view the text and select a portion, and the process of selecting the selected text satisfies the user's need for free choice. Once the user selects at least a portion of the text from the text to be processed, a summary text with a high degree of matching to the selected text is obtained, improving the convenience for the user to obtain the summary text.

[0058] Please see Figure 2 , Figure 2 This is a flowchart illustrating another embodiment of the abstract generation method of this application, the method comprising:

[0059] S201: Get the selected text selected by the user in the text to be processed.

[0060] Display at least one first candidate in the text to be processed, wherein each first candidate corresponds to a text block.

[0061] In one implementation scenario, before obtaining the selected text chosen by the user in the text to be processed, the method further includes: displaying at least one first candidate option in the text to be processed; wherein each first candidate option corresponds to a text block.

[0062] Specifically, please refer to Figure 3 , Figure 3 yes Figure 2 A schematic diagram of an application scenario corresponding to step S201 shows that at least one first candidate (e.g., ...) is set in the text to be processed. Figure 3 The text is divided into rectangles and columns, and the first candidate is displayed in the text to be processed, with each first candidate followed by a text block. The text block contains at least a portion of the content of the text to be processed.

[0063] Furthermore, when the user selects the first candidate (e.g.) Figure 3 The checkmarked rectangle will select all text in the text block following the first candidate selected by the user, thus improving the convenience of text selection for the user.

[0064] In one application method, at least one first candidate option is generated based on the semantics corresponding to the natural paragraphs in the text to be processed, and all generated first candidate options are displayed in the text to be processed; wherein, when the text block corresponding to the first candidate option includes multiple natural paragraphs, the multiple natural paragraphs in the text block are semantically related.

[0065] Specifically, the semantics corresponding to each natural paragraph in the text to be processed are obtained. Based on the semantics corresponding to each natural paragraph, natural paragraphs with semantic relevance are determined. Thus, any natural paragraph that has no semantic relevance to other natural paragraphs or multiple natural paragraphs that have consecutive semantic relevance are taken as text blocks.

[0066] Furthermore, a corresponding first candidate option is generated for the text block so that when the user selects the first candidate option, they can obtain multiple natural paragraphs with semantic relevance, and adapt to the situation where the probability of consecutive natural paragraphs in regular text having semantic relevance is relatively high.

[0067] In one application scenario, the semantics of each natural paragraph are obtained based on a pre-trained semantic model.

[0068] In a specific application scenario, such as Figure 3 As shown, the first candidate in the text to be processed corresponds to two consecutive natural paragraphs that are semantically related, while the other two candidate options each correspond to a natural paragraph that is semantically relatively independent.

[0069] In another application, at least one first candidate is generated based on the semantics of the text content and the number of characters in the text to be processed, and all generated first candidate options are displayed in the text to be processed; wherein, the content in the text block corresponding to the first candidate option has semantic relevance, and the total number of characters in the text block corresponding to the first candidate option is greater than the character count threshold.

[0070] Specifically, please refer to Figure 4 , Figure 4 yes Figure 2 The following is an application scenario diagram of another implementation corresponding to step S201: the semantics of the text content in the text to be processed are obtained, and the text to be processed is divided into at least one text block with a total number of characters greater than the character count threshold, so that the obtained text block has semantic relevance and has a sufficient number of characters.

[0071] Furthermore, set a first candidate for the text block (such as...). Figure 4 (The gray circle in the text), where the first candidate is not limited to being set before a natural paragraph, so that when the user selects the first candidate, they can obtain content with a total number of characters greater than the character count threshold and semantic relevance, so that more accurate summary text can be extracted after the content in the text block is selected.

[0072] In one application scenario, text segmentation position prediction is performed based on a neural network sequence labeling framework. This involves sequentially inputting sentences from the text and predicting whether each sentence is at the segmentation boundary. During the training and prediction process, the segment length information features from the segmentation boundary of the most recent historical moment to the current moment, the cue word features of the current sentence, and the contextual semantic expression of the current sentence are fused together to achieve the effect of segmenting text based on semantic understanding and effectively constraining the paragraph length.

[0073] In another application, at least one first candidate option is generated based on the semantic groups corresponding to keywords and key sentences in the text to be processed, and all generated first candidate options are displayed in the text to be processed; wherein, at least part of the content in the text block corresponding to the first candidate option corresponds to the same semantic group.

[0074] Specifically, please refer to Figure 5 , Figure 5 yes Figure 2 The diagram illustrates an application scenario of another implementation method corresponding to step S201. The text to be processed is segmented into words and sentences. Based on the semantic groups corresponding to keywords and key sentences in the text to be processed, the text to be processed is divided into at least one text block, so that at least some of the content in the obtained text block corresponds to the same semantic group.

[0075] Furthermore, set a first candidate for the text block (such as...). Figure 5 (The black triangle in the text), where the first candidate is not limited to being set before a natural paragraph, so that when the user selects the first candidate, they can get at least some of the content with the same meaning group, so that the content in the text block can be extracted more accurately after the content is selected.

[0076] In one application scenario, a large amount of text with keywords, key sentences extracted, and semantic groups segmented is collected in advance. A joint prediction hierarchical model for keyword and key sentence extraction is constructed. The hidden features of the sentence are extracted by fusing contextual information, and the distribution features of the word-level keywords and sentence-level key sentences are extracted after the combination. A sentence-level segmentation prediction model that integrates keyword and key sentence information is constructed. Then, the trained joint prediction hierarchical model is used to segment keywords and key sentences, and the segmentation prediction model is used to divide the text to be processed into at least one text block.

[0077] In another application, in response to obtaining auxiliary content that matches the text to be processed, auxiliary text is extracted from the auxiliary content, at least one first candidate is generated based on the auxiliary text, and all generated first candidate options are displayed in the text to be processed; wherein the content of the text block corresponding to at least one first candidate is related to the auxiliary text.

[0078] Specifically, when auxiliary content uploaded by the user that matches the text to be processed is obtained, auxiliary text is extracted from the auxiliary content. The auxiliary content includes at least one of images, handwritten text, and speech. Thus, the auxiliary text can be extracted from images or handwritten text, or obtained by converting speech into text.

[0079] Furthermore, the semantics of the auxiliary text are extracted, and the text to be processed is segmented based on the semantics of the auxiliary text to obtain at least one text block. The content of at least one text block matches the semantics of the auxiliary text. A first candidate is set for the text block. The first candidate is not limited to being set before a natural paragraph, so that when the user selects the first candidate, he / she can select a text block that matches the auxiliary text in the uploaded auxiliary content, thereby improving the matching degree between the summary text and the user's needs.

[0080] In a specific application scenario, the text to be processed is the text obtained by transcribing speech related to a meeting, and the auxiliary content is a picture taken in the meeting scene input by the user. The picture contains the key content required by the user. OCR recognition is performed on the picture to extract the text on the picture as auxiliary text. The text to be processed is segmented based on the semantics of the auxiliary text to obtain at least one text block that matches the semantics of the auxiliary text, so as to highly match the key content required by the user.

[0081] Furthermore, obtaining the selected text selected by the user in the text to be processed includes: in response to at least one first candidate being selected, treating the content of the text blocks corresponding to all selected first candidate options as the selected text.

[0082] Understandably, when at least one first candidate is selected, the content of the text block corresponding to all selected first candidates is treated as selected text. Users can select first candidates in any triggering manner, thereby treating the content of the text block corresponding to the user's selected first candidate as selected text.

[0083] It should be noted that any of the methods disclosed in the above implementation scenarios, which involve generating a first candidate option and providing it to the user for selection, thereby obtaining the content of the text block corresponding to the first candidate option selected by the user and obtaining the selected text, can also be used in the previous embodiment.

[0084] S202: Based on at least a portion of the content in the selected text, obtain reference text from the unselected text that is semantically related to the selected text, wherein the unselected text includes text other than the selected text in the text to be processed.

[0085] Specifically, based on at least a portion of the content in the selected text, text that is semantically related to the selected text in the unselected text is obtained as reference text, wherein the unselected text is text other than the selected text in the text to be processed.

[0086] In one application method, the semantic information corresponding to the selected text is obtained, and the content that matches the semantic information of the selected text is extracted from the unselected text. The text in the unselected text that has a semantic relationship with the selected text is obtained as reference text. Thus, the reference text is obtained based on the complete semantic information corresponding to the selected text, thereby improving the reliability of the obtained reference text.

[0087] In another application, the extracted content is obtained from the selected text, and the semantic information of the extracted content is obtained. Content that matches the semantic information of the extracted content is extracted from the unselected text. The text in the unselected text that has semantic relevance to the selected text is obtained as reference text. Thus, the reference text is obtained based on the semantic information of the extracted content, which improves the efficiency of obtaining reference text and reduces the computational power consumption of obtaining reference text.

[0088] In one application scenario, the selected text and the unselected text are fed into a pre-trained semantic extraction model. The output of the semantic extraction model is obtained, and the semantic information corresponding to the selected text and the unselected text is obtained. The content that matches the semantic information of the selected text is extracted from the unselected text to obtain the reference text.

[0089] In another application scenario, a first prompt is generated based on the selected text. The first prompt is then input into a large language model, which extracts the refined content and semantic information of the refined content from the selected text based on the first prompt. The unselected text is then fed into a pre-trained semantic extraction model, and the output of the semantic extraction model is obtained to obtain the semantic information corresponding to the unselected text. Finally, content that matches the semantic information of the refined content is extracted from the unselected text to obtain the reference text.

[0090] In a specific application scenario, in order to make users clearly aware that the summary text is generated together with the reference text in addition to the selected text, the reference text is highlighted in the text to be processed after the reference text is obtained.

[0091] In another specific application scenario, in order to make users clearly aware that in addition to the selected text, there is also reference text and that users can make a selection, a radio button is set in front of the reference text after it is obtained, and all reference texts have a corresponding select all box so that users can freely choose the reference text to use.

[0092] S203: Obtain the summary text based on the selected text and the reference text.

[0093] Specifically, based on the selected text and the reference text, key content related to the selected text is extracted to obtain the summary text, thereby expanding the reference range of the generated summary text and reducing the probability that the summary text is not accurate due to the user omitting content that is semantically related to the selected text.

[0094] In one application method, the first key sentence in the selected text and the second key sentence in the reference text are extracted. The text composed of all the first key sentences and some of the second key sentences is obtained to obtain the summary text. Thus, the summary text is generated with the first key sentence as the main content, which improves the matching degree between the summary text and the selected text selected by the user.

[0095] In another application, the first key sentence in the selected text and the second key sentence in the reference text are extracted to obtain a text composed of at least part of the first key sentence and at least part of the second key sentence, thus obtaining a summary text. This improves the accuracy of the summary text by combining the first and second key sentences.

[0096] In one application scenario, a first weight is set for the selected text, and a second weight is set for the reference text. The first key sentence in the selected text and the second key sentence in the reference text are extracted. Based on the first weight and the first key sentence, as well as the second weight and the second key sentence, the summary text is obtained.

[0097] In a specific application scenario, the first weight of the selected text is greater than the second weight of the reference text, thus using the second key sentence to assist the first key sentence in generating the summary text.

[0098] In another specific application scenario, the first weight of the selected text is equal to the second weight of the reference text, thus combining the first and second key sentences to generate the summary text.

[0099] In one implementation scenario, based on at least a portion of the content in the selected text, a reference text that is semantically related to the selected text is obtained from the unselected text, including: extracting refined content from the selected text; based on the semantics of the extracted content, obtaining a reference text that is semantically related to the extracted content from the unselected text; and generating and displaying prompt information corresponding to the reference text.

[0100] Specifically, key content is extracted from the selected text as refined content, and the semantics of the refined content are obtained so that key content can be obtained from the selected text first, thereby obtaining semantics with a higher degree of matching with the selected content.

[0101] Furthermore, based on the semantics of the extracted content, reference texts that are semantically related to the extracted content are extracted from the unselected text. Prompt information is generated for the reference texts and displayed to the user so that the user clearly understands that the summary text is generated together with the reference texts in addition to the selected text.

[0102] In one application scenario, to facilitate user confirmation of extracted content, after extracting the extracted content from the selected text, the extracted content is highlighted and a confirmation box for the extracted content is generated. The confirmation box can be dragged and selected so that users can freely select the extracted content from the selected text. Then, when the user confirms the extracted content in the confirmation box, the content in the confirmation box is taken as the final extracted content.

[0103] In a specific application scenario, after obtaining the reference text, the position of the reference text in the text to be processed is highlighted. The highlighting can be achieved by using any background color or by making the text bold, so that the user can perceive the position of the reference text.

[0104] Optionally, the prompt message includes at least one second candidate option corresponding to the reference text; after generating and displaying the prompt message corresponding to the reference text, it further includes: in response to all second candidate options being confirmed or cancelled, using the reference text corresponding to the confirmed second candidate option as the confirmation text.

[0105] Specifically, the prompt message includes at least one second candidate text corresponding to the reference text, so that the user can freely select the reference text to be adopted. When all second candidates are confirmed or cancelled, the reference text corresponding to the confirmed second candidate text is used as the confirmation text.

[0106] Furthermore, based on the selected text and the reference text, a summary text is obtained, including: based on the selected text and the confirmed text, a summary text is obtained.

[0107] Specifically, after obtaining the confirmation text from the user, a summary text is generated based on the selected text and the confirmation text. This ensures that the reference scope of the summary text is more closely aligned with the user's needs, thereby improving the user experience.

[0108] In a specific application scenario, the second candidate options include a "select all" and "select all" option for all reference texts at all locations, as well as a "select one" and "select one" option for each reference text at each location, allowing users to freely select reference texts at different locations. Once all reference texts at all locations are confirmed or cancelled, the reference text confirmed by the user is obtained as the confirmation text.

[0109] S204: Display the summary text.

[0110] Specifically, the obtained summary text is displayed. The display method is the same as in the previous embodiment, and will not be repeated here.

[0111] Unlike the previous embodiments, this method sets at least one first candidate option in the text to be processed, followed by a corresponding text block. When a user selects the first candidate option, the content of the text block corresponding to the selected first candidate option is taken as the selected text, thereby improving the user's convenience in selecting text. After the user selects at least a portion of the text in the text to be processed, the selected text is obtained as the selected text. Based on at least a portion of the content in the selected text, text that is semantically related to the selected text is obtained from the unselected text as reference text. The unselected text refers to text other than the selected text in the text to be processed. Based on the selected text and the reference text, a summary text is obtained, thereby expanding the reference range for generating the summary text and reducing the probability of low accuracy of the summary text due to the user omitting content that is semantically related to the selected text. After obtaining the summary text, it is displayed. Therefore, the user only needs to select a portion of the text that the user notices to obtain a summary text with a high degree of matching with the selected text, improving the user's convenience in obtaining the summary text.

[0112] Please see Figure 6 , Figure 6 This is a flowchart illustrating another embodiment of the abstract generation method of this application, the method comprising:

[0113] S601: In response to receiving voice data input by the user, convert the voice data into text to be processed.

[0114] Specifically, when user-input voice data is obtained, the voice data is converted into text to obtain the text to be processed, thereby adapting to the scenario where the user needs to extract a summary from the voice data.

[0115] In one application scenario, voice data refers to data uploaded by users on smart terminal applications. The voice data uploaded by users is converted into text to be processed, so that users can upload voice data collected in any scenario and obtain text to be processed.

[0116] In another application scenario, voice data is data collected by smart terminal applications during use. When the application ends, the collected voice data is converted into text to be processed, so that the user can obtain the text corresponding to the voice data collected during the application's use after the application ends.

[0117] In another application scenario, voice data is collected by a smart terminal that integrates text processing and voice acquisition functions. When the user selects to convert the voice data to text on the smart terminal, the voice data collected in this instance is converted into text to be processed, so that the user can freely collect voice data in different scenarios and freely choose whether to transcribe it.

[0118] In a specific application scenario, the voice data is the voice data collected by the online conferencing application on the smart terminal during the meeting. After the online meeting ends, the voice data collected this time is converted into text to be processed.

[0119] In another specific application scenario, voice data refers to the voice data collected by the user after triggering the recording option during an online or offline meeting on a smart terminal. When the user chooses to convert the voice into text, the voice data collected this time is converted into text to be processed.

[0120] S602: Display the text to be processed in the first area.

[0121] Specifically, please refer to Figure 7 , Figure 7 yes Figure 6 A schematic diagram of an application scenario corresponding to step S602, wherein the first region is Figure 7 The dashed box on the left side of the first area displays the text to be processed.

[0122] Optionally, in addition to the text converted from speech data, the text to be processed also includes the speaker of the corresponding content in the text. When the speaker matches the voiceprint database, the speaker is clearly indicated in the text to be processed.

[0123] S603: Obtain at least a portion of the text to be processed selected by the user in the first area, and display the selected content in the first area as note text in the second area, wherein the second area does not overlap with the first area.

[0124] For details, please continue reading Figure 7 When a user selects at least a portion of the text to be processed in the first area, the selected content in the first area is displayed as note text in the second area, where the second area is... Figure 7 The dashed box on the right side of the text, and the second area does not overlap with the first area, so that when the user makes a selection in the text to be processed, the selected content can be displayed separately, making the selected content more intuitive.

[0125] In one application scenario, a split-screen operation is performed on the display interface to obtain a non-overlapping first area and a second area. The text to be processed is displayed in the first area. When the user selects part of the content in the first area, the selected content in the first area is displayed in the second area so that the user can view all the content selected in the first area in the second area.

[0126] S604: Obtain at least one note text selected by the user in the second area, and treat the selected note text in the second area as the selected text.

[0127] Specifically, when a user selects at least one note text in the second area, the selected note text in the second area is treated as the selected text. This allows the user to pre-select some content in the first area as note text for centralized display in the second area. The user can then filter the note text in the second area and use the selected note text in the second area as the user's confirmed selected text. This allows the user to compare and refer to the selected text, resulting in a more accurate selection of text. Consequently, when generating summary text, the user can obtain a summary text that better matches their needs.

[0128] Optionally, the second area provides users with options to add, delete, and modify the note text, so that users can adjust the note text if the text to be processed is inaccurate.

[0129] S605: Obtain the summary text based at least on the selected text.

[0130] Optionally, obtaining at least a portion of the content of the text to be processed selected by the user in the first area, and displaying the selected content in the first area as note text in the second area, further includes: obtaining a timestamp matching each note text; wherein the timestamp is related to the speech segment matched by the note text in the speech data; using the timestamp matching the note text corresponding to the selected text as a reference timestamp, and based on at least a portion of the content in the selected text, obtaining reference text that is semantically related to the selected text from unselected text within a preset time range from the reference timestamp.

[0131] Specifically, the timestamp of the audio segment matched in the audio data for each note text is obtained, and the timestamp of each note text is saved.

[0132] Optionally, the timestamp of each note text can be displayed in the second area so that users know the actual time when each note text was generated.

[0133] Furthermore, when the selected text is the selected note text in the second area, the timestamp of the note text corresponding to the selected text is used as the reference timestamp. From the unselected text within a preset time range from the reference timestamp, content that is semantically related to at least part of the content in the selected text is searched to obtain the reference text that is semantically related to the selected text, thereby improving the efficiency of obtaining the reference text and reducing the amount of processing required to obtain the reference text.

[0134] In one specific application scenario, the preset time range is the transcribed text corresponding to the voice data within five minutes before and after the reference timestamp. In other specific application scenarios, the preset time range can also be a custom time range such as ten minutes before and after. This application does not impose any specific restrictions on this.

[0135] It is understood that, after obtaining the reference text, the step of obtaining the summary text based at least on the selected text includes: obtaining the summary text based on the selected text and the reference text.

[0136] Specifically, the method for obtaining the summary text based on the selected text and the reference text can be found in the previous embodiment, and will not be repeated here.

[0137] S606: Display the summary text.

[0138] Specifically, step S606 can be found in the relevant content of any of the above embodiments, and will not be repeated here.

[0139] Furthermore, after displaying the summary text, the process also includes: generating an index relationship between the summary text and the corresponding selected text; obtaining the speech segment of the selected text corresponding to the summary text in the speech data; and generating the speech paragraph corresponding to the summary text based on the speech segment.

[0140] Specifically, the selected text corresponding to the abstract text is obtained, and an index relationship between the abstract text and the corresponding selected text and reference text is generated so that when users select the corresponding abstract text, they can find the source of the abstract text, making it more convenient and intuitive for users to confirm whether the abstract text is accurate.

[0141] Furthermore, based on the timestamp of the selected text corresponding to the summary text, the corresponding audio segments are obtained from the audio data. The audio segments are cut and spliced ​​to generate audio segments corresponding to the summary text. This allows users to select and play the audio segments corresponding to the summary text at the summary text location, or users can share the summary text and the audio segments corresponding to the source of the summary text. This enables the recipients of the summary text and its corresponding audio segments to obtain more concise audio and protect the content in the original audio data that is not suitable for sharing.

[0142] Optionally, when the summary text also corresponds to a reference text, the process after displaying the summary text further includes: generating the index relationship between the summary text and the corresponding selected text and reference text; obtaining the speech segments of the selected text and reference text corresponding to the summary text in the speech data; and generating the speech paragraphs corresponding to the summary text based on the speech segments.

[0143] Specifically, the selected text and reference text corresponding to the abstract text are obtained, and an index relationship between the abstract text and the corresponding selected text and reference text is generated so that when users select the corresponding abstract text, they can find the source of the abstract text, making it more convenient and intuitive for users to confirm whether the abstract text is accurate.

[0144] Furthermore, based on the timestamps of the selected text and reference text corresponding to the summary text, the corresponding audio segments are obtained from the audio data. The audio segments are then cut and spliced ​​to generate audio segments corresponding to the summary text. This allows users to select and play the audio segments corresponding to the summary text at the summary text location, or users can share the summary text and the audio segments corresponding to the source of the summary text. This enables the recipients of the summary text and its corresponding audio segments to obtain more concise audio and protects the content in the original audio data that is not suitable for sharing.

[0145] Optionally, when the selected text corresponding to the summary text and the reference text match with a clearly identified speaker, a speaker prompt is added to the corresponding audio segment in the final generated audio segment so that the user can clearly know the actual speaker of the corresponding audio segment.

[0146] Unlike the previous embodiments, when user-input voice data is obtained, the voice data is converted into text to obtain the text to be processed. This adapts to scenarios where users need to extract summaries from voice data. The text to be processed is displayed in a first area, and the content selected by the user in the first area is displayed as note text in a second area. This allows users to pre-select some content from the first area as note text, which is then displayed centrally in the second area. The note text is then filtered in the second area, and the selected note text is then confirmed as the user's selected text. This allows users to compare and refer to the selected text, resulting in a more accurate selection. Consequently, when generating the summary text, a summary text that better matches the user's needs can be obtained. When the Chinese text is the selected note text in the second area, the timestamp of the note text corresponding to the selected text is used as the reference timestamp. From the unselected text within a preset time range from the reference timestamp, content that is semantically related to at least part of the content in the selected text is searched to obtain the reference text that is semantically related to the selected text. This improves the efficiency of obtaining the reference text and reduces the amount of processing required. After obtaining the summary text based on the selected text and the reference text, the summary text is post-processed to generate the index relationship between the summary text and the corresponding selected text and reference text, as well as the audio segment corresponding to the summary text. This makes it easier for users to find the source of the summary text and share the summary text and its corresponding audio segment.

[0147] Please see Figure 8 , Figure 8 This is a flowchart illustrating another embodiment of the abstract generation method of this application, the method comprising:

[0148] S801: In response to receiving voice data input by the user, convert the voice data into text to be processed.

[0149] Specifically, this step is the same as step S601 above. For details, please refer to the corresponding embodiments. This application will not repeat them here.

[0150] S802: Display the text to be processed in the third area.

[0151] Specifically, please refer to Figure 9 , Figure 9 yes Figure 8 A schematic diagram of an application scenario corresponding to step S802, wherein the third region is... Figure 9 The text to be processed will be displayed in the third area on the left side of the middle section.

[0152] S803: Get the selected text selected by the user in the text to be processed.

[0153] S804: Obtain the summary text based at least on the selected text.

[0154] Specifically, the detailed process of steps S803-S804 can be found in any of the above embodiments, and will not be repeated here.

[0155] For a specific application scenario, please refer to [link / reference]. Figure 9 The user-selected text is represented by a solid line box, while semantically related reference text is represented by a dashed line box. However, the actual display method is not limited to that described in this scenario; it is only used here for illustrative purposes. The summary text obtained based on the selected text and reference text is as follows: Figure 9 The text states: "In today's world, a new technological revolution and industrial transformation are poised to take off, the international landscape is undergoing profound changes, and the imbalance in development has not fundamentally changed. The Asia-Pacific region is the most dynamic and promising economic region in the world, and it is also a globally recognized important engine of world economic growth."

[0156] S805: Display the summary text in the third area using an editable box.

[0157] For details, please continue reading Figure 9 The system uses an editable box to display the summary text in the third area, allowing users to obtain the displayed summary text from the continuation of the text to be processed, and to add, delete, and modify the displayed summary text, thus improving the convenience of comparing the summary text and the text to be processed.

[0158] S806: In response to detecting that the user has selected the send option, the content in the editable box is displayed in the fourth area, wherein the fourth area does not overlap with the third area.

[0159] Specifically, when the user selects the send option, the content in the editable box is displayed in the fourth area, which does not overlap with the third area. This allows the officially generated summary text to be displayed separately from the text to be processed, so that the user can view all the final generated summary texts at the same time.

[0160] For a specific application scenario, please refer to Figure 10 and Figure 11 , Figure 10 yes Figure 8 A schematic diagram of an application scenario corresponding to step S806 in one embodiment. Figure 11 yes Figure 8 An application scenario diagram of another implementation method corresponding to step S806 is shown, such as... Figure 10 As shown, when the user sends an option at the location corresponding to the dashed box (i.e. Figure 10 The "Insert Notes" area will display the content from the editable box in the fourth area, such as... Figure 11 As shown, the content in the editable box will be displayed. Figure 11 The position within the dashed box, where, Figure 11 The text above the dashed box is the summary text that has been generated and sent to the fourth area. In this way, all confirmed summary texts can be displayed together in the fourth area for users to view in a unified manner.

[0161] It is understandable that, such as Figure 10 As shown, the editable box in the third area may include other options such as copy and highlight in addition to the send option, and this application does not impose specific restrictions on this.

[0162] Optionally, displaying the content in the editable box in the fourth area includes: obtaining a target summary template that matches the target semantics based on the target semantics corresponding to the content in the editable box; filling the content in the editable box into the corresponding position in the target summary template to obtain the target summary text; and displaying the target summary text in the fourth area.

[0163] Specifically, when a user selects to send content in the editable box, the target semantics corresponding to the content in the editable box are obtained, a target summary template matching the target semantics is obtained, and the content in the editable box is filled into the corresponding position in the target summary template to obtain the final target summary text. The target summary text is then displayed in the fourth area so that all target summary texts obtained by the user in the fourth area correspond to their respective target summary templates, enabling the user to distinguish target summary texts with different semantics and improving the convenience for the user to view and share target summary texts.

[0164] In a specific application scenario, the summary templates corresponding to different semantic types include at least the summary templates corresponding to the discussion content, conclusion content, and to-do content. Then, when the target semantics corresponding to the content in the editable box matches any semantic type, the content in the corresponding editable box is added to the corresponding target summary template.

[0165] Furthermore, after displaying the summary text, the process also includes: generating an index relationship between the summary text and the corresponding selected text; obtaining the speech segment of the selected text corresponding to the summary text in the speech data; and generating the speech paragraph corresponding to the summary text based on the speech segment.

[0166] Optionally, when the summary text also corresponds to a reference text, the process after displaying the summary text further includes: generating an index relationship between the summary text and the corresponding selected text and reference text; obtaining the speech segments of the selected text and reference text corresponding to the summary text in the speech data; and generating a speech paragraph corresponding to the summary text based on the speech segments. For details of the above process, please refer to the content of the previous embodiment; this application will not repeat it further.

[0167] Unlike the previous embodiments, an editable box is used to display the summary text in the third area. This allows users to obtain the displayed summary text from the continuation of the text to be processed and to add, delete, or modify the displayed summary text, improving the ease with which users can compare the summary text and the text to be processed. When the user selects the send option, the content in the editable box is displayed in the fourth area. The fourth area does not overlap with the third area, thus displaying the officially generated summary text separately from the area displaying the text to be processed. This allows users to view all the finally generated summary texts together. Furthermore, all target summary texts obtained in the fourth area correspond to their respective target summary templates, enabling users to distinguish target summary texts with different semantics and improving the ease with which users can view and share target summary texts.

[0168] Please see Figure 12 , Figure 12 This is a schematic diagram of an embodiment of the abstract generation device of this application. The abstract generation device 120 includes an acquisition module 122, a generation module 124 and a display module 126. The acquisition module 122 is used to acquire the selected text selected by the user in the text to be processed; the generation module 124 is used to obtain the abstract text based at least on the selected text; and the display module 126 is used to display the abstract text.

[0169] The above solution involves obtaining the selected text from the text to be processed after the user selects at least a portion of the text. Based on this selected text, a summary text is generated, ensuring that the summary text is at least related to the selected text. In other words, the reference range for generating the summary text includes at least the selected text. After obtaining the summary text, it is displayed. Therefore, the user only needs to view the text and select a portion, and the process of selecting the selected text satisfies the user's need for free choice. Once the user selects at least a portion of the text from the text to be processed, a summary text with a high degree of matching to the selected text is obtained, improving the convenience for the user to obtain the summary text.

[0170] Optionally, the acquisition module 122 is further configured to acquire reference text that is semantically related to the selected text from the unselected text based on at least a portion of the content in the selected text; wherein, the unselected text includes text other than the selected text in the text to be processed; the generation module is further configured to 124 obtain the summary text based on the selected text and the reference text.

[0171] Optionally, the display module 126 is further configured to display at least one first candidate in the text to be processed; wherein each first candidate corresponds to a text block; the acquisition module 122 is further configured to, in response to at least one first candidate being selected, take the content in the text blocks corresponding to all selected first candidate as the selected text.

[0172] Optionally, the acquisition module 122 is further configured to generate at least one first candidate option based on the semantics corresponding to the natural paragraphs in the text to be processed, and display all the generated first candidate options in the text to be processed; wherein, when the text block corresponding to the first candidate option includes multiple natural paragraphs, the multiple natural paragraphs in the text block are semantically related.

[0173] Optionally, the acquisition module 122 is further configured to generate at least one first candidate based on the semantics of the text content in the text to be processed and the number of characters in the text content, and to display all the generated first candidate in the text to be processed; wherein the content in the text block corresponding to the first candidate has semantic relevance, and the total number of characters in the text block corresponding to the first candidate is greater than the character count threshold.

[0174] Optionally, the acquisition module 122 is further configured to generate at least one first candidate based on the semantic groups corresponding to the keywords and key sentences in the text to be processed, and display all the generated first candidate in the text to be processed; wherein at least part of the content in the text block corresponding to the first candidate corresponds to the same semantic group.

[0175] Optionally, the acquisition module 122 is further configured to, in response to acquiring auxiliary content that matches the text to be processed, extract auxiliary text from the auxiliary content, generate at least one first candidate option based on the auxiliary text, and display all the generated first candidate options in the text to be processed; wherein the content in the text block corresponding to at least one first candidate option is related to the auxiliary text.

[0176] Optionally, the acquisition module 122 is also used to acquire extracted content from the selected text, and based on the semantics of the extracted content, acquire reference text from the unselected text that has a semantic relationship with the extracted content; and generate and display prompt information corresponding to the reference text.

[0177] Optionally, the prompt information includes at least one second candidate corresponding to the reference text; the generation module 124 is further configured to, in response to all second candidates being confirmed or cancelled, use the reference text corresponding to the confirmed second candidate as the confirmation text; and obtain the summary text based on the selected text and the confirmation text.

[0178] Optionally, the acquisition module 122 is further configured to convert the voice data into text to be processed in response to acquiring the voice data input by the user; the display module 126 is further configured to display the text to be processed in a first area; the acquisition module 122 is further configured to acquire at least a portion of the text to be processed selected by the user in the first area, and the display module 126 is further configured to display the selected content in the first area as note text in a second area; wherein, the second area does not overlap with the first area; the acquisition module 122 is further configured to acquire at least one note text selected by the user in the second area, and to display the selected note text in the second area as selected text.

[0179] Optionally, the acquisition module 122 is further configured to acquire the timestamp of each note text match; wherein the timestamp is related to the speech segment matched by the note text in the speech data; the generation module 124 is further configured to use the timestamp of the note text match corresponding to the selected text as a reference timestamp, and based on at least part of the content in the selected text, obtain the summary text from the unselected text within a preset time range away from the reference timestamp, based on the selected text and the reference text.

[0180] Optionally, the acquisition module 122 is further configured to convert the voice data into text to be processed in response to acquiring the voice data input by the user; the display module 126 is further configured to display the text to be processed in a third area; display summary text in the third area using an editable box; and display the content in the editable box in a fourth area in response to detecting that the user has selected the send option; wherein the fourth area does not overlap with the third area.

[0181] Optionally, the display module 126 is also used to obtain a target summary template that matches the target semantics based on the target semantics corresponding to the content in the editable box; fill the content in the editable box into the corresponding position in the target summary template to obtain the target summary text; and display the target summary text in the fourth area.

[0182] Optionally, the generation module 124 is further configured to generate the index relationship between the summary text and the corresponding selected text; obtain the speech segment of the selected text corresponding to the summary text in the speech data; and generate the speech paragraph corresponding to the summary text based on the speech segment.

[0183] Please see Figure 13 , Figure 13This is a schematic diagram of the structure of an embodiment of the smart terminal of this application. The smart terminal 130 includes a display screen 1300, a memory 1301, and a processor 1302. The display screen 1300 and the memory 1301 are respectively coupled to the processor 1302. The display screen 1300 is used at least to display content to the user and allow the user to select content. The memory 1301 stores program instructions (not identified). The processor 1302 is used to execute the program instructions to implement the summary generation method described in any of the above embodiments. For related explanations, please refer to the detailed description of the above method embodiments, which will not be repeated here.

[0184] The above solution can improve the ease of obtaining abstract text and the accuracy of the abstract text.

[0185] Please see Figure 14 , Figure 14 This is a schematic diagram of one embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 140 stores program instructions 1400, which, when executed by a processor, implement the digest generation method described in any of the above embodiments. For further details on related content, please refer to the detailed description of the above method embodiments; it will not be repeated here.

[0186] The above solution can improve the ease of obtaining abstract text and the accuracy of the abstract text.

[0187] It should be noted that the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0188] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0189] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0190] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for generating abstracts, characterized in that, The summary generation method includes: The selected text chosen by the user in the text to be processed is obtained; the text to be processed is obtained by converting the voice data input by the user, and at least part of the content in the text to be processed selected by the user is note text, and the note text selected by the user is the selected text. At least based on the selected text, a summary text is obtained; the summary text is obtained based on the selected text and a reference text, the reference text being obtained based on the following steps: obtaining the timestamp matching each note text; wherein, the timestamp is related to the speech segment matched by the note text in the speech data; using the timestamp matching the note text corresponding to the selected text as a reference timestamp, and based on at least a portion of the content in the selected text, obtaining a reference text that is semantically related to the selected text from unselected text within a preset time range from the reference timestamp; The summary text is displayed.

2. The method according to claim 1, characterized in that, Before obtaining the summary text based at least on the selected text, the method further includes: Based on at least a portion of the selected text, reference text that is semantically related to the selected text is obtained from the unselected text; wherein, the unselected text includes text other than the selected text in the text to be processed; The process of obtaining the summary text based at least on the selected text includes: The summary text is obtained based on the selected text and the reference text.

3. The method according to claim 1 or 2, characterized in that, Before obtaining the selected text chosen by the user in the text to be processed, the process also includes: At least one first candidate option is displayed in the text to be processed; wherein each first candidate option corresponds to a text block; The step of obtaining the selected text chosen by the user in the text to be processed includes: In response to at least one of the first candidate options being selected, the content of the text blocks corresponding to all the selected first candidate options is taken as the selected text.

4. The method according to claim 3, characterized in that, Displaying at least one first candidate option in the text to be processed includes: Based on the semantics of the natural paragraphs in the text to be processed, at least one first candidate option is generated, and all the generated first candidate options are displayed in the text to be processed. Wherein, when the text block corresponding to the first candidate includes multiple natural paragraphs, the multiple natural paragraphs in the text block are semantically related.

5. The method according to claim 3, characterized in that, Displaying at least one first candidate option in the text to be processed includes: Based on the semantics of the text content in the text to be processed and the number of characters in the text content, at least one first candidate option is generated, and all the generated first candidate options are displayed in the text to be processed. Wherein, the content in the text block corresponding to the first candidate option is semantically related, and the total number of characters in the text block corresponding to the first candidate option is greater than the character count threshold.

6. The method according to claim 3, characterized in that, Displaying at least one first candidate option in the text to be processed includes: Based on the semantic groups corresponding to the keywords and key sentences in the text to be processed, at least one first candidate option is generated, and all the generated first candidate options are displayed in the text to be processed. Among them, at least part of the content in the text block corresponding to the first candidate option corresponds to the same semantic group.

7. The method according to claim 3, characterized in that, Displaying at least one first candidate option in the text to be processed includes: In response to obtaining auxiliary content that matches the text to be processed, auxiliary text is extracted from the auxiliary content, at least one first candidate option is generated based on the auxiliary text, and all the generated first candidate options are displayed in the text to be processed; In this case, the content of at least one text block corresponding to the first candidate is related to the auxiliary text.

8. The method according to claim 2, characterized in that, The step of obtaining reference text that is semantically related to the selected text from the unselected text based on at least a portion of the selected text includes: Extracted content is obtained from the selected text, and based on the semantics of the extracted content, reference text that is semantically related to the extracted content is obtained from the unselected text. Generate and display the prompt information corresponding to the reference text.

9. The method according to claim 8, characterized in that, The prompt information includes at least one second candidate option corresponding to the reference text; After generating and displaying the prompt information corresponding to the reference text, the method further includes: In response to all second candidates being confirmed or cancelled, the reference text corresponding to the confirmed second candidate shall be used as the confirmation text; The process of obtaining the summary text based on the selected text and the reference text includes: The summary text is obtained based on the selected text and the confirmed text.

10. The method according to claim 1, characterized in that, Before obtaining the selected text chosen by the user in the text to be processed, the process also includes: In response to receiving voice data input by the user, the voice data is converted into the text to be processed; The text to be processed is displayed in the first area; At least a portion of the text to be processed selected by the user in the first area is obtained, and the selected content in the first area is displayed as note text in the second area; wherein the second area does not overlap with the first area; The step of obtaining the selected text chosen by the user in the text to be processed includes: Obtain at least one note text selected by the user in the second area, and use the selected note text in the second area as the selected text.

11. The method according to claim 1, characterized in that, Before obtaining the selected text chosen by the user in the text to be processed, the process also includes: In response to receiving voice data input by the user, the voice data is converted into the text to be processed; The text to be processed is displayed in the third area; The presentation of the summary text includes: The summary text is displayed in the third area using an editable box; In response to detecting that the user has selected the send option, the content of the editable box is displayed in the fourth area; wherein the fourth area does not overlap with the third area.

12. The method according to claim 11, characterized in that, The step of displaying the content of the editable box to the fourth area includes: Based on the target semantics corresponding to the content in the editable box, obtain a target summary template that matches the target semantics; Fill the content in the editable box into the corresponding position in the target summary template to obtain the target summary text; The target summary text is displayed in the fourth area.

13. The method according to claim 9 or 11, characterized in that, After displaying the summary text, the method further includes: Generate the index relationship between the summary text and the corresponding selected text; Obtain the audio segment of the selected text corresponding to the summary text in the audio data; Based on the audio segment, the audio paragraph corresponding to the summary text is generated.

14. A summary generation apparatus, characterized in that, The summary generation device includes: The acquisition module is used to acquire the selected text chosen by the user in the text to be processed; the text to be processed is obtained by converting the voice data input by the user, and at least part of the content in the text to be processed selected by the user is note text, and the note text selected by the user is the selected text; A generation module is used to obtain summary text based at least on the selected text; the summary text is obtained based on the selected text and reference text, and the reference text is obtained based on the following steps: obtaining a timestamp matching each note text; wherein, the timestamp is related to the speech segment matched by the note text in the speech data; using the timestamp matching the note text corresponding to the selected text as a reference timestamp, and based on at least a portion of the content in the selected text, obtaining reference text that is semantically related to the selected text from unselected text within a preset time range away from the reference timestamp; The display module is used to display the summary text.

15. A smart terminal, characterized in that, include: The device includes a display screen, a memory, and a processor, wherein the display screen and the memory are respectively coupled to the processor, the display screen is used at least to display content to a user and to allow the user to select content, the memory stores program instructions, and the processor is used to execute the program instructions to implement the summary generation method as described in any one of claims 1 to 13.

16. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the summary generation method as described in any one of claims 1-13.

Citation Information

Patent Citations

  • Text summary generation method and device, equipment and storage medium

    CN114328899A

  • Interactive document reading

    WO2005103946A1