Information summarization method and device based on specific task, equipment and medium
By training a small model to extract fine-grained information for specific tasks and combining it with a large model for information summarization, it solves the problems of domain adaptability, text information fragmentation, computing resources, and real-time performance, and generates high-quality information summaries suitable for specific tasks of processing long texts.
Patent Information
- Application Number
- CN202510611570.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Existing information summarization technologies have shortcomings in domain adaptability, text information fragmentation, long text processing capabilities, computing resources and real-time performance. Especially when processing patents, legal documents and other fields, it is difficult to effectively capture specific task information, resulting in information redundancy, omissions and excessive consumption of computing resources.
By training a small model to extract the fine-grained information of a specific task with information labels for a specific task, and using regular expressions and sentence integrity judgment models, combined with the method of a large model, the fine-grained information phrases are matched with the original text, and the large model is combined to summarize the information and generate high-quality summary content.
Through the collaborative work of small and large models, key information for specific tasks can be quickly extracted, redundant information and information loss can be reduced, and processing efficiency and computing resource utilization can be improved.
Smart Images

Figure CN120633649A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a method, device, equipment and medium for summarizing information based on a specific task. Background Art
[0002] Information summarization is an important technology in natural language processing. It aims to extract or summarize content that can summarize the main information of the entire text or a specific task from a large amount of text, and express it in concise, accurate and clear language. It can be used for medical report analysis, legal document processing, financial data interpretation, patent document abstract generation, etc.
[0003] Existing information summarization techniques can generally be divided into two categories: extractive summarization and generative summarization.
[0004] Extractive summarization selects the most important sentences, paragraphs, or phrases from the original text and directly concatenates them into a summary without changing the original text's expression. Traditional keyword extraction methods, such as TF-IDF (Term Frequency-Inverse Document Frequency) and TextRank, extract key information by calculating the weight of words and the importance of sentences, or by using deep neural networks, such as recurrent neural networks (RNNs), to automatically learn the structure and content of the text and extract key information. Although the results do not result in information loss or inaccuracy, they cannot effectively summarize long content. The extracted information lacks coherence and is prone to information redundancy.
[0005] Generative summarization uses deep learning models to reorganize language and generate new summaries based on understanding the original text content. Generative summarization typically relies on neural network models (such as Transformer, BERT, GPT, etc.), using an encoder-decoder architecture to understand the semantic relationships of the text and generate more fluent and natural summaries. Its advantages are that it can better summarize long content, avoid information redundancy, and make the summarized information more coherent and consistent with human expression habits. Compared with extractive summarization, generative summarization can rewrite sentence structure, improve text readability, and support customized summary styles for different scenarios. However, because the text is generated after understanding, it may introduce certain information biases or false information. It has certain limitations in processing very long texts and consumes a lot of computing resources.
[0006] Therefore, although many effective methods have been proposed for information summarization tasks in the field of natural language processing, information summarization still faces the following challenges in practical applications:
[0007] Insufficient domain adaptability: Most pre-trained models are trained primarily on general corpora, making it difficult to capture the professional terminology, technical details, and inherent logic in specific fields (such as patent texts and legal documents). Directly using pre-trained models to summarize information for specific tasks will not yield and generate relevant information that meets the specific task.
[0008] Text information is fragmented and redundant: During implementation, relevant phrases for the same specific task (such as technical efficacy and technical issues in patent texts, and judgment results and basis in legal documents) may appear in different locations in the original text. For example, keywords for technical issues and technical effects may appear in multiple places in a patent or paper. The extractive summary model will extract all of this content when performing the task, which may result in repeated information appearing multiple times and causing information redundancy. If these contents are directly presented as summary results, the true purpose of information summary cannot be achieved.
[0009] Insufficient processing capabilities for long texts: Whether using extractive summarization based on small models or generative summarization based on large models, processing long texts is a challenging task, especially in fields such as medicine, law, and patents, where the text content is both large and complex. When summarizing long texts, the model may not be able to effectively capture the context of the entire article, resulting in the omission of important information or loss of contextual relationships. Furthermore, processing extremely long texts places high demands on computing resources.
[0010] The conflict between computing resources and real-time performance: Generative summarization techniques typically utilize models with large parameter counts, which typically require significant computing resources. Large models, in particular, require significant amounts of video memory and computing power during inference, making real-time processing a significant bottleneck. Furthermore, in some high-concurrency, low-latency application scenarios, summarization methods based on large models are slow and cannot meet the real-time requirements of the system. If real-time performance issues were addressed by processing data offline, the time cost of processing the data offline would be significant, as would the cost of updating the model and data later.
[0011] The quality of generation is difficult to control: The results obtained using extractive summarization technology lack coherence and poor readability. Generative summarization optimizes this deficiency to a certain extent, but it may generate overly general summary information or repeatedly generate the same content. It has weak control over details, especially in the extraction of key information about specific tasks. There may be omissions, which affects the conciseness of the summary and the accuracy of the information. It may also introduce some biased or false information. Summary of the Invention
[0012] In order to overcome the problems existing in the related art, the present disclosure provides a method, device, equipment and medium for summarizing information based on a specific task to solve the technical problems in the related art.
[0013] One or more embodiments of this specification provide a method for summarizing information based on a specific task, including the steps of:
[0014] Obtain pre-collected natural language text data, label it with specific task labels, and obtain labeled training data;
[0015] Input the training data into the small model until the model converges, obtaining a small model for extracting fine-grained information about a specific task;
[0016] Input the text to be extracted into the trained small model to extract fine-grained information phrases;
[0017] The obtained fine-grained information phrase is matched with the text of the information to be extracted through a matching algorithm, thereby obtaining a target sentence containing the fine-grained information phrase;
[0018] The target sentence is input into the big model, and information summary is performed according to the preset prompt words related to the specific task to obtain summary information about the specific task.
[0019] Furthermore, the small model selects an information extraction model, and performs fine-grained information learning and extraction training for a specific task based on the specified number of training iterations, batch size, learning rate, and optimizer through labeled training data to obtain a small model for extracting fine-grained information about a specific task.
[0020] Furthermore, the method further includes the following steps: based on the fine-grained information extraction result of the small model, if the fine-grained information phrase is not obtained, the text of the information to be extracted is input into the large model to perform information summarization.
[0021] Furthermore, the obtained fine-grained information phrase is matched with the text of the information to be extracted through a matching algorithm to obtain a target sentence containing the fine-grained information phrase, which specifically includes the following steps:
[0022] Based on the obtained fine-grained information phrases, a target sentence containing the fine-grained information phrases is obtained from the original text through regular expression matching;
[0023] Based on the obtained target sentence, the sentence completeness judgment model is used to determine whether the target sentence contains the subject and action continuation participles. If so, the target sentence is retained. Otherwise, the target sentence is merged with the sentence in front of it and then input into the sentence completeness judgment model for judgment process.
[0024] Furthermore, the steps include:
[0025] For the target sentences obtained, if the number of target sentences exceeds the preset value, they are spliced in sequence according to the position order of the target sentences in the original text and then input into the large model; otherwise, they are spliced in sequence according to the position order of the target sentences in the original text and output as the final summary result.
[0026] One or more embodiments of this specification provide an information device based on a specific task, including:
[0027] The data processing module is used to obtain pre-collected natural language text data, label it with specific task labels, and obtain labeled training data;
[0028] A model training module is used to input training data into the small model until the model converges, thereby obtaining a small model for extracting fine-grained information about a specific task;
[0029] The fine-grained information module is used to input the text of the information to be extracted into the trained small model to extract fine-grained information phrases;
[0030] A sentence matching module is used to match the fine-grained information phrase obtained by the fine-grained information module with the text of the information to be extracted through a matching algorithm, thereby obtaining a target sentence containing the fine-grained information phrase;
[0031] The information summarization module is used to input the target sentence into the large model, perform information summary according to the preset prompt words related to the specific task, and obtain summary information about the specific task.
[0032] Furthermore, the clause matching module includes a matching submodule, a clause analysis submodule and a clause expansion submodule;
[0033] A matching submodule is used to obtain a target sentence containing a fine-grained information phrase from the original text by matching the obtained fine-grained information phrase through a regular expression;
[0034] The clause analysis submodule is used to determine whether the target clause contains subject and action continuation participles based on the obtained target clause using the sentence integrity judgment model. If so, the target clause is retained; otherwise, the clause expansion submodule is called;
[0035] The clause expansion submodule is used to merge the target clause with the clause in front of it and send it to the clause analysis submodule.
[0036] Furthermore, it also includes a sentence processing module, which is used to judge based on the obtained target sentences. If the number of target sentences exceeds a preset value, the target sentences are spliced in sequence according to the position order of each target sentence in the original text and then input into the large model; otherwise, the target sentences are spliced in sequence according to the position order of the target sentences in the original text and output as the final summary result.
[0037] One or more embodiments of this specification provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described task-based information summarization methods when executing the computer program.
[0038] One or more embodiments of this specification provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-described information summarization methods based on specific tasks.
[0039] The present disclosure provides a method, device, equipment and medium for information summarization based on a specific task, which has the advantage of training a small model with text data marked with a specific task label to obtain a model that can extract fine-grained information about the specific task, and then using the fine-grained information extracted by the small model as keywords to match the original text to obtain corresponding sentences, which are then input into the large model for information summarization; this embodiment takes advantage of the small model in processing efficiency, can quickly extract key information of fine-grained specific tasks with low computing resource consumption, and can ensure the extraction of core information from a wide and comprehensive perspective at the fine-grained level, while the large model, based on the broad and comprehensive results obtained by the small model extraction, deeply understands and summarizes the information with high quality from a global perspective, integrates some redundant information, rewrites the sentence structure, and generates summary content with higher readability; through the collaborative work of the two models, it can ensure the semantic accuracy of the summary information while improving processing efficiency and effectively reducing redundant information and information loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0041] Figure 1 A flowchart of a method for summarizing information based on a specific task provided for one or more embodiments of this specification;
[0042] Figure 2A block diagram of an information summarization device based on a specific task provided in one or more embodiments of this specification;
[0043] Figure 3 A schematic diagram of the structure of a computer device provided in one or more embodiments of this specification. DETAILED DESCRIPTION
[0044] In order to help those skilled in the art better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0045] The present invention will be described in detail below with reference to specific implementation methods and the accompanying drawings.
[0046] Method Example
[0047] According to an embodiment of the present invention, a method for summarizing information based on a specific task is provided, such as Figure 1 FIG. 1 is a flow chart of a method for summarizing information based on a specific task provided by this embodiment. The method for summarizing information based on a specific task according to an embodiment of the present invention includes the following steps:
[0048] Step S1: Obtain pre-collected natural language text data, mark it with specific task labels, and obtain labeled training data.
[0049] In step S2, the training data is input into the small model until the model converges, thereby obtaining a small model for extracting fine-grained information about a specific task.
[0050] Step S3: Input the text of the information to be extracted into the trained small model to extract fine-grained information phrases.
[0051] Step S4: Matching the obtained fine-grained information phrase with the text of the information to be extracted through a matching algorithm to obtain a target sentence containing the fine-grained information phrase.
[0052] In step S5, the target sentence is input into the large model, and information summarization is performed according to preset prompt words related to the specific task to obtain summary information about the specific task.
[0053] This embodiment provides a task-specific information summarization method. To obtain specific information from a text, a small model is trained on text data labeled with a specific task tag to obtain a model capable of extracting fine-grained information about the specific task. During the use of the small model, because the fine-grained information extracted by the small model does not reflect the complete meaning of the sentence, the content summarized by the large model based on the extracted fine-grained information deviates significantly from the semantics of the original text. In this embodiment, the fine-grained information extracted by the small model is used as keywords to match the original text to obtain corresponding sentences, which are then input into the large model for information summarization. This embodiment takes advantage of the processing efficiency of the small model, can quickly extract key information of the fine-grained specific task with low computing resource consumption, and can also ensure the comprehensive and extensive extraction of core information at the fine-grained level. Based on the comprehensive results extracted by the small model, the large model deeply understands and summarizes the information from a global perspective, integrates some redundant information, rewrites the sentence structure, and generates a more readable summary. The collaborative work of the two models can ensure the semantic accuracy of the summary information while also improving processing efficiency and effectively reducing redundant information and information loss.
[0054] In this embodiment, the acquired natural language text data is also preprocessed, and the preprocessing may include cleaning and denoising.
[0055] First, remove unnecessary formatting information to ensure only the core text content is retained. For example, a combination of regular expressions and an HTML parsing library (BeautifulSoup) can be used to remove HTML tags. Special characters such as , <, etc. can be processed using Python's HTML parsing library, ultimately returning the clean text after preprocessing.
[0056] Furthermore, in step S1, the specific steps of marking a specific task label are as follows:
[0057] Use data annotation tools to label task-related phrases in the preprocessed text with predefined labels to obtain training data.
[0058] In one embodiment, for example, in the patent field, the specific task is to extract beneficial effects from the patent application text, then it is necessary to manually annotate the technical effect phrases in the patent application text, such as "eliminating discomfort" and "preventing bumps" and other technical effect phrases; the label name can be customized according to your own task, such as adding the "technical_effect" label. After the annotation is completed, the annotated data can be divided into a training set and a test set for training the model, wherein text annotation can be implemented using the doccano tool.
[0059] In this embodiment, the small model can select an information extraction model, such as the UIE (Unified Information Extraction) model, the BERT (Bidirectional Encoder Representations from Transformers) model, and perform fine-grained information learning and extraction training for a specific task based on the specified number of training iterations, batch size, learning rate, and optimizer, thereby obtaining a small model for extracting fine-grained information about a specific task.
[0060] In one embodiment, the following steps are further included:
[0061] Step S31: Based on the fine-grained information extraction results of the small model, determine whether a fine-grained information phrase is obtained. If not, input the text of the information to be extracted into the large model, and perform information summary in combination with the pre-set prompt words to obtain the information summary result; if a fine-grained information phrase is obtained, execute step S4.
[0062] This embodiment is preferred because the small model extracts fine-grained information of a specific task, that is, phrases related to the specific task. The phrases may not contain subject information and the sentence meaning is incomplete. Therefore, the semantic information expressed is incomplete, which makes the semantics likely to deviate from the original context. The information extracted is provided to the large model for understanding and summary, which will cause information semantic deviation and context information loss. For example, the current specific task is to extract technical effects from the text. For example, the extracted relevant fine-grained information is "ensuring high safety", etc., but the original text is "front bumper mounting structure, which can effectively reduce the deformation of the vehicle body during offset collision and ensure high safety". Therefore, it can be seen that the fine-grained information has lost the entity that realizes the technical effect in the original text. Therefore, in order to improve the accuracy of the summarized information, it is necessary to obtain the original text corresponding to the fine-grained information and provide it to the large model to achieve information summary. In this way, the original context of the fine-grained information is retained, which greatly reduces information loss. Therefore, step S4 specifically includes the following steps:
[0063] Step S41 : Based on the obtained fine-grained information phrase, a target sentence containing the fine-grained information phrase is matched from the original text by using a regular expression.
[0064] In this embodiment, the original sentence containing the fine-grained information phrase is obtained through regular expression matching. Compared with the similarity matching method, similarity matching requires vectorization of both the original text and the extracted phrase. This process is time-consuming, resulting in slow matching speed and matching of a lot of redundant information.
[0065] Step S42: Based on the acquired target sentence, use the sentence integrity judgment model to determine whether the target sentence contains the subject and action continuation participles, that is, determine whether the sentence meaning of the target sentence is complete. If it does, retain the target sentence, otherwise execute step S43.
[0066] Step S43: merge the target clause with the clause in front of it, and go to step S42.
[0067] In this embodiment, the sentence completeness judgment model can use an LLM (Large Language Model) model, combined with a sentence completeness judgment prompt word method to enable the LLM model to judge whether the sentence is complete; this technology is a mature technology well known to those skilled in the art and will not be described in detail here.
[0068] In another embodiment, the sentence completeness judgment model can also be a sequence classification model fine-tuned based on training data. The specific fine-tuning process is to obtain a certain amount of training text corpus, use the doccano tool to perform a binary "complete / incomplete" labeling on the corpus to ensure the data quality of model training, and then select a model specifically for sequence classification, such as BERT, RoBERTa model, etc., to fine-tune the model based on the sentence completeness judgment task, and use indicators such as precision (Precision), recall (Recall) and F1 score to evaluate the performance of the fine-tuned model. After fine-tuning the model, the sentence completeness judgment model is obtained. In this embodiment, for the obtained target sentence, the execution subject of the technical effect may not be in the matching sentence, but based on the expression method of Chinese, the execution subject is generally in the previous sentence. Therefore, the semantic integrity of the extracted sentence is achieved through the above expansion.
[0069] In one example, the obtained fine-grained information phrases are as follows:
[0070] "Reducing deformation of the upper body"
[0071] “Effectively reduce deformation of the upper body”
[0072] "Ensuring high security"
[0073] “It can enhance its rigidity”
[0074] "Deformation near the cutouts a and b can be suppressed"
[0075] "It simplifies assembly tools and work"
[0076] "Prevents damage, etc."
[0077] "It can effectively reduce the deformation of the upper body"
[0078] “It has a significant impact on occupant protection.”
[0079] Furthermore, the target sentence obtained in step S42 is as follows:
[0080] "By transmitting and dispersing it to the vehicle body, it effectively absorbs the impact of the vehicle body and reduces deformation of the upper body."
[0081] "A front bumper mounting structure that can effectively reduce deformation of the upper body during offset collisions, ensuring high safety"
[0082] "Furthermore, by forming the flange into a thick plate portion, its rigidity can be enhanced and deformation of the flange portion can be suppressed."
[0083] "In this case, by cutting only a certain portion of the front bumper reinforcement, deformation near the cut portions a and b can be suppressed."
[0084] "When the bolts are tightened, the front bumper reinforcement can be assembled without a jig or the like to support it, simplifying assembly tools and work."
[0085] "It can prevent damage, etc., and can effectively reduce deformation of the upper body"
[0086] "It ensures high safety in offset collisions and has a significant effect on occupant protection."
[0087] Therefore, it can be seen that some target clauses do not contain the execution subject for realizing the technical effect. For example, the target clause "can prevent damage, etc., and can effectively reduce the deformation of the upper vehicle body" has an action-taking object, but no execution subject, resulting in incomplete semantics. Therefore, the above expansion is implemented through step S43 to ensure the semantic integrity of the target clause.
[0088] In this embodiment, the target sentences containing fine-grained information phrases obtained through the matching algorithm may be numerous. If these sentences are directly spliced and presented to the reader, the purpose of information summary cannot be achieved. Therefore, in step S5, the large model performs information summary based on the obtained target sentences. Before the target sentences are input into the large model, the following steps are also included:
[0089] Step A1, based on the target sentences determined in step S4 or step S42, determine whether the number of target sentences exceeds a preset value. If so, execute step A2; otherwise, go to step A3;
[0090] Step A2: splice the acquired target sentences according to their position order in the original text to obtain a spliced text, and input it into the large model.
[0091] In step A3, the obtained target sentences are sequentially spliced according to their positions in the original text and output as the final summary result.
[0092] Through the above steps, the number of target sentences obtained is judged to determine whether they need to be input into the large model for information summary. If the number does not exceed the preset value, they are directly output. If the number of target sentences is insufficient, there is no need to call the large model for further information summary, which saves computer computing power. Since the contents of the target sentences are the most critical content related to specific tasks extracted by the small model, there is not much redundancy in directly outputting them to readers. Only when the number of information target sentences exceeds a certain number, information summary is achieved through the large model to achieve better information acquisition effect.
[0093] In this embodiment, the target sentence quantity determination threshold may be set to 3-10.
[0094] In this embodiment, the large model combines preset prompt words corresponding to specific tasks, summarizes the input information, and returns it to the user.
[0095] In one example, assuming that the current specific task is to summarize the technical effects in the patent application text, the corresponding prompt words are as follows:
[0096] "Please extract and summarize the technical advantages of the following text information. The technical advantages should be described concisely and coherently, highlighting the advantages, innovations, and specific effects of the invention compared to the existing technology. Avoid listing bullet points. The output should be a smooth paragraph that clearly conveys improvements in performance, efficiency, accuracy, etc. Please enclose the summarized technical advantages in [ ]."
[0097] Sample input:\n
[0098] {}\n
[0099] Sample output:
[0100] This invention introduces new XX technology, significantly resolving XX problems in existing technologies. It not only significantly increases processing speed and reduces processing time, but also optimizes the XX process, improving system accuracy and stability while also reducing energy consumption and increasing overall system reliability, providing a more efficient and environmentally friendly solution for practical applications.
[0101] In this embodiment, the large model can use models such as GLM (generalize linear model), DeepSeek, GPT (Generative Pre-trained Transformer), and LLaMA (Large Language Model Meta AI).
[0102] The task-specific information summarization method provided in this embodiment receives raw text as input. The raw text can be any document requiring information summary, such as patent documents, legal contracts, medical reports, and the text content can be of any length. After preprocessing the text, fine-grained information extraction is performed using a small model fine-tuned for a specific extraction task (e.g., technical issues, technical effects, case decision basis, treatment plan and treatment basis in medical cases). The small model trained for the specific task can identify fine-grained information phrases for the specific task. To obtain complete semantic information text, the raw text is further analyzed based on this fine-grained information. Regular expressions are used to match target sentences containing fine-grained information phrases from the original text. A sentence integrity judgment model is then used to determine whether the target sentence contains the subject of the action performer, ensuring the integrity of the overall information expression. By accurately matching the sentences in the original text, the system accurately obtains a relatively complete original text expression. If the target sentences exceed the preset number, the sentences are sequentially spliced into paragraphs and passed to the large model for information summarization. If the target sentences do not exceed the preset number, the target sentences are spliced in the order in which they appear in the original text as the final result. The method of this embodiment combines the advantages of small models and large models, aiming to improve the efficiency and quality of information summarization tasks. It is particularly suitable for scenarios that require processing massive, large-scale text data or semantically complex text data, such as medical report analysis, legal document processing, financial data interpretation, and patent document abstract generation. By rationally utilizing the small model's efficient data processing capabilities and the large model's powerful semantic understanding and generation capabilities, the present invention can significantly improve processing speed and reduce computing resource consumption while ensuring the accuracy of the semantic information of the summary information. It is widely used in various information processing and intelligent document analysis tasks.
[0103] Device embodiment
[0104] According to an embodiment of the present invention, a device for summarizing information based on a specific task is provided, such as Figure 2 FIG. 1 is a block diagram of a task-specific information summarization device according to an embodiment of the present invention. The task-specific information summarization device according to an embodiment of the present invention includes:
[0105] The data processing module 10 is used to obtain pre-collected natural language text data, mark specific task labels, and obtain labeled training data.
[0106] The model training module 20 is used to input training data into the small model until the model converges, thereby obtaining a small model for extracting fine-grained information about a specific task.
[0107] The fine-grained information module 30 is used to input the text of the information to be extracted into the trained small model to extract fine-grained information phrases.
[0108] The sentence matching module 40 is configured to match the fine-grained information phrases obtained by the fine-grained information module 30 with the text of the information to be extracted through a matching algorithm to obtain a target sentence containing the fine-grained information phrases.
[0109] The information summarizing module 50 is used to input the target sentence into the large model, perform information summarization according to preset prompt words related to the specific task, and obtain summary information about the specific task.
[0110] The information summarization device based on specific tasks provided in this embodiment uses text data marked with specific task tags through the data processing module 10 to train a small model in order to obtain specific information in the text, and obtains a model that can extract fine-grained information about the specific task. During the use of the small model, since the fine-grained information extracted by the small model does not reflect the complete sentence meaning, the content summarized by the large model based on the extracted fine-grained information deviates too much from the semantics of the original text. In this embodiment, the fine-grained information extracted by the small model is used as a keyword to match the original text, and the corresponding sentence is obtained through the sentence matching module 40, and then input into the information summarization module 50. The information summarization module 50 summarizes the information through the large model. This embodiment takes advantage of the small model's processing efficiency, enabling it to quickly extract fine-grained key information for specific tasks with low computing resource consumption, while ensuring broad and comprehensive extraction of core information from a fine-grained level. The large model, based on the broad and comprehensive results obtained by the small model extraction, conducts in-depth understanding and high-quality summarization of information from a global perspective, integrates some redundant information, rewrites sentence structures, and generates highly readable summary content. Through the collaborative work of the two models, while ensuring the semantic accuracy of the summary information, it can also improve processing efficiency and effectively reduce redundant information and information loss.
[0111] This embodiment also includes a text preprocessing module for preprocessing the acquired natural language text data. The preprocessing may include cleaning and denoising. In this embodiment, the small model can select an information extraction model, such as the UIE model or the BERT model. Using labeled training data, the small model performs task-specific fine-grained information learning and extraction training based on a specified number of training iterations, batch size, learning rate, and optimizer, thereby obtaining a small model for extracting fine-grained information about the specific task.
[0112] In one embodiment, a first judgment execution module is further included, which is used to determine whether a fine-grained information phrase is obtained after completing the fine-grained information extraction based on the small model. If not, the text of the information to be extracted is input into the large model, and the information summary is performed in combination with the pre-set prompt words to obtain the information summary result. If a fine-grained information phrase is obtained, the sentence matching module 40 is called.
[0113] In this embodiment, the small model extracts fine-grained information of a specific task, that is, phrases related to the specific task. The phrases may not contain subject information and the meaning of the sentence is incomplete. Therefore, the semantic information expressed is incomplete, resulting in the semantics likely to deviate from the original context. The information extracted is provided to the large model for understanding and summarization, which will cause information semantic deviation and loss of context information. Therefore, the sentence matching module 40 is configured to include a matching submodule, a sentence analysis submodule, and a sentence expansion submodule.
[0114] The matching submodule is used to obtain a target sentence containing the fine-grained information phrase from the original text through a regular expression based on the obtained fine-grained information phrase.
[0115] This embodiment obtains the original sentence containing the fine-grained information phrase through regular expression matching. Compared with the similarity matching method, similarity matching requires vectorization of both the original text and the extracted phrase. This process is time-consuming, resulting in slow matching speed and matching a lot of redundant information.
[0116] The clause analysis submodule is used to determine whether the target clause contains the subject and action continuation participles based on the obtained target clause using the sentence integrity judgment model, that is, to determine whether the sentence meaning of the target clause is complete. If it does, the target clause is retained; otherwise, the clause expansion submodule is called;
[0117] The clause expansion submodule is used to merge the target clause with the clause in front of it and send it to the clause analysis submodule.
[0118] In this embodiment, there may be many target sentences containing fine-grained information phrases in the text obtained by matching through a matching algorithm. If these sentences are directly spliced and presented to the reader, the purpose of information summary cannot be achieved. Therefore, a sentence processing module is also configured to realize the judgment and processing of the number of target sentences. The specific configuration is used to: if it is judged that the number of target sentences exceeds the preset value, the target sentences are spliced in sequence according to the position order of the original text and then input into the large model; otherwise, the target sentences are spliced in sequence according to the position order of the original text as the final summary result output.
[0119] The embodiment of the present invention is an apparatus embodiment corresponding to the above-mentioned method embodiment. The specific operations of the processing steps of each module can be understood by referring to the description of the method embodiment, and will not be repeated here.
[0120] like Figure 3 As shown, the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the information summarization method based on a specific task in the above embodiment is implemented.
[0121] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the information summarizing method based on a specific task in the above-mentioned embodiment, or implements the information summarizing method based on a specific task in the above-mentioned embodiment when the computer program is executed by a processor.
[0122] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0123] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person of ordinary skill in the art can understand and implement it without making any creative efforts.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and the contents not described in detail in the specification of the present invention are common knowledge to those skilled in the art.
Claims
1. A task-specific information summarization method, characterized in that: Including steps: Obtain pre-collected natural language text data, label it with specific task labels, and obtain labeled training data; Input the training data into the small model until the model converges, obtaining a small model for extracting fine-grained information about a specific task; Input the text to be extracted into the trained small model to extract fine-grained information phrases; The obtained fine-grained information phrase is matched with the text of the information to be extracted through a matching algorithm, thereby obtaining a target sentence containing the fine-grained information phrase; The target sentence is input into the big model, and information summary is performed according to the preset prompt words related to the specific task to obtain summary information about the specific task.
2. The task-based information summarization method according to claim 1, wherein: The small model selects an information extraction model, and performs fine-grained information learning and extraction training for a specific task based on the specified number of training iterations, batch size, learning rate, and optimizer through labeled training data to obtain a small model for extracting fine-grained information about a specific task.
3. The task-based information summarization method according to claim 1, wherein: Also includes the steps: Based on the fine-grained information extraction results of the small model, if no fine-grained information phrases are obtained, the text of the information to be extracted is input into the large model to perform information summarization.
4. The task-based information summarization method according to claim 1, wherein: The obtained fine-grained information phrase is matched with the text of the information to be extracted through a matching algorithm to obtain a target sentence containing the fine-grained information phrase, which specifically includes the following steps: Based on the obtained fine-grained information phrases, a target sentence containing the fine-grained information phrases is obtained from the original text through regular expression matching; Based on the obtained target sentence, the sentence completeness judgment model is used to determine whether the target sentence contains the subject and action continuation participles. If so, the target sentence is retained. Otherwise, the target sentence is merged with the sentence in front of it and then input into the sentence completeness judgment model for judgment process.
5. The task-based information summarization method according to claim 1 or 4, characterized in that: Also includes the steps: For the target sentences obtained, if the number of target sentences exceeds the preset value, they are spliced in sequence according to the position order of the target sentences in the original text and then input into the large model; otherwise, they are spliced in sequence according to the position order of the target sentences in the original text and output as the final summary result.
6. An information summarization device based on a specific task, characterized in that: include: The data processing module is used to obtain pre-collected natural language text data, label it with specific task labels, and obtain labeled training data; A model training module is used to input training data into the small model until the model converges, thereby obtaining a small model for extracting fine-grained information about a specific task; The fine-grained information module is used to input the text of the information to be extracted into the trained small model to extract fine-grained information phrases; A sentence matching module is used to match the obtained fine-grained information phrase with the text of the information to be extracted through a matching algorithm, thereby obtaining a target sentence containing the fine-grained information phrase; The information summarization module is used to input the target sentence into the large model, perform information summary according to the preset prompt words related to the specific task, and obtain summary information about the specific task.
7. The information summarization device based on a specific task according to claim 6, characterized in that: The clause matching module includes a matching submodule, a clause analysis submodule and a clause expansion submodule; A matching submodule is used to obtain a target sentence containing a fine-grained information phrase from the original text by matching the obtained fine-grained information phrase through a regular expression; The clause analysis submodule is used to determine whether the target clause contains subject and action continuation participles based on the obtained target clause using the sentence integrity judgment model. If so, the target clause is retained; otherwise, the clause expansion submodule is called; The clause expansion submodule is used to merge the target clause with the clause in front of it and send it to the clause analysis submodule.
8. The task-specific information summarization device according to claim 6 or 7, characterized in that: It also includes a sentence processing module, which is used to determine based on the obtained target sentences that if the number of target sentences exceeds a preset value, the target sentences are spliced in sequence according to the position order of each target sentence in the original text and then input into the large model; otherwise, the target sentences are spliced in sequence according to the position order of the target sentences in the original text and output as the final summary result.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for summarizing information based on a specific task according to any one of claims 1 to 5 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for summarizing information based on a specific task according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Long text matching method combining noise filtering and divide-and-conquer strategy
CN117216189A
Text generation method and device, electronic equipment and readable storage medium
CN119226455A
Prompt-based ESG report text analysis method and system
CN119441474A
Document abstract generation method and device, storage medium and electronic device
CN119739854A
Document information extraction method and system, and electronic device and storage medium
WO2025060686A1