Abstract generation method and device combined with large model technology

By combining large-scale model technology with question-answer pairs and knowledge graphs to optimize summary generation, the problem of poor logical structure in summaries in the coal mining field has been solved, thus improving the quality and efficiency of summaries.

CN121980025APending Publication Date: 2026-05-05CCTEG COAL MINING RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CCTEG COAL MINING RES INST
Filing Date
2025-12-30
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional automatic summary generation technology has poor logic in the field of coal mining, and cannot deeply mine professional terms and logic, resulting in a lot of redundancy and errors in the summary.

Method used

The large model technique is used to generate the original summary by generating a large model from multiple summaries. The high-scoring backup summaries are selected to generate a large model. The large model is then adjusted by combining question-answer pairs and knowledge graphs to optimize the generation of the target summary.

Benefits of technology

It significantly improves the quality and efficiency of scientific and technological literature abstracts in the coal industry, and the logic is more in line with the needs of the vertical field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121980025A_ABST
    Figure CN121980025A_ABST
Patent Text Reader

Abstract

The invention provides an abstract generation method and device combined with a large model technology, and the method comprises the steps: determining a plurality of abstract generation large models which are used for generating original abstracts of scientific and technical literatures of the coal industry; selecting a corresponding standby abstract generation large model when the score is greater than a set threshold value from the original abstracts generated by the plurality of abstract generation large models; obtaining question and answer pairs of the scientific and technical literature, and adjusting the standby abstract generation large model through the question and answer pairs to obtain a target abstract generation large model; obtaining the target abstract to generate a corresponding standby abstract when the large model processes the scientific and technical literature of the coal industry; and adjusting the standby abstract based on a knowledge graph corresponding to the scientific and technical literature to obtain a target abstract. Therefore, the abstract is continuously optimized and is close to the vertical field of coal mining through the adjustment of the question and answer pair on the backup abstract generation large model and the adjustment of the knowledge graph on the backup abstract, and the quality and efficiency of the scientific and technical literature abstract are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of abstract generation technology, and in particular to an abstract generation method and apparatus that combines large model technology. Background Technology

[0002] In the digital age, automatic summary generation technology for knowledge bases is a mature natural language processing technology. Traditional automatic summary generation technology is divided into two types: extractive and generative text summarization. It identifies the core sentences and phrases in the text and automatically generates summary text. However, in the field of coal mining, due to the highly specialized nature of the vertical field, the summaries generated by traditional automatic summary technology for technical reports have extremely poor logic and cannot deeply mine the professional vocabulary and logic of scientific and technological literature, resulting in a lot of redundancy and errors in the summaries. Summary of the Invention

[0003] This disclosure provides a method and apparatus for abstract generation combining large model techniques, to at least partially solve one of the technical problems in related technologies. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a summary generation method combining large model technology is provided. The method includes: determining a plurality of summary generation large models, the summary generation large models being used to generate original summaries of scientific and technological literature in the coal industry; selecting backup summary generation large models from the original summaries generated by the plurality of summary generation large models whose scores are greater than a set threshold; obtaining question-answer pairs of the scientific and technological literature, and adjusting the backup summary generation large models through the question-answer pairs to obtain a target summary generation large model; obtaining backup summaries corresponding to the target summary generation large model when processing the scientific and technological literature in the coal industry; and adjusting the backup summaries based on the knowledge graph corresponding to the scientific and technological literature to obtain the target summary.

[0004] According to a second aspect of the present disclosure, a summary generation apparatus combining large model technology is provided. The apparatus includes: a determining module, configured to determine multiple large summary generation models, the large summary generation models being used to generate original summaries of scientific and technological literature in the coal industry; a selecting module, configured to select backup large summary generation models from the original summaries generated by the multiple large summary generation models whose scores are greater than a set threshold; a first adjusting module, configured to obtain question-answer pairs of the scientific and technological literature and adjust the backup large summary generation models through the question-answer pairs to obtain a target large summary generation model; an obtaining module, configured to obtain backup summaries corresponding to the target large summary generation model when processing the scientific and technological literature in the coal industry; and a second adjusting module, configured to adjust the backup summaries based on the knowledge graph corresponding to the scientific and technological literature to obtain the target summary.

[0005] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement a summary generation method combining large model techniques as described in the first aspect of the present disclosure.

[0006] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform a summary generation method incorporating large model techniques as described in the first aspect of the present disclosure.

[0007] The technical solutions provided by the embodiments of this disclosure bring at least the following beneficial effects: In this technical solution, multiple large-scale abstract generation models are identified and used to generate original abstracts of scientific and technological literature in the coal industry. Backup large-scale abstract generation models are selected from the original abstracts generated by these models, provided their scores exceed a set threshold. Question-answer pairs from the scientific and technological literature are obtained, and these backup models are adjusted based on these pairs to obtain the target large-scale abstract generation model. Backup abstracts corresponding to the target large-scale abstract generation model are then obtained. Finally, the backup abstracts are adjusted based on the knowledge graph corresponding to the scientific and technological literature to obtain the target abstract. Thus, by adjusting the backup large-scale abstract generation models through question-answer pairs and by adjusting the backup abstracts through the knowledge graph, the abstracts are continuously optimized, closely aligned with the vertical field of coal mining, and significantly improve the quality and efficiency of scientific and technological literature abstracts.

[0008] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0009] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0010] Figure 1 This is a schematic flowchart of the summary generation method combining large model technology shown in the first embodiment of this disclosure; Figure 2 This is a schematic diagram illustrating the principle of the summary generation method combining large model technology as shown in the second embodiment of this disclosure; Figure 3 This is a schematic diagram of the structure of the summary generation apparatus combining large model technology shown in the third embodiment of this disclosure; Figure 4 This is a schematic diagram of the structure of an electronic device shown in an exemplary embodiment of the present disclosure. Detailed Implementation

[0011] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0012] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0013] It should be noted that the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution disclosed herein are all carried out with the consent of the user, and all comply with the provisions of relevant laws and regulations, and do not violate public order and good morals.

[0014] The following description, with reference to the accompanying drawings, describes a summary generation method and apparatus combining large model techniques according to embodiments of the present disclosure.

[0015] Figure 1 This is a flowchart illustrating the summary generation method combining large model technology as shown in the first embodiment of this disclosure.

[0016] like Figure 1 As shown, the summary generation method combining large model techniques includes the following steps: Step 101: Determine multiple abstract generation models, which are used to generate original abstracts of scientific and technological literature in the coal industry.

[0017] It should be noted that scientific and technological literature includes, but is not limited to, technical reports, papers, and patents in the field of coal mining.

[0018] In this embodiment, multiple general-purpose large models (such as GPT4, Deepseek, etc.) can be called via an external application programming interface (API) as large models for abstract generation to process scientific and technological literature in the coal industry, so as to obtain the corresponding original abstracts. This section focuses on enhancing the results of the original abstracts, generating multiple original abstracts from a single scientific document to increase abstract diversity. Specifically, by calling various mainstream and commonly used models in the coal industry, the number of multiple original abstracts generated is as follows: The sum of the scores for several general models is .in, , Softmax is a function that transforms a set of numerical values ​​into a probability distribution.

[0019] As an example, taking the knowledge service platform of the coal industry as an example, multiple general-purpose large models can be called through the API of the knowledge service platform to quickly and accurately generate the original abstracts of scientific and technological literature in the coal industry. For example, three general-purpose large models, GPT3.5, Deepseek R1, and Q-wen 7B, are externally connected.

[0020] Step 102: Select backup summary generation models from the original summaries generated by multiple summary generation models whose scores are greater than a set threshold.

[0021] In this embodiment of the disclosure, to optimize the summary and improve its quality, as an example, the following steps are taken: First, the first score of the original summaries generated by each summary generation model is obtained. This first score is calculated using a text quality assessment index between the original summary and the backup summary. Second, the second score of the original summaries generated by each summary generation model is obtained. This second score is calculated using a preset user rating mechanism, the number of views, adoptions, and discussions of question-and-answer pairs. The average score and weight of each summary generation model are obtained, with the average score being the average of the scores of multiple original summaries generated by each model. Based on the first score, the second score, the activity parameters of the knowledge service platform corresponding to the question-and-answer pairs, the average score, and the weight, the score of the original summaries generated by each summary generation model is calculated. Backup summary generation models are selected when the score of the original summary is greater than a set threshold. The backup summary is the summary generated by the target summary generation model, which is obtained by selecting the summary generation model through scoring and adjusting it using question-and-answer pairs. Furthermore, the number and types of summary generation models called via API, the number and proportion of generated summaries, can be adjusted according to actual needs, based on the score of the original summary.

[0022] It should be noted that the activity parameters of the knowledge service platform corresponding to the question-and-answer pair are calculated using the monthly active users of the knowledge service platform, the number of views (c) of the question-and-answer pair, the number of adoptions (d) and the number of discussions (e).

[0023] To clarify the scoring method in the original abstract, for example, comparing the original abstract with the backup abstracts, the original abstracts are generated from scientific literature data processed by the knowledge service platform. By default, the top n abstracts with the highest scores are selected. If the abstracts have not yet been scored, n abstracts are randomly selected, denoted as abstract(n). The generated abstract is denoted as text. Then, the first score is calculated using ROUGE. ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is a type of automated text quality assessment metric in the field of natural language generation, especially commonly used for evaluating results in tasks such as summarizing, generative question answering, and dialogue. The first score... The calculation method is as follows:

[0024] Where ROUGE-1 is used to count the number of unary terms that overlap between the original abstract and the prepared abstract, denoted as ROUGE-1(text,abstract(n)). As A in the above formula, the adjustment parameter is set to... ROUGE-2 is used to count the number of overlapping binary terms between the original abstract and the alternative abstract, denoted as ROUGE-2(text,abstract(n)). As B in the above formula, the adjustment parameter is set to... ROUGE-L is used to calculate the longest common subsequence (LCS) between the original abstract and the alternative abstract, denoted as ROUGE-L(text,abstract(n)). C is used in the above formula, and the adjustment parameter is set to... Simultaneously calculate the TF-IDF of the original summary and the alternative summary. TF-IDF (Term Frequency–Inverse Document Frequency) is a feature extraction method in text representation and information retrieval, denoted as TF-IDF(text, abstract(n)). Let D be the value in the above formula, and set the adjustment parameter to... The cosine similarity between the two abstracts is calculated and denoted as cosine_similarity(text,abstract(n)). E is used as the parameter in the above formula, and the adjustment parameter is set to... BLEU (Bilingual Evaluation Alternative) measures the similarity between the machine's original summary and the alternative summary, denoted as BLEU(text,abstract(n)). F is used in the above formula, and the adjustment parameter is set to... . These are adjustment parameters set according to actual conditions. The larger the value, the more important the corresponding technical parameter.

[0025] Second rating The calculation method is as follows: For example, human evaluation is used as the second score. Specifically, a coal mining expert rating mechanism is adopted as the user rating mechanism, and the average score of all expert user ratings is taken. As a weighted average, the more active the users of the knowledge service platform, the higher the value of this evaluation. Here, we calculate the monthly active users (f) of the knowledge service platform, and denote the activity level parameter as... ,in,

[0026] .

[0027] The average score of the multiple original summaries generated by each summary generation model is denoted as . The weights of the corresponding summary generation large model are denoted as... , .

[0028] In summary, the formula for calculating the score of the original abstract can be: The score reflects the quality of the original abstract; the higher the score, the better the quality. The score P is fed back to the target abstract generation model and the backup abstract generation model for model fine-tuning.

[0029] Step 103: Obtain question-and-answer pairs from scientific and technological literature, and adjust the backup abstract generation model using the question-and-answer pairs to obtain the target abstract generation model.

[0030] In this embodiment of the disclosure, in order to select high-quality question-answer pairs for optimizing the backup abstract generation model, as an example, multiple initial question-answer pairs corresponding to scientific and technological documents are obtained from a knowledge service platform; the selection probability of each initial question-answer pair is determined based on the number of views, adoptions, and discussions of each initial question-answer pair; and question-answer pairs with a selection probability greater than a set probability threshold are used as question-answer pairs for scientific and technological documents.

[0031] For example, a knowledge service platform generates question-and-answer pairs about a specific scientific document. Let c be the number of views for each question-and-answer pair, d be the number of views, and e be the number of discussions. Set a threshold. As the probability of selection, when When the probability is greater than the set probability threshold, i.e., when the following formula is satisfied:

[0032] Select high-quality question-and-answer pairs of the required scientific and technological literature.

[0033] Step 104: Obtain the backup abstract corresponding to the target abstract when generating a large model to process scientific and technological literature in the coal industry.

[0034] It should be noted that scientific and technological literature can also include journal articles, patents, conference reports, etc. in the fields of coal mining, safety, electromechanical equipment maintenance, carbon emission control, etc.

[0035] Step 105: Adjust the backup abstracts based on the knowledge graph corresponding to the scientific and technological literature to obtain the target abstract.

[0036] In this embodiment of the disclosure, in order to enhance the accuracy, completeness, and logic of the summary and make it more in line with the information needs of the coal industry, as an example, a knowledge graph corresponding to the scientific and technological literature is obtained; the text of the coal industry in the alternative summary is determined; and the knowledge graph is used as a vector to sort and adjust the text of the coal industry to obtain the target summary.

[0037] Among them, the knowledge graph corresponding to scientific and technological literature is a graph-based tool that uses scientific and technological literature as the core data source and reveals the knowledge structure, research hotspots, and development trends of a discipline through visualization technology. The initial data is automatically generated.

[0038] In summary, multiple large-scale abstract generation models were identified and used to generate original abstracts of scientific and technological literature in the coal industry. Backup abstract generation models were selected from the original abstracts generated by these models, with scores exceeding a set threshold. Question-answer pairs from the scientific and technological literature were obtained, and these backup models were adjusted to obtain the target abstract generation model. Backup abstracts corresponding to the target model were then obtained. Finally, the backup abstracts were adjusted based on the knowledge graph corresponding to the scientific and technological literature to obtain the target abstract. Thus, by adjusting the backup abstract generation models through question-answer pairs and the backup abstracts through the knowledge graph, the abstracts are continuously optimized, closely aligned with the vertical field of coal mining, and significantly improve the quality and efficiency of scientific and technological literature abstracts.

[0039] To clearly describe the summary generation method that combines large model techniques, such as Figure 2 As shown, Figure 2This is a schematic diagram illustrating the principle of a summary generation method combining large-scale model technology, as shown in an embodiment of this disclosure. Specifically, the fine-tuning data generation module is used to process scientific and technological literature in the coal industry by determining multiple large-scale summary generation models to generate original summaries. Simultaneously, an evaluation generation module calculates scores in the original summaries to select backup large-scale summary generation models whose scores exceed a set threshold. Meanwhile, question-and-answer pairs from the scientific and technological literature are obtained and fed back to the fine-tuning data generation module as fine-tuning data to adjust the backup large-scale summary generation models for the coal mining vertical domain, resulting in a target large-scale summary generation model. Then, backup summaries corresponding to the target large-scale summary generation model processing scientific and technological literature in the coal industry are obtained. Simultaneously, based on the graph data generation module, the knowledge graph corresponding to the scientific and technological literature is obtained, and the text in the coal industry is sorted and adjusted using the graph data to obtain the target summary. Thus, through vector sorting of the knowledge graph and fine-tuning of the large-scale model, the quality, efficiency, and logical consistency of the scientific and technological literature summaries are significantly improved.

[0040] Corresponding to the summary generation method combining large model technology provided in the above embodiments, this disclosure also provides a summary generation apparatus combining large model technology. Since the summary generation apparatus combining large model technology provided in this disclosure corresponds to the summary generation method combining large model technology provided in the above embodiments, the implementation of the summary generation method combining large model technology is also applicable to the summary generation apparatus combining large model technology provided in this disclosure, and will not be described in detail in this disclosure.

[0041] Figure 3 This is a schematic diagram of the structure of the summary generation apparatus combining large model technology shown in the third embodiment of this disclosure.

[0042] like Figure 3 As shown, the summary generation device 300 combining large model technology includes: a determining module 301, a selecting module 302, a first adjusting module 303, an acquiring module 304, and a second adjusting module 305.

[0043] The system includes: a determining module 301 for determining multiple abstract generation models, which are used to generate original abstracts of scientific and technological literature in the coal industry; a selecting module 302 for selecting backup abstract generation models from the original abstracts generated by the multiple abstract generation models when the scores are greater than a set threshold; a first adjusting module 303 for obtaining question-answer pairs of the scientific and technological literature and adjusting the backup abstract generation models based on the question-answer pairs to obtain a target abstract generation model; an obtaining module 304 for obtaining backup abstracts corresponding to the target abstract generation model when processing the scientific and technological literature in the coal industry; and a second adjusting module 305 for adjusting the backup abstracts based on the knowledge graph corresponding to the scientific and technological literature to obtain the target abstract.

[0044] As one possible implementation of this disclosure, the selection module 302 is specifically used for: obtaining a first score of the original summary generated by each of the summary generation models; wherein the first score is calculated using a text quality assessment index between the original summary and the backup summary; obtaining a second score of the original summary generated by each of the summary generation models; wherein the second score is calculated using a preset user rating mechanism, the number of views, adoptions, and discussions of the question-answer pair; obtaining the average score and weight of each of the summary generation models, wherein the average score is the average of the scores of multiple original summaries generated by each of the summary generation models; calculating the score of the original summary generated by each of the summary generation models based on the first score, the second score, the activity parameter of the knowledge service platform corresponding to the question-answer pair, the average score, and the weight; and selecting the backup summary generation model corresponding to the original summary whose score is greater than a set threshold.

[0045] As one possible implementation of this disclosure, the activity parameter of the knowledge service platform corresponding to the question-and-answer pair is calculated using the monthly active users of the knowledge service platform, the number of views, adoptions, and discussions of the question-and-answer pair.

[0046] As one possible implementation of this disclosure, the first adjustment module 303 is specifically used for: obtaining multiple initial question-and-answer pairs corresponding to the scientific and technological documents from the knowledge service platform; determining the selection probability of each initial question-and-answer pair based on the number of views, adoptions, and discussions of each initial question-and-answer pair; and using question-and-answer pairs with selection probabilities greater than a set probability threshold as question-and-answer pairs for scientific and technological documents.

[0047] As one possible implementation of this disclosure, the second adjustment module 305 is specifically used for: obtaining the knowledge graph corresponding to the scientific and technological literature; determining the text of the coal industry in the spare summary; and sorting and adjusting the text of the coal industry using the knowledge graph as a vector to obtain the target summary.

[0048] This embodiment of the abstract generation apparatus, incorporating large-scale model technology, determines multiple large-scale abstract generation models for generating original abstracts of scientific and technological literature in the coal industry. It selects backup large-scale abstract generation models from the original abstracts generated by these models, choosing those with scores exceeding a set threshold. It then obtains question-and-answer pairs of the scientific and technological literature and adjusts these backup models to obtain a target large-scale abstract generation model. Finally, it obtains backup abstracts corresponding to the target large-scale abstract generation model when processing scientific and technological literature in the coal industry. Based on the knowledge graph corresponding to the scientific and technological literature, it adjusts these backup abstracts to obtain the target abstract. Thus, by adjusting the backup large-scale abstract generation models through question-and-answer pairs and by adjusting the backup abstracts through the knowledge graph, the abstracts are continuously optimized, closely aligned with the vertical field of coal mining, significantly improving the quality and efficiency of scientific and technological literature abstracts.

[0049] In an exemplary embodiment, an electronic device is also proposed.

[0050] The electronic devices include: processor; Memory used to store processor-executable instructions; The processor is configured to execute instructions to implement the summary generation method combining large model techniques as proposed in any of the foregoing embodiments.

[0051] As an example, Figure 4 This is a schematic diagram of the structure of an electronic device 400 as shown in an exemplary embodiment of this disclosure, as follows: Figure 4 As shown, the above-mentioned electronic device 400 may further include: The memory 410 and processor 420 are connected by a bus 430, which connects the different components (including the memory 410 and the processor 420). The memory 410 stores a computer program, and when the processor 420 executes the program, it implements the summary generation method combining large model technology as described in the embodiments of this disclosure.

[0052] Bus 430 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0053] Electronic device 400 typically includes a variety of electronic device readable media. These media can be any available media that can be accessed by electronic device 400, including volatile and non-volatile media, removable and non-removable media.

[0054] Memory 410 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 440 and / or cache memory 450. The server may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 460 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 4 Not shown; usually referred to as a "hard drive"). Although Figure 4 As not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 430 via one or more data media interfaces. Memory 410 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.

[0055] A program / utility 480 having a set (at least one) of program modules 470 may be stored, for example, in memory 410. Such program modules 470 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 470 typically perform the functions and / or methods described in the embodiments of this disclosure.

[0056] Electronic device 400 can also communicate with one or more external devices 490 (e.g., keyboard, pointing device, display 491, etc.), and with one or more devices that enable a user to interact with electronic device 400, and / or with any device that enables electronic device 400 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 492. Furthermore, electronic device 400 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 493. As shown, network adapter 493 communicates with other modules of electronic device 400 via bus 430. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0057] The processor 420 performs various functional applications and data processing by running programs stored in the memory 410.

[0058] It should be noted that the implementation process and technical principles of the electronic device in this embodiment are explained in the foregoing description of the summary generation method combining large model technology in the embodiments of this disclosure, and will not be repeated here.

[0059] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions that can be executed by a processor of an electronic device to perform the summary generation method combining large model techniques proposed in any of the above embodiments. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0060] In an exemplary embodiment, a computer program product is also provided, including a computer program / instructions that, when executed by a processor, implement the summary generation method combining large model technology proposed in any of the above embodiments.

[0061] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0062] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A summarization method combining large model techniques, characterized in that, The method includes: Multiple abstract generation models are identified, which are used to generate original abstracts of scientific and technological literature in the coal industry. Select backup summary generation models from the original summaries generated by multiple summary generation models whose scores are greater than a set threshold; The question-and-answer pairs of the scientific and technological documents are obtained, and the backup abstract generation model is adjusted based on the question-and-answer pairs to obtain the target abstract generation model. The target abstract is used to generate a backup abstract for processing scientific and technological literature in the coal industry using a large model. The backup abstract is adjusted based on the knowledge graph corresponding to the scientific and technological literature to obtain the target abstract.

2. The method according to claim 1, characterized in that, The selection of backup summary generation models from the original summaries generated by multiple summary generation models, where the score is greater than a set threshold, includes: Obtain the first score of the original summary generated by each of the aforementioned summary generation models; wherein the first score is calculated using a text quality assessment index between the original summary and the alternative summary. Obtain a second score for the original summaries generated by each of the aforementioned summary generation models; wherein the second score is calculated using a preset user rating mechanism, the number of views, adoptions, and discussions of the question-answer pair; Obtain the average score and weight of each of the above-mentioned summary generation models, wherein the average score is the average score of multiple original summaries generated by each of the above-mentioned summary generation models; Based on the first score, the second score, the activity parameters of the corresponding knowledge service platform, the average score and the weight, the score of the original summary generated by each summary generation model is calculated. The backup summary corresponding to the original summary whose score is greater than a set threshold is selected to generate a large model.

3. The method according to claim 2, characterized in that, The activity parameter of the knowledge service platform corresponding to the question-and-answer pair is calculated by the monthly active users of the knowledge service platform, the number of views, adoptions, and discussions of the question-and-answer pair.

4. The method according to claim 2, characterized in that, The question-and-answer pairs for obtaining the scientific and technological documents include: Obtain multiple initial question-and-answer pairs corresponding to the scientific and technological documents from the knowledge service platform; The probability of selecting each initial question-and-answer pair is determined based on the number of views, adoptions, and discussions. Question-answer pairs with a probability greater than a set probability threshold will be selected as question-answer pairs for scientific and technological literature.

5. The method according to claim 1, characterized in that, The adjustment of the backup abstract based on the knowledge graph corresponding to the scientific and technological literature to obtain the target abstract includes: Obtain the knowledge graph corresponding to the scientific and technological documents; Determine the text related to the coal industry in the alternative summary; The knowledge graph is used as a vector to organize and adjust the text of the coal industry to obtain the target summary.

6. A summary generation device combining large model technology, characterized in that, The device includes: A determination module is used to determine multiple abstract generation models, which are used to generate original abstracts of scientific and technological literature in the coal industry. The selection module is used to select backup summary generation models from the original summaries generated by multiple summary generation models when the score is greater than a set threshold. The first adjustment module is used to obtain the question-and-answer pairs of the scientific and technological documents, and adjust the backup abstract generation model through the question-and-answer pairs to obtain the target abstract generation model. The acquisition module is used to acquire the backup summary corresponding to the target summary generation model when processing the scientific and technological literature of the coal industry; The second adjustment module is used to adjust the backup abstract based on the knowledge graph corresponding to the scientific and technological literature to obtain the target abstract.

7. The apparatus according to claim 6, characterized in that, The selection module is specifically used for: Obtain the first score of the original summary generated by each of the aforementioned summary generation models; wherein the first score is calculated using a text quality assessment index between the original summary and the alternative summary. Obtain a second score for the original summaries generated by each of the aforementioned summary generation models; wherein the second score is calculated using a preset user rating mechanism, the number of views, adoptions, and discussions of the question-answer pair; Obtain the average score and weight of each of the above-mentioned summary generation models, wherein the average score is the average score of multiple original summaries generated by each of the above-mentioned summary generation models; Based on the first score, the second score, the activity parameters of the corresponding knowledge service platform, the average score and the weight, the score of the original summary generated by each summary generation model is calculated. The backup summary corresponding to the original summary whose score is greater than a set threshold is selected to generate a large model.

8. The apparatus according to claim 7, characterized in that, The activity parameter of the knowledge service platform corresponding to the question-and-answer pair is calculated by the monthly active users of the knowledge service platform, the number of views, adoptions, and discussions of the question-and-answer pair.

9. The apparatus according to claim 7, characterized in that, The first adjustment module is specifically used for: Obtain multiple initial question-and-answer pairs corresponding to the scientific and technological documents from the knowledge service platform; The probability of selecting each initial question-and-answer pair is determined based on the number of views, adoptions, and discussions. Question-answer pairs with a probability greater than a set probability threshold will be selected as question-answer pairs for scientific and technological literature.

10. The apparatus according to claim 6, characterized in that, The second adjustment module is specifically used for: Obtain the knowledge graph corresponding to the scientific and technological documents; Determine the text related to the coal industry in the alternative summary; The knowledge graph is used as a vector to organize and adjust the text of the coal industry to obtain the target summary.