Work report generation method and system based on large model

Through the large-model-based work report generation method, the full process automation from user input titles to generating a complete work report is achieved, and the problems of limited intelligence level and insufficient automation in the existing technology are solved, and the generation efficiency and report quality are significantly improved, and it can quickly adapt to policy environment changes and data volume growth.

CN120068814APending Publication Date: 2025-05-30SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510197925.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

It is difficult for the existing technology to achieve efficient and accurate generation of work reports, especially in the rapidly changing policy environment and the scenario of massive data growth. The level of intelligence is limited, and the internal connection between complex policy information and data cannot be fully understood. The degree of automation is insufficient, making it difficult to achieve full process automation.

Method used

The method of generating work report based on big models is adopted, and the knowledge base is pre-constructed and managed, combining keyword big models, deep learning models and paragraph big models, to achieve the full process automation from user input titles to generating a complete work report. The method includes steps such as semantic understanding and keyword extraction, structured outline generation and paragraph loop generation.

Benefits of technology

It significantly improves the efficiency of the production of work reports, enhances the logic and depth of the report, reduces the degree of manual participation, reduces labor costs and correction costs caused by human errors, and can quickly adapt to policy environment changes and data volume growth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068814A_ABST
    Figure CN120068814A_ABST
Patent Text Reader

Abstract

The invention provides a work report generation method and system based on a large model, and belongs to the technical field of automatic office, and the method comprises the steps: pre-constructing a knowledge base for storing work report generation related data information, and carrying out the management; the method comprises the following steps: training a keyword large model for generating keywords, a deep learning model for generating an outline and a paragraph large model for generating paragraphs in advance; in response to a title input by a user, performing semantic understanding and keyword extraction by using a keyword large model, and then generating keywords related to the title; in response to a title and a keyword input by a user, generating a structured outline of the work report in combination with the knowledge graph and the deep learning model; and combining the structured outline with data information in the knowledge base, and circularly generating paragraphs by using a paragraph large model until the generation of the work report is completed. By integrating an artificial intelligence technology and a natural language processing technology, full-process automation of work report generation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of automated office, and particularly relates to a method and system for generating work reports based on large models. Background Art

[0002] In today's digital age, it has become a daily necessity for all walks of life to produce work summary reports, and the efficiency and quality of writing work summary reports have become increasingly important. The traditional way of writing work reports mainly relies on manual completion by staff. They need to spend a lot of time and energy collecting information from a large amount of policy documents, meeting records, work data and other materials, and then writing and editing. The whole process of making work summary reports is not only inefficient, but also easily affected by human factors, making it difficult to ensure the accuracy and consistency of the content of work reports.

[0003] With the development of artificial intelligence technology, natural language processing technology has brought new possibilities for automated report generation. However, existing technical solutions have many deficiencies and cannot meet the requirements for generating work reports, especially in scenarios that need to cope with a rapidly changing policy environment and the growth of a large amount of data. On the one hand, the level of intelligence is limited, unable to fully understand the internal relationship between complex policy information and data, resulting in reports with low depth and poor logic. On the other hand, the degree of automation is low, and it is difficult to achieve full-process automation from data collection to report generation, and still requires a large amount of manual participation. Currently, existing technical solutions cannot meet the requirements for efficiently and accurately generating work reports. Summary of the Invention

[0004] In a first aspect, an embodiment of the present application provides a method for generating a work report based on a large model, including the following steps: S1. Pre-build a knowledge base for storing data and materials related to work report generation and manage it; S2. Pre-train a keyword large model for generating keywords, a deep learning model for generating outlines, and a paragraph large model for generating paragraphs; S3. In response to the title input by the user, use the keyword large model for semantic understanding and keyword extraction to generate keywords related to the title; S4. In response to the title and keywords input by the user, combine the knowledge graph and the deep learning model to generate a structured outline of the work report; S5. Combine the structured outline with the data and materials in the knowledge base, and use the paragraph large model to repeatedly generate paragraphs until the generation of the work report is completed.

[0005] Further, the specific steps of step S1 are as follows: S11. Initialize the knowledge base and determine the knowledge types according to the types of work reports; the knowledge types include policy documents related to work reports, meeting records, industry data materials, and reference materials; S12. Set configuration permissions for the knowledge base and enable or disable the automatic knowledge collection method through the management user; S13. When the automatic collection method is enabled, use web crawlers to automatically collect policy documents, meeting records, and industry data materials related to work reports from the Internet, and manually upload materials when the automatic collection method is disabled; S14. According to the knowledge types, use large models and the automated tool RPA to intelligently screen and organize the collected materials, and store the identified materials as knowledge in the knowledge base; S15. Classify the materials in the knowledge base and establish an index, and provide a retrieval interface after classification is completed; S16. Respond to the user's retrieval request through the retrieval interface, perform a retrieval operation according to the retrieval request and the index of the knowledge base, and obtain the retrieval result; If the retrieval result does not meet the user's needs, go to step S17; If the retrieval result meets the user's needs, go to step S18; S17. Receive the user's feedback information, optimize the retrieval method, and then return to step S16; S18. Display the retrieval result to the user and provide access and download interfaces; S19. Regularly audit and maintain the knowledge base, and update the materials in the knowledge base when the knowledge base needs to be updated.

[0006] Further, the specific steps of step S2 are as follows: S21. Pre-collect text data related to work reports, covering the fields and topics involved in work reports, construct a first data set, and collect the titles and corresponding keywords related to work reports in historical work reports and policy documents as task data; S22. Use the first data set to pre-train the model of the Transformer architecture, and use the language modeling task to optimize during the pre-training process to obtain the pre-trained model; S23. Connect a classification layer to the output of the pre-trained model, set the titles in the task data as the input, and set the keywords as the output to complete the fine-tuning target setting, and then use the task data for fine-tuning training, and optimize in combination with the loss function to obtain the keyword large model; S24. Pre-collect the text data of historical work reports, and the text data of the historical work reports includes historical titles, keywords, work report contents, and work report outlines, and construct a second data set; S25. Select a bidirectional LSTM network as the RNN model, define the input layer and output layer of the RNN model, and add an activation function. Input the historical titles, keywords, and work report contents in the second dataset into the input layer, and input the historical work report outlines in the training dataset into the output layer; S26. Use the tensorflow framework to train the RNN model, and use the Adam optimizer to adjust the learning rate during the training process. Use the cosine similarity function to calculate the correlation between the title and keywords, and calculate the difference between the outline generated by the RNN model and the historical work report outlines in the second dataset through the cross-entropy loss function. When the difference meets the requirements, obtain the deep learning model; S27. Pre-collect text data related to work reports, covering the fields and topics involved in work reports. After cleaning and tokenizing the collected data, construct a vocabulary table and map each word to a unique index to obtain the third dataset. Also, collect data related to the paragraphs of work reports and organize them according to the lowest-level outline objectives to obtain the fourth dataset; the paragraph-related data includes structured outlines, titles, keywords, reference materials, and paragraph contents, and the lowest-level outline objectives include titles, keywords, reference materials, and expected generated paragraphs; reference materials are used to assist the paragraph large model in learning and generating paragraph contents that conform to the actual situation; S28. Select a large language model based on the Transformer architecture as the base model, and then use the masked language model and the next prediction task to pre-train the paragraph large model in combination with the third dataset. Calculate the total loss through the cross-entropy loss function during the pre-training process until the total loss meets the requirements; S29. Set the fine-tuning objective to generate paragraph contents based on the titles, keywords, and reference materials of the lowest-level outline objectives. Adjust the pre-trained paragraph large model using the fourth dataset according to the fine-tuning objective, and use the cross-entropy loss function to calculate the difference between the paragraphs generated by the paragraph large model and the paragraphs in the fourth dataset until the difference meets the requirements to obtain the final paragraph large model.

[0007] Further, the specific steps of step S3 are as follows: S31. Pre-use the faiss vector database to store keywords, construct a keyword knowledge base, and create an ES index; S32. After preprocessing the title input by the user, input it into the keyword large model for semantic understanding, and extract the initial keywords; S33. Determine whether the keyword data volume is sufficient; If so, go to step S36; If not, go to step S34; S34. Use the Elasticsearch tool to retrieve the ES index in the keyword knowledge base in combination with the title, and further extract keywords; S35. Determine whether the number of keywords is sufficient; If so, go to step S36; If not, go to step S37; S36. Use the BGE tool to perform cross-language retrieval of relevant documents according to the title, extract cross-language keywords, and return to step S35; S37. Merge the keywords, and then match the keywords with the knowledge base to obtain the final keyword list.

[0008] Furthermore, the specific steps of step S4 are as follows: S41. Use the neo4j library of python to construct a knowledge graph of the historical work report outline in advance and store it; S42. Obtain the title and extracted keywords input by the user, preprocess the title and keywords, and use the Elasticsearch tool to query the knowledge graph; If the query result meets the requirements, generate a preliminary outline and go to step S46; If the query result does not meet the requirements, go to step S43; S43. Input the title and keywords into the deep learning model to output the relevance between the title and keywords, and obtain the candidate outline; S44. Receive user feedback and determine whether the candidate outline is valid; If so, go to step S46; If not, go to step S45; S45. Adjust the candidate outline according to the user feedback, and respond to the user's confirmation of the adjustment of the candidate outline, and return to step S44; S46. Merge and optimize the candidate outline or the preliminary outline, and perform tree-structured display to obtain the keywords of the lowest-level outline, and retrieve reference materials, which are used to provide specific content support for subsequent paragraph generation; If the reference materials are sufficient, go to step S48; If the reference materials are insufficient, go to step S47; S48. Further collect materials, and after the user reviews and confirms, return to step S46; S47. Display the reference materials and generate the structured outline of the final work report.

[0009] Furthermore, the specific steps of step S5 are as follows: S51. Initialize the paragraph generation parameters according to the structured outline, retrieve the data and reference materials in the knowledge base according to the structured outline, and determine whether the reference materials and data are sufficient; If so, go to step S53; If not, go to step S52; S52. Supplement the materials and data, and return to step S51; S53. Use the paragraph large model to start paragraph generation in combination with the reference materials and data, apply the multi-head attention mechanism to generate the initial paragraph draft, and verify whether the quality of the initial paragraph draft meets the requirements; If so, go to step S55; If not, go to step S54; S54. Regenerate the paragraph and return to step S53; S55. Optimize the initial paragraph draft to obtain the final paragraph draft, and determine whether all paragraphs have been completed; If so, go to step S57; If not, go to step S56; S56. Start generating the next paragraph and return to step S51; S57. Integrate all paragraphs to generate a draft report and provide it to the user; S58. Receive user feedback; If the user is satisfied, generate the final work report and end; If the user is not satisfied, go to step S59; S59. Adjust the draft report according to the user feedback and return to step S58.

[0010] In a second aspect, the embodiments of the present application further provide a work report generation system based on a large model, including: A knowledge base management module that manages the knowledge base storing data and materials related to work report generation, and is provided with retrieval and document management interfaces; A keyword generation module that, in response to the title input by the user, uses the large model to perform semantic understanding and keyword extraction and then generates keywords related to the title; An outline generation module that, in response to the title and keywords input by the user, combines the knowledge graph and the deep learning model to generate a structured outline of the work report; A paragraph generation module that, according to the structured outline, combines the data and materials in the knowledge base, and uses the large model to generate paragraphs in a loop until the generation of the work report is completed.

[0011] Further, the knowledge base management module includes: A knowledge base initialization unit that initializes the knowledge base; The knowledge collection method configuration unit sets configuration permissions for the knowledge base and enables or disables the automatic knowledge collection method through the management user. The knowledge collection unit, when the automatic collection method is enabled, uses web crawlers to automatically collect policy documents, meeting records, and industry data related to work reports from the Internet, and manually uploads materials when the automatic collection method is disabled. The knowledge base storage unit, according to the knowledge type, uses large models and the automated tool RPA to intelligently screen and organize the collected materials, stores the identified materials as knowledge in the knowledge base, and classifies and indexes them. The knowledge retrieval unit responds to the user's retrieval request through the retrieval interface, performs retrieval operations according to the retrieval request and the index of the knowledge base, and obtains the retrieval results. The knowledge base maintenance unit regularly audits and maintains the knowledge base, and updates the materials in the knowledge base when the knowledge base needs to be updated.

[0012] Furthermore, the keyword generation module includes: The keyword knowledge base construction unit pre-uses the faiss vector database to store keywords, constructs a keyword knowledge base, and creates an ES index. The keyword extraction unit preprocesses the title input by the user and then inputs it into the keyword large model for semantic understanding to extract initial keywords. The keyword knowledge base retrieval unit uses the Elasticsearch tool to retrieve the ES index in the keyword knowledge base in combination with the title to further extract keywords. The keyword cross-language retrieval unit, when the number of keywords is insufficient, uses the BGE tool to perform cross-language retrieval of relevant documents according to the title to extract cross-language keywords. The keyword merging unit merges the keywords, then matches the keywords with the knowledge base to obtain the final keyword list.

[0013] Furthermore, the outline generation module includes: The knowledge graph construction unit pre-uses the neo4j library of python to construct a knowledge graph of historical work report outlines and stores it. The knowledge graph query unit obtains the title and extracted keywords input by the user, preprocesses the title and keywords, and then uses the Elasticsearch tool to query the knowledge graph. The candidate outline generation unit, when the query results do not meet the requirements, inputs the title and keywords into the deep learning model to output the relevance between the title and keywords, and obtains the candidate outline. The outline tree-shaped display unit merges and optimizes the candidate outline or preliminary outline, and performs tree-shaped structured display to obtain the keywords of the lowest-level outline, and retrieves reference materials; The structured outline generation unit displays the reference materials and generates the structured outline of the final work report when there is sufficient reference material; The paragraph generation module includes: The knowledge base retrieval unit initializes the paragraph generation parameters according to the structured outline, retrieves the data and reference materials in the knowledge base according to the structured outline, and judges whether the reference materials and data are sufficient; The initial paragraph generation unit starts paragraph generation using the paragraph large model in combination with reference materials and data, applies the multi-head attention mechanism to generate the initial paragraph, and verifies whether the quality of the initial paragraph meets the requirements; The report draft generation unit integrates all paragraphs to generate a report draft and provides it to the user when all paragraphs are completed.

[0014] As can be seen from the above technical solutions, the present invention has the following advantages: In the method and system for generating a work report based on a large model provided by the present application, through automated data collection, keyword extraction, outline generation, and paragraph generation, the time of manual operation is significantly reduced, and the generation efficiency of the work report is remarkably improved; by using the large model and deep learning technology, the internal connection between policy information and data can be understood more accurately, and the logic and depth of the report content can be improved; the degree of manual participation is reduced, and the labor cost and the correction cost caused by human errors are reduced; it can quickly adapt to the changes in the policy environment and the growth of data volume, update the knowledge base and model in a timely manner, and ensure that the generated report meets the latest requirements. Brief Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a schematic flow chart of the method for generating a work report based on a large model of the present invention.

[0017] Figure 2 It is a schematic diagram of the system for generating a work report based on a large model of the present invention. Detailed Description of the Embodiment

[0018] In the following detailed description of the specific steps of the large model-based work report generation method, various embodiments of the present disclosure will be more comprehensively described. The present disclosure can have various embodiments, and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents, and / or alternative solutions that fall within the spirit and scope of the various embodiments of the present disclosure.

[0019] Exemplarily speaking, in the context of today's digital age, writing work summary reports has become an indispensable part of the daily work in various industries, and the importance of its efficiency and quality has become increasingly prominent. The traditional work report writing mode highly relies on manual operations. Staff need to invest a huge amount of time and energy to collect information from numerous policy documents, meeting records, and work data, and then write and edit. This process is not only inefficient but also vulnerable to human factors, making it difficult to guarantee the accuracy and consistency of the content of the work report.

[0020] With the development of artificial intelligence technology, natural language processing technology provides a new opportunity for the automated generation of work reports. However, the existing technical solutions still have many defects and are difficult to meet the actual needs of work report generation. Especially in the context of rapid policy changes and a sharp increase in data volume, the problem is particularly prominent. On the one hand, the level of intelligence is still insufficient, unable to deeply understand and grasp complex policy information and the internal relationships between data, resulting in reports lacking depth and poor logic. On the other hand, the automated process is still imperfect, unable to achieve full-chain automated processing from data collection to report generation, and still requires a large amount of manual intervention. Therefore, the existing technical solutions are difficult to meet the urgent need for efficient and accurate generation of work reports.

[0021] In response to the above problems, this embodiment provides a large model-based work report generation method. By integrating artificial intelligence technology and natural language processing technology, it realizes the full-process automation of work report generation; not only improves the generation efficiency of work reports but also enhances the quality of the reports; by pre-constructing and managing the knowledge base and combining the application of keyword large models, deep learning models, and paragraph large models, it can quickly respond to user needs and generate structured and logically strong work reports.

[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] Please refer to Figure 1The figure shows a flowchart of a method for generating a work report based on a large model in a specific embodiment. The method includes the following steps: S1. Pre-construct a knowledge base for storing data and materials related to work report generation and manage it; It should be noted that by centrally managing various policy documents, meeting records, industry data and materials, etc. in the knowledge base, it is ensured that accurate data and materials can be obtained at any time during the report generation process; by classifying, indexing and retrieving the knowledge base, it is convenient to quickly locate the required materials, improve the efficiency of information acquisition, and reduce the time cost of searching for information in a large amount of data; S2. Pre-train a keyword large model for generating keywords, a deep learning model for generating outlines, and a paragraph large model for generating paragraphs; It should be noted that by pre-training the keyword large model, the deep learning model and the paragraph large model, the model's ability to understand the internal relationship between policy information and data is improved, and the depth and logic of the work report are enhanced; S3. In response to the title input by the user, use the keyword large model to perform semantic understanding and keyword extraction to generate keywords related to the title; It should be noted that by using the large model to perform semantic understanding and keyword extraction on the title, the core points of the work report can be quickly grasped, providing key clues for outline generation and work report content writing; by using the large model to extract keywords, the title can also be analyzed from multiple angles, and keywords with strong representativeness and relevance can be extracted; through the keywords, the theme direction of the work report can be clarified to ensure that the generated work report meets the user's needs; S4. In response to the title and keywords input by the user, combine the knowledge graph and the deep learning model to generate a structured outline of the work report; It should be noted that the knowledge graph can help the deep learning model understand the relationship between different concepts, so as to generate a logical and systematic outline. The outline can plan the structure and content level of the work report, making the work report well-organized; by combining the title and keywords through the deep learning model, an outline framework that meets the requirements can be generated according to historical data and semantic relationships, reducing the time for manual outline conception; S5. Combine the structured outline with the data and materials in the knowledge base, and use the paragraph large model to generate paragraphs in a loop until the generation of the work report is completed; It should be noted that by using the paragraph large model and generating paragraphs in a loop, the automatic generation of the report content is realized, and the generation efficiency is improved.

[0024] In this embodiment, by pre - constructing and managing a knowledge base, data support is provided for the generation of work reports; by pre - training keyword large models, deep learning models, and paragraph large models, the model's understanding ability of the internal relationship between policy information and data is improved; the full - process automation from the user - input title to the generation of a complete work report is realized, reducing manual intervention and improving the generation efficiency.

[0025] Furthermore, as a refinement and extension of the specific implementation manner of the above - mentioned embodiment, in order to fully illustrate the specific implementation process in this embodiment, another method for generating a work report based on a large model is provided. This method includes the following steps: S1. Pre - construct a knowledge base for storing data related to the generation of work reports and manage it. The specific steps of step S1 are as follows: S11. Initialize the knowledge base and determine the knowledge types according to the types of work reports. The knowledge types include policy documents related to work reports, meeting records, industry data, and reference materials; S12. Set configuration permissions for the knowledge base and determine whether to enable the automatic knowledge collection method through the management user; S13. When the automatic collection method is enabled, use web crawlers to automatically collect policy documents, meeting records, and industry data related to work reports from the Internet. When the automatic collection method is disabled, manually upload materials; Exemplarily, use Scrapy to crawl policy websites and filter noise in combination with a large model; The implementation code is as follows: import scrapy class GovReportSpider(scrapy.Spider): def parse(self, response): text = response.css("div.content::text").get() if self.model.filter(text):# Call the large model for screening yield {"content": text} Use spaCy for entity recognition and extract key fields; S14. According to the knowledge types, and through a large model and the automation tool RPA, intelligently screen and organize the collected materials, and store the materials identified as knowledge in the knowledge base; Exemplarily, for automatic file reading, using RPA tools can simulate manual operations to open relevant report files (such as PDF or Word documents), read the file content, and extract key information; Use libraries such as pyautogui to simulate mouse clicks and keyboard inputs to open the file; Use the PyPDF2 and python-docx libraries to read the file content; The relevant code is as follows: Open the relevant report file: RPA Robot.open_file(file_path) Read the file content: content = RPA Robot.read_content(file_path) Extract key information: title, content = RPA Robot.extract_info(content) For automatic recognition, by combining RPA with large models and using natural language processing techniques to identify key knowledge such as entities, events, and relationships in the document; The code framework for calling the large model to analyze the extracted text is as follows: Extract text: text = RPA Robot.extract_text(file_content) Call the large model: entities, relations = NLP Model.analyze(text) As an automated input tool for the knowledge base management module, the RPA tool automatically stores the data collected from the Internet and documents into the knowledge base; the RPA tool combines with keyword generation to automatically extract keywords from the document and update the keyword knowledge base; S15. Classify the materials in the knowledge base and establish an index, and provide a retrieval interface after the classification is completed; S16. Respond to the user's retrieval request through the retrieval interface, perform a retrieval operation according to the retrieval request and the index of the knowledge base, and obtain the retrieval result; If the retrieval result does not meet the user's requirements, go to step S17; If the retrieval result meets the user's requirements, go to step S18; S17. Receive the user feedback information, optimize the retrieval method, and then return to step S16; S18. Display the retrieval result to the user and provide access and download interfaces; S19. Regularly audit and maintain the knowledge base, and update the materials in the knowledge base when the knowledge base needs to be updated; It should be noted that by initializing the knowledge base and determining the knowledge types, a clear direction is provided for the construction and management of the knowledge base. By configuring permissions and knowledge collection methods, the security of data is ensured while flexible data collection channels are provided. By intelligently screening and sorting materials, the quality of data in the knowledge base is improved. By establishing indexes and providing retrieval interfaces, users can quickly obtain the required materials. By regularly auditing and maintaining the knowledge base, the timeliness and accuracy of data are ensured; S2. Pre-train the keyword large model for generating keywords, the deep learning model for generating outlines, and the passage large model for generating paragraphs. The specific steps of step S2 are as follows: S21. Pre-collect text data related to work reports, covering the fields and topics involved in work reports, construct the first data set, and collect the titles and corresponding keywords related to work reports in historical work reports and policy documents as task data; S22. Use the first data set to pre-train the model with the Transformer architecture, and optimize it using the language modeling task during the pre-training process to obtain the pre-trained model; Specifically, use the model with the Transformer architecture of the GPT series or BERT. The Transformer architecture captures long-range dependencies in the text through the self-attention mechanism; In the pre-training stage, use the language modeling task, taking the masked language model MLM as an example, to optimize the model. Taking the model with the Transformer architecture of BERT as an example, construct the loss function:

[0026] where N is the sequence length, is the probability of predicting the i-th word given the previous words; by minimizing this loss function, the model learns the statistical laws and semantic representations of the language; S23. Connect a classification layer to the output of the pre-trained model, set the titles in the task data as the input, and set the keywords as the output to complete the fine-tuning target setting. Then use the task data for fine-tuning training, optimize it in combination with the loss function to obtain the keyword large model; It should be noted that for the keyword extraction task, it can be regarded as a multi-label classification problem. In the fine-tuning stage, use the cross-entropy loss function to optimize the model to minimize the difference between the predicted keywords and the true keywords. The formula of the cross-entropy loss function is as follows:

[0027] where M is the number of keywords as the number of categories, is the true label (0 or 1, indicating whether it is this keyword), is the probability predicted by the pre-trained model; S24. Pre-collect the text data of historical work reports, where the text data of the historical work reports includes historical titles, keywords, work report contents, and work report outlines, and construct a second data set; S25. Select a bidirectional LSTM network as the RNN model, define the input layer and output layer of the RNN model, and add an activation function. Input the historical titles, keywords, and work report contents in the second data set to the input layer, and input the historical work report outlines in the training set to the output layer; It should be noted that the activation function ReLU is selected; S26. Use the tensorflow framework to train the RNN model, and use the Adam optimizer to adjust the learning rate during the training process. Use the cosine similarity function to calculate the correlation between the title and the keyword, and calculate the difference between the outline generated by the RNN model and the historical work report outlines in the second data set through the cross-entropy loss function. When the difference meets the requirements, obtain the deep learning model; It should be noted that the following formula is used to calculate the correlation CS:

[0028] where A is the title vector and B is the keyword vector, is the dot product of the title vector and the keyword vector, is the norm of the title vector, is the norm of the title vector; Calculate the difference between the outline generated by the RNN model and the historical work report outlines in the second data set through the following cross-entropy loss function:

[0029] where, is the true label, is the predicted probability of the model; S27. Pre-collect the text data related to work reports, covering the fields and topics involved in the work reports. After cleaning and tokenizing the collected data, construct a vocabulary and map each word to a unique index to obtain a third data set, and collect the data related to the paragraphs of the work reports and organize them according to the lowest-level outline objectives to obtain a fourth data set; the data related to the paragraphs includes structured outlines, titles, keywords, reference materials, and paragraph contents, and the lowest-level outline objectives include titles, keywords, reference materials, and expected generated paragraphs; the reference materials are used to assist the paragraph large model to learn and generate paragraph contents that conform to the actual situation; It should be noted that cleaning the collected data includes removing noise and incorrect characters; S28. Select a large language model based on the Transformer architecture as the base model, and then use the masked language model and the next sentence prediction task, combined with the third dataset, to perform pre-training on the paragraph large model. During the pre-training process, calculate the total loss through the cross-entropy loss function until the total loss meets the requirements; It should be noted that when selecting a large language model based on the Transformer architecture, the attention mechanism is used to capture semantic information in different aspects; Randomly input some words in the sequence to the masked language model MLM, and let the model predict these masked words. The cross-entropy loss function is constructed as follows:

[0030] where, is the masked word, is the input sequence except for the masked word, is the probability predicted by the model, and mp is the set of positions of the masks; The next sentence prediction task NSP judges whether two sentences are consecutive, and also uses the cross-entropy loss function as follows:

[0031] where, is the true label, is the probability predicted by the next sentence prediction task NSP; The total loss function for pre-training is:

[0032] where, and are hyperparameters for balancing the weights of the two tasks of the masked language model MLM and the next sentence prediction task NSP; S29. Set the fine-tuning objective to generate paragraph content based on the titles, keywords, and reference materials of the lowest-level outline objectives. Use the fourth dataset to adjust the pre-trained paragraph large model according to the fine-tuning objective, and use the cross-entropy loss function to measure the difference between the paragraphs generated by the paragraph large model and the paragraphs in the fourth dataset until the difference meets the requirements to obtain the final paragraph large model; It should be noted that the cross-entropy loss function is used to measure the difference between the paragraphs generated by the pre-trained model and the true paragraphs; taking the true word sequence as as an example, the word sequence generated by the pre-trained model is , and the loss function is:

[0033] where, is the size of the vocabulary, is the one - hot encoding of the real word at time step t, is the probability predicted after the model is pre - trained; It should be noted that through data and pre - training and fine - tuning, the model can be adapted to the specific task of work report generation. By selecting appropriate model architectures and training methods, such as bidirectional LSTM networks and Transformer architectures, the advantages of the model are utilized, and the quality of outline and paragraph generation is improved. And by using a variety of loss functions and optimization methods, the effectiveness of model training is ensured; S3. In response to the title input by the user, use the keyword large - model to perform semantic understanding and keyword extraction, and then generate keywords related to the title; The specific steps of step S3 are as follows: S31. Pre - store keywords in the faiss vector database in advance, construct a keyword knowledge base, and create an ES index; S32. After obtaining the title input by the user and performing pre - processing, input it into the keyword large - model for semantic understanding, and extract the initial keywords; S33. Judge whether the keyword data volume is sufficient; If so, go to step S36; If not, go to step S34; S34. Use the Elasticsearch tool to retrieve the ES index in the keyword knowledge base in combination with the title, and further extract keywords; S35. Judge whether the number of keywords is sufficient; If so, go to step S36; If not, go to step S37; S36. Use the BGE tool to perform cross - language retrieval of relevant documents according to the title, extract cross - language keywords, and return to step S35; S37. Merge the keywords, and then match the keywords with the knowledge base to obtain the final keyword list; It should be noted that by constructing a keyword knowledge base and creating an index, it is convenient for the storage and retrieval of keywords. By extracting keywords through multiple methods such as model extraction, knowledge base retrieval, and cross - language retrieval, the comprehensiveness and accuracy of keywords are improved; By merging and matching keywords, the final keyword list is obtained, providing a data basis for subsequent outline generation; S4. In response to the title and keywords input by the user, combine the knowledge graph and the deep - learning model to generate a structured outline of the work report; The specific steps of step S4 are as follows: S41. Pre - construct a knowledge graph of the historical work report outline using the neo4j library of python and store it; S42. Obtain the title entered by the user and the extracted keywords, preprocess the title and keywords, and then use the Elasticsearch tool to query the knowledge graph; If the query result meets the requirements, generate a preliminary outline and proceed to step S46; If the query result does not meet the requirements, proceed to step S43; S43. Input the title and keywords into the deep learning model to output the relevance between the title and keywords, and obtain the candidate outline; S44. Receive user feedback and determine whether the candidate outline is valid; If yes, proceed to step S46; If no, proceed to step S45; S45. Adjust the candidate outline according to the user feedback, respond to the user's confirmation of the adjustment of the candidate outline, and return to step S44; S46. Merge and optimize the candidate outline or the preliminary outline, and perform a tree-structured display to obtain the keywords of the lowest-level outline, and retrieve reference materials, which are used to provide specific content support for subsequent paragraph generation; If there is sufficient reference material, proceed to step S48; If there is insufficient reference material, proceed to step S47; Exemplarily, the tree-structured display of the outline is as follows: First-level title (such as "Economic Development") ├─ Second-level title (such as "GDP Growth") │└─ Third-level title (such as "Data in 2023") └─ Second-level title (such as "Employment Policy") S48. Further collect materials, and after the user reviews and confirms, return to step S46; S47. Display the reference materials and generate the structured outline of the final work report; It should be noted that by constructing and querying the knowledge graph, using the structured information of the knowledge graph to assist in generating a reasonable preliminary outline; combining the deep learning model to generate a candidate outline and adjusting it according to user feedback improves the applicability of the outline; merging, optimizing and tree-structured displaying the outline makes the outline clear and convenient for subsequent paragraph generation; S5. Combine the structured outline with the data and materials in the knowledge base, and use the paragraph large model to generate paragraphs in a loop until the generation of the work report is completed; The specific steps of step S5 are as follows: S51. Initialize the paragraph generation parameters according to the structured outline, retrieve the data and reference materials in the knowledge base according to the structured outline, and determine whether the reference materials and data are sufficient; If yes, proceed to step S53; If not, go to step S52; It should be noted that Elasticsearch (ES) and BGE are used for material knowledge retrieval. The retrieved knowledge is further segmented and then ranked by the Reranker model to improve the recall accuracy, ensuring that the materials referenced by each outline are highly relevant articles; The Reranker model is used to re-rank the documents retrieved by es and BGE to increase the priority of relevant documents; Recall rate = (number of relevant documents retrieved) / (total number of relevant documents) The formula is expressed as: Recall = TP / (TP + FN) Among them, TP represents true positive cases, that is, relevant documents that are retrieved, and FN represents false negative cases, that is, relevant documents that are not retrieved; S52. Supplement materials and data, and return to step S51; S53. Use the paragraph large model to start paragraph generation in combination with reference materials and data, apply the multi-head attention mechanism to generate the initial paragraph draft, and verify whether the quality of the initial paragraph draft meets the requirements; If so, go to step S55; If not, go to step S54; S54. Regenerate the paragraph and return to step S53; S55. Optimize the initial paragraph draft to obtain the final paragraph draft, and determine whether all paragraphs have been completed; If so, go to step S57; If not, go to step S56; S56. Start generating the next paragraph and return to step S51; S57. Integrate all paragraphs to generate a draft report and provide it to the user; S58. Receive user feedback; If the user is satisfied, generate the final work report and end; If the user is not satisfied, go to step S59; S59. Adjust the draft report according to the user feedback and return to step S58; It should be noted that retrieving materials and data according to the structured outline ensures the data support for paragraph generation; using the paragraph large model and the multi-head attention mechanism to generate the initial paragraph draft, and performing quality verification and optimization improves the coherence of the paragraph; by integrating the paragraph to generate the draft report and adjusting according to the user feedback, the personalized needs of the user are thus met.

[0034] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0035] As Figure 2 shown below are embodiments of a large model-based work report generation system provided by the embodiments of the present disclosure. This system and the large model-based work report generation methods of the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiments of the large model-based work report generation system, reference can be made to the embodiments of the large model-based work report generation method above.

[0036] The system includes: A knowledge base management module that manages a knowledge base storing data and materials related to work report generation, and is provided with retrieval and document management interfaces; A keyword generation module that, in response to a title input by a user, uses a large model for semantic understanding and keyword extraction to generate keywords related to the title; An outline generation module that, in response to a title and keywords input by a user, combines a knowledge graph and a deep learning model to generate a structured outline of the work report; A paragraph generation module that, according to the structured outline, combines the data and materials in the knowledge base, and uses the large model to generate paragraphs in a loop until the generation of the work report is completed; It should be noted that generating paragraphs in a loop according to the structured outline until the work report is completed realizes the automation and efficiency of report generation.

[0037] In this embodiment, the knowledge base management module ensures the effective management and utilization of data, the keyword generation module accurately extracts keywords to clarify the core points for work report generation; the outline generation module generates a structured outline to construct a framework for the work report, and the paragraph generation module realizes the automatic generation of the content of the work report, improving the generation efficiency and quality.

[0038] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another large model-based work report generation system is provided. This system includes: A knowledge base management module that manages a knowledge base storing data and materials related to work report generation, and is provided with retrieval and document management interfaces; A keyword generation module that, in response to a title input by a user, uses a large model for semantic understanding and keyword extraction to generate keywords related to the title; An outline generation module that, in response to a title and keywords input by a user, combines a knowledge graph and a deep learning model to generate a structured outline of the work report; The paragraph generation module generates paragraphs in a loop using a large model according to the structured outline, in combination with the data in the knowledge base, until the generation of the work report is completed.

[0039] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating a work report based on a large model, characterized in that: The steps include: S1. Pre-build and manage a knowledge base for storing data related to work report generation; S2. Pre-training a keyword model for generating keywords, a deep learning model for generating outlines, and a paragraph model for generating paragraphs; S3. In response to the title input by the user, generate keywords related to the title after semantic understanding and keyword extraction using the keyword big model; S4. In response to the title and keywords input by the user, generate a structured outline of the work report by combining the knowledge graph and the deep learning model; S5. Combine the structured outline with the data in the knowledge base, and use the paragraph model to generate paragraphs in a loop until the work report is completed.

2. The method for generating a work report based on a large model according to claim 1, characterized in that: The specific steps of step S1 are as follows: S11. Initialize the knowledge base and determine the knowledge type according to the work report type; the knowledge type includes policy documents related to the work report, meeting minutes, industry data and reference materials; S12. Set configuration permissions for the knowledge base and configure whether the automatic knowledge collection mode is enabled through management users; S13. When the automatic collection mode is on, use the web crawler mode to automatically collect policy documents, meeting minutes and industry data related to the work report from the Internet, and manually upload the data when the automatic collection mode is off; S14. According to the type of knowledge, the collected data is intelligently screened and sorted through the big model and the automation tool RPA, and the identified data as knowledge is stored in the knowledge base; S15. Classify and index the information in the knowledge base, and provide a retrieval interface after the classification is completed; S16. Respond to the user's search request through the search interface, perform a search operation based on the search request and the knowledge base index, and obtain the search results; If the search results do not meet the user's needs, go to step S17; If the search results meet the user's needs, go to step S18; S17. After receiving user feedback information and optimizing the search method, return to step S16; S18. Display the search results to the user and provide an access and download interface; S19. Review and maintain the knowledge base regularly, and update the information in the knowledge base when the knowledge base needs to be updated.

3. The method for generating a work report based on a large model according to claim 2, characterized in that: The specific steps of step S2 are as follows: S21. Collect text data related to the work report in advance, covering the fields and topics involved in the work report, construct the first data set, and collect titles and corresponding keywords related to the work report in historical work reports and policy documents as task data; S22. Pre-training the model of the Transformer architecture using the first data set, and optimizing using the language modeling task during the pre-training process to obtain a pre-trained model; S23. Connect a classification layer to the output of the pre-trained model, set the title in the task data as input, and set the keyword as output, complete the fine-tuning target setting, and then use the task data for fine-tuning training, combined with the loss function for optimization, to obtain the keyword large model; S24. Pre-collect text data of historical work reports, wherein the text data of the historical work reports includes historical titles, keywords, work report contents, and work report outlines, and construct a second data set; S25. Select a bidirectional LSTM network as the RNN model, define the input layer and output layer of the RNN model, and add an activation function, so that the historical titles, keywords, and work report contents in the second data set correspond to the input layer, and the historical work report outlines in the training set correspond to the output layer; S26. Use the tensorflow framework to train the RNN model, and use the Adam optimizer to adjust the learning rate during the training process, use the cosine similarity function to calculate the correlation between the title and the keyword, and use the cross entropy loss function to calculate the difference between the outline generated by the RNN model and the outline of the historical work report in the second data set, and obtain the deep learning model when the difference meets the requirements; S27. Collect text data related to the work report in advance, covering the fields and topics involved in the work report, clean and segment the collected data, then construct a vocabulary, and map each word to a unique index to obtain a third data set, and collect data related to the paragraphs of the work report and organize them according to the lowest level outline goals to obtain a fourth data set; the paragraph-related data includes a structured outline, title, keywords, reference materials, and paragraph content, and the lowest level outline goals include titles, keywords, reference materials, and paragraphs to be generated; the reference materials are used to assist the paragraph large model in learning and generating paragraph content that conforms to the actual situation; S28. Select a large language model based on the Transformer architecture as the basic model, then use the masked language model and the next prediction task, combined with the third data set to pre-train the paragraph large model, and calculate the total loss through the cross entropy loss function during the pre-training process until the total loss meets the requirements; S29. Set the fine-tuning target to generate paragraph content according to the title, keywords, and reference materials of the lowest-level outline target, use the fourth data set to adjust the pre-trained paragraph model according to the fine-tuning target, and use the cross-entropy loss function to compare the difference between the paragraphs generated by the paragraph model and the paragraphs in the fourth data set until the difference meets the requirements, and obtain the final paragraph model.

4. The method for generating a work report based on a large model according to claim 3, characterized in that: The specific steps of step S3 are as follows: S31. Use the faiss vector database to store keywords in advance, build a keyword knowledge base, and create an ES index; S32. After obtaining the title input by the user and preprocessing it, the keyword model is input for semantic understanding and the initial keywords are extracted; S33. Determine whether the amount of keyword data is sufficient; If yes, go to step S36; If not, proceed to step S34; S34. Use the Elasticsearch tool to search the ES index in the keyword knowledge base in combination with the title to further extract keywords; S35. Determine whether the number of keywords is sufficient; If yes, go to step S36; If not, proceed to step S37; S36. Use the BGE tool to perform cross-language retrieval of relevant documents based on the title, extract cross-language keywords, and return to step S35; S37. Merge the keywords, and then match the keywords with the knowledge base to obtain a final keyword list.

5. The method for generating a work report based on a large model according to claim 4, characterized in that: The specific steps of step S4 are as follows: S41. Use the neo4j library of python to construct the knowledge graph of the outline of historical work reports in advance and store it; S42. Obtain the title input by the user and the extracted keywords, and use the Elasticsearch tool to query the knowledge graph after preprocessing the title and keywords; If the query result meets the requirements, a preliminary outline is generated and the process proceeds to step S46; If the query result does not meet the requirements, go to step S43; S43. Input the title and keywords into the deep learning model to output the correlation between the title and the keywords, and obtain the candidate outline; S44. Receive user feedback and determine whether the candidate outline is valid; If yes, go to step S46; If not, proceed to step S45; S45. Adjust the candidate outline according to user feedback, and respond to the user's confirmation of the candidate outline adjustment, and return to step S44; S46. Merge and optimize the candidate outlines or preliminary outlines, and display them in a tree structure, obtain the keywords of the lowest level outline, and retrieve reference materials, which are used to provide specific content support when generating subsequent paragraphs; If the reference material is sufficient, proceed to step S48; If the reference material is insufficient, proceed to step S47; S48. Further collect materials, and after the user's review and confirmation, return to step S46; S47. Display reference materials and generate a structured outline of the final work report.

6. The method for generating a work report based on a large model according to claim 5, characterized in that: The specific steps of step S5 are as follows: S51. Initialize paragraph generation parameters according to the structured outline, retrieve data and reference materials in the knowledge base according to the structured outline, and determine whether the reference materials and data are sufficient; If yes, go to step S53; If not, proceed to step S52; S52. Supplement materials and data, and return to step S51; S53. Use the large paragraph model combined with reference materials and data to start paragraph generation, apply the multi-head attention mechanism to generate the first draft of the paragraph, and verify whether the quality of the first draft of the paragraph meets the requirements; If yes, go to step S55; If not, proceed to step S54; S54. Regenerate the paragraph and return to step S53; S55. Optimize the first draft of the paragraph to obtain the final draft of the paragraph, and determine whether all paragraphs have been completed; If yes, go to step S57; If not, proceed to step S56; S56. Start the next paragraph generation and return to step S51; S57. Integrate all paragraphs to generate a draft report and provide it to the user; S58. Receive user feedback; If the user is satisfied, a final work report is generated and the process ends; If the user is not satisfied, proceed to step S59; S59. Adjust the draft report according to user feedback and return to step S58.

7. A work report generation system based on a large model, characterized in that: include: The knowledge base management module manages the knowledge base that stores data related to work report generation, and has a retrieval and document management interface; A keyword generation module, in response to a title input by a user, generates keywords related to the title after semantic understanding and keyword extraction using a large model; The outline generation module generates a structured outline of the work report in response to the title and keywords input by the user, combining the knowledge graph and deep learning model; The paragraph generation module generates paragraphs based on the structured outline, combined with the data in the knowledge base, and uses a large model to loop through the generation of the work report until it is complete.

8. The large model-based work report generation system according to claim 7, characterized in that: The knowledge base management module includes: A knowledge base initialization unit is used to initialize the knowledge base; The knowledge collection mode configuration unit sets the configuration permissions for the knowledge base and configures whether the automatic knowledge collection mode is enabled through management users; The knowledge collection unit, when the automatic collection mode is turned on, uses a web crawler to automatically collect policy documents, meeting minutes, and industry data related to the work report from the Internet, and manually uploads the data when the automatic collection mode is turned off; The knowledge base storage unit intelligently screens and organizes the collected data according to the knowledge type and through the big model and automation tool RPA, stores the identified data as knowledge in the knowledge base, and classifies and indexes them; The knowledge retrieval unit responds to the user's retrieval request through the retrieval interface, performs the retrieval operation according to the retrieval request and the index of the knowledge base, and obtains the retrieval result; The knowledge base maintenance unit regularly reviews and maintains the knowledge base, and updates the information in the knowledge base when the knowledge base needs to be updated.

9. The work report generation system based on a large model according to claim 8 is characterized in that: The keyword generation module includes: The keyword knowledge base construction unit uses the Faiss vector database to store keywords in advance, builds the keyword knowledge base, and creates an ES index; The keyword extraction unit obtains the title input by the user and performs preprocessing, then inputs the keyword model for semantic understanding and extracts the initial keywords; The keyword knowledge base retrieval unit uses the Elasticsearch tool to search the ES index in the keyword knowledge base in combination with the title to further extract keywords; The keyword cross-language search unit uses the BGE tool to perform cross-language search of relevant documents based on titles and extract cross-language keywords when the number of keywords is insufficient; The keyword merging unit merges the keywords and then matches the keywords with the knowledge base to obtain the final keyword list.

10. The work report generation system based on a large model according to claim 9, characterized in that: The outline generation module includes: The knowledge graph construction unit uses the neo4j library of python to construct the knowledge graph of the outline of historical work reports in advance and stores it; The knowledge graph query unit obtains the title input by the user and the extracted keywords, and uses the Elasticsearch tool to query the knowledge graph after preprocessing the title and keywords; The candidate outline generation unit, when the query result does not meet the requirements, inputs the title and keywords into the deep learning model to output the correlation between the title and the keywords, and obtains the candidate outline; The outline tree display unit merges and optimizes the candidate outlines or preliminary outlines, displays them in a tree structure, obtains the keywords of the lowest level outline, and retrieves reference materials; The structured outline generation unit displays the reference materials and generates a structured outline of the final work report when there are sufficient reference materials; The paragraph generation module includes: A knowledge base retrieval unit initializes paragraph generation parameters according to the structured outline, retrieves data and reference materials in the knowledge base according to the structured outline, and determines whether the reference materials and data are sufficient; The paragraph draft generation unit uses the paragraph model combined with reference materials and data to start paragraph generation, applies the multi-head attention mechanism to generate paragraph drafts, and verifies whether the quality of the paragraph draft meets the requirements; The report draft generation unit integrates all the paragraphs to generate a report draft and provides it to the user when all the paragraphs are completed.

Citation Information

Cited By

  • Self-supervision method and system for promoting chat robot to utilize unintegrated service

    CN120849563A

  • Method and system for generating intelligent insight report based on AI large model

    CN121328504A