Method and system for automatically generating document based on template and large model

The method of automatically generating meeting documents by using templates and large models solves the problems of inconsistent formatting and inaccurate content caused by manual writing, and achieves efficient and consistent document generation and verification.

CN121009874APending Publication Date: 2025-11-25TERMINUSBEIJING TECH CO LTD

Patent Information

Application Number
CN202511062498.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In existing technologies, the generation of meeting documents relies on manual writing, which leads to inconsistencies in format and difficulty in ensuring the accuracy of content, resulting in low efficiency.

Method used

An automatic generation method based on templates and large models is adopted. By parsing multimodal files to extract structured data, dynamically combining templates, calling large models to generate project background and content summaries, and performing intelligent cross-document consistency verification and error correction, the final standardized meeting review documents are formed.

Benefits of technology

Ensuring consistent document formatting and accurate content improves generation efficiency, reduces the time cost of manual operations, and enables automated document merging and intelligent verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009874A_ABST
    Figure CN121009874A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and system for automatically generating a document based on a template and a large model, and relates to the technical field of document generation, the method comprises the following steps: step S1, a user inputs information through a browser interface and uploads a to-be-reviewed file; step S2, analyzing the uploaded multi-modal file and extracting structured data; s3, outputting a standardized template instance file according to the decision tree dynamic combination template; s4, calling a large model to generate a project background and a file content abstract; s5, the generated content is rendered and combined with the attachment to form a complete conference review file; s6, executing cross-document consistency intelligent verification and executing intelligent error correction; and S7, generating and storing a final document with version traceability information. According to the method, on the basis of the document template and the large model capacity, document formats are unified, content is accurately generated, and the document writing efficiency is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of document generation, and in particular to a method and system for automatically generating a document based on a template and a large model. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, intelligent means are widely used in various industries, including network media, public security, and office scenarios. As a typical application form of artificial intelligence technology in the office field, intelligent conference systems rely on AI capabilities to realize the automation of conference management and information processing, not only providing strong support for organizational decision-making, but also playing an important role in improving office efficiency and standardizing office processes.

[0003] When organizing a file review or discussion meeting, the meeting organizer needs to prepare the meeting document for the review file. Currently, this document is mainly prepared manually, which has many drawbacks. On the one hand, in the case of multiple preparations or different personnel preparations, it is difficult to ensure the uniformity of the document format, and different versions of the document differ significantly in font, font size, paragraph spacing, and other layout elements, affecting the standardization and professionalism of the enterprise document. On the other hand, manual preparation is often based on past document modifications, which can easily result in content omissions or errors. More importantly, manual preparation requires reading each file and manually preparing the background and abstract, which not only consumes a lot of time and effort, but also is extremely inefficient. The meeting document for the review file of an enterprise has fixed format requirements. In terms of content composition, the cover page covers information such as the main title, meeting topic, reporting department, and reporter. The content page includes sections such as the meeting topic, opening remarks, basic information (related to the project background, main content, and solicitation of opinions), items for review, and attachments. In terms of style, each document section has specific requirements for font, font size, font color, alignment, etc. Among them, the content section varies depending on the situation of each meeting, especially the project background and main content sections. The project background needs to be summarized and refined based on the contents of all the files to be reviewed at the meeting, and the main content needs to list the names and content summaries of each document to be reviewed, and the attachments need to add all the original contents of the documents to be reviewed in order.

[0004] A Chinese patent with publication number CN120340497A discloses an intelligent meeting automatic meeting recording and summary generation method, including the following steps: real-time capture of multi-speaker voice stream through directional microphone array, generation of original text stream with timing markers based on real-time voiceprint clustering, and synchronous extraction of intent intensity parameters in the voice stream; intent-driven dynamic segmentation processing of the original text stream; agenda-aware summary block generation on the segmented text, including extracting decision statements that meet the semantic density threshold as summary core blocks; and assembling the summary core blocks into a structured summary document according to the hierarchical structure of the agenda template. Compared with the traditional linear transcription method based on speech recognition, the invention significantly improves the mapping accuracy between the content of the speech and the identity of the participants, providing a solid foundation for semantic segmentation and decision tracing.

[0005] A Chinese patent with publication number CN110493019B discloses a conference minutes automatic generation method, device, equipment and storage medium. The method includes: after receiving a conference minutes recording request sent by a group in a communication application program that is currently in session at a request time point, obtaining the session content information of each participant in the group after the request time point; determining whether the session is ended according to the session interruption duration; if the session is ended, recording the end time point of the session end, and obtaining text information after preprocessing the session content information; identifying the text information through a semantic recognition model to obtain a semantic recognition result; extracting theme content in the semantic recognition result through a document theme generation model; generating a conference minutes according to the obtained conference parameters and theme content, and sending the conference minutes to the conference host after confirming the conference minutes, and then sending the conference minutes to each participant in the participant list. The method provided by the invention can improve work efficiency.

[0006] The above patents all have the problem proposed in the background art: the existing solutions are generally manually operated by the conference organizer, and are based on the documents used in previous meetings for writing. The specific process is as follows: collate the list of files to be reviewed, fill in the relevant information of the current meeting, write the project background content, write the content summary of each file, paste the original content of each file to be reviewed as an attachment, and finally complete the preparation of the conference document for the review file. This process completely relies on manual operation, and it is difficult to guarantee the consistency of the format and the accuracy of the content, and the efficiency is low. SUMMARY

[0007] The technical problem to be solved by the present application is to provide a method for automatically generating documents based on templates and large models.

[0008] To achieve the above-mentioned purpose, the technical solution adopted by the present application is:

[0009] In a first aspect, the embodiments of the present application provide a method for automatically generating a document based on a template and a large model, comprising the following steps:

[0010] Step S1, a user inputs information and uploads a to-be-reviewed file through a browser interface;

[0011] Step S2, parsing the uploaded multi-modal file and extracting structured data;

[0012] Step S3, dynamically combining templates according to a decision tree to output a standardized template instance file;

[0013] Step S4, calling a large model to generate a project background and a file content summary;

[0014] Step S5, rendering the generated content and merging with attachments to form a complete conference review file;

[0015] Step S6, performing intelligent checking for cross-document consistency and executing intelligent error correction;

[0016] Step S7, generating a final document with version trace information and storing it.

[0017] Further, the step S1 specifically comprises:

[0018] receiving a dynamic template identifier and an associated decision rule set selected by a user through a browser interface, wherein the decision rule set is defined by a Boolean logic expression;

[0019] receiving template variable values filled by a user and multi-modal to-be-reviewed files uploaded by the user, and the supported file types include PDF, DOCX, PPT, and image scans.

[0020] Further, the step S2 specifically comprises:

[0021] establishing a file type routing mechanism to distribute files to corresponding parsing engines according to file extensions and content characteristics;

[0022] The parsing engine performs hierarchical parsing, including: applying a layout segmentation model to PDF files, implementing OOXML deep parsing to DOCX files, and using end-to-end OCR table reconstruction for image scans;

[0023] extracting structured data, including text paragraphs, table data, and image titles, from the parsing results, and constructing a unified JSON intermediate representation;

[0024] The OCR table reconstruction specifically comprises the following steps:

[0025] inputting an original table image to generate a candidate structure;

[0026] Calculate the cell alignment and the empty cell rate respectively;

[0027] Construct a table reconstruction optimization function, and the calculation formula is:

[0028] Γ(T)=argmin[0.7·D cell (T,S)+0.3·Empty rate (S)]

[0029] Wherein, Γ(T) represents the table reconstruction optimization function, T represents the original table area, S represents the candidate table structure, D cell represents the cell alignment, and Empty rate represents the empty cell rate;

[0030] The candidate table structure output that makes the table reconstruction optimization function minimum is selected as the optimal table structure.

[0031] Further, the step S3 specifically comprises:

[0032] Recursively search for candidate components from a pre-constructed tree structure template library, wherein the candidate components include: core components including: title page, approval process table and risk analysis matrix; conditional components including: high-amount approval module and emergency plan chapter; nested sub-templates including: legal clause appendix and expert review opinion table;

[0033] Perform conditional function matching on each candidate component, specifically including: verifying whether the user input variables meet the component activation conditions, checking the enterprise approval system requirements, confirming whether the preposed components have been activated, and outputting the component matching degree score;

[0034] Perform hierarchical decision-making according to the matching degree score, wherein when the component matching degree score≥0.8, strong activation is performed, the component is immediately inserted and the position is locked; when 0.6≤component matching degree score<0.8, weak activation is performed, a placeholder is added and manual confirmation is prompted; when the component matching degree score<0.6, no activation is performed, and the component is skipped; generate a component activation log: record the decision basis and variable traceability;

[0035] Assign style weights according to component priority, wherein the component priority is divided into 5 levels, priority 1 component is a legal force core component, including contract signing page, approval signature column and seal area; priority 2 component is a decision key content component, including risk analysis matrix, financial data table and resolution clause; priority 3 component is a business logic necessary component, including project timeline, technical parameter table and resource allocation diagram; priority 4 component is an auxiliary explanation component, including data source annotation, term explanation box and reference appendix; priority 5 component is a replaceable component, including decorative header and footer, background watermark and promotional slogan;

[0036] Performing three-layer style fusion and outputting a standardized template instance file, wherein the base layer applies enterprise global style specifications; the component layer mixes component self-contained styles according to weights; and the dynamic adjustment layer detects conflicting attributes and assigns final values according to weights.

[0037] Further, the step S4 specifically includes:

[0038] Building knowledge-enhanced prompt words, including: retrieving relevant clauses, historical cases, and industry specifications from an enterprise private knowledge base, generating semantic vector encoding, converting user-filled variable values into structured feature vectors, integrating multi-modal analysis output text and table data, and generating an abstract vector;

[0039] Calling a large language model with domain knowledge enhancement to process prompt words and generate project background and file content abstracts;

[0040] Adding a confidence score to the AI-generated content, and triggering a manual review flag when the confidence score is below a threshold.

[0041] Further, the step S5 specifically includes:

[0042] Using the Docxtemplater engine to render the dynamically combined template with variable values and AI-generated content to generate an initial DOCX document;

[0043] Calling a document merging component to attach user-uploaded files to the end of the initial document according to a preset sorting rule to form a complete meeting review document.

[0044] Further, the step S6 specifically includes:

[0045] Using a three-level scanning mechanism to scan document elements, detect and mark all format abnormalities that deviate from standard values;

[0046] Based on an entity association engine, extracting key data from multiple documents, constructing a conflict detection matrix to identify numerical contradictions, term differences, and logical conflicts, and visualizing the distribution of conflict points through a heat map;

[0047] According to the error type, performing intelligent error correction, wherein format errors call template style reset, content conflicts trigger manual confirmation processes, compliance issues associate with the knowledge base to automatically replace outdated clauses and add desensitization processing;

[0048] Outputting a visual audit view containing a hierarchical problem list, a three-dimensional health score, and executable modification instructions, and synchronously recording a version bloodline map for traceability analysis.

[0049] Further, the step S7 specifically includes:

[0050] Construct a document generation graph to record the source of each data item, mark the content conversion path and timestamp;

[0051] Generate a difference visualization report to distinguish AI-generated content from manually modified parts by color;

[0052] Store the final document in cloud storage, generate a unique file ID and access link, and return the downloadable file package to the user interface.

[0053] In a second aspect, the embodiments of the present application also provide a system for automatically generating documents based on templates and large models, which is used to implement the method for automatically generating documents based on templates and large models in the first aspect, and includes:

[0054] A dynamic template management engine for storing a tree structure template component library and a versioned decision rule set, and performing conditional trigger type template combination and style conflict arbitration;

[0055] A multi-modal data fusion module for realizing heterogeneous file parsing and structured data extraction through layout segmentation, table reconstruction, and entity linking technology;

[0056] An AI generation unit with knowledge enhancement for fusing enterprise knowledge base and risk matrix to generate a draft of review opinions with confidence control;

[0057] A cross-document consistency checker for implementing three-order cross-document intelligent checking of format specifications, content conflicts, and compliance clauses, and performing minimal intervention correction on detected errors;

[0058] A visual version traceability interface for realizing document full life cycle management through data bloodline graph recording of content source and conversion path;

[0059] An intelligent evaluation and repair module for evaluating document health and performing repair on checking errors.

[0060] Further, the intelligent evaluation and repair module includes:

[0061] A document health evaluation dashboard for displaying a three-dimensional health score in real time and generating a list of improvement suggestions;

[0062] An automatic repair engine for performing minimal change repair on checking errors and retaining repair history records;

[0063] An intelligent recommendation module for recommending relevant clauses based on context and actively prompting missing information items.

[0064] Compared with the prior art, the present application has the following beneficial effects:

[0065] 1、The application uses document template technology to dynamically render documents, fundamentally ensuring the format uniformity of the generated final document. Regardless of how many times the document generation operation is performed or initiated by different personnel, the font, font size, paragraph, and other layout elements of the document remain consistent, while ensuring accurate content and avoiding errors and omissions caused by manual writing.

[0066] 2、The application introduces large model technology to generate project background and content summary, which greatly improves work efficiency compared with manual writing. Large models can quickly process large amounts of file content and generate high-quality backgrounds and summaries, reducing the time cost of manual reading and writing.

[0067] 3、The application uses document merging technology to merge document content, achieving automatic addition and integration of attachments without manual copying and pasting, further improving work efficiency and making the conference document generation process more efficient and convenient.

[0068] It should be understood that the content described in the summary section is not intended to limit the key or important features of the embodiments of the application, nor to limit the scope of the application. Other features of the application will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0069] The above and other features, advantages, and aspects of the embodiments of the present application will become more apparent with reference to the following detailed description when considered in conjunction with the accompanying drawings. The drawings are intended to better understand the present application and do not limit the present application. In the drawings, the same or similar reference numerals refer to the same or similar elements, wherein:

[0070] Figure 1 A flowchart of a method for automatically generating a document based on a template and a large model according to an embodiment of the present application;

[0071] Figure 2 A schematic diagram of the technical framework of an embodiment of the present application;

[0072] Figure 3 A schematic diagram of a user operation page according to an embodiment of the present application;

[0073] Figure 4 A system structure diagram according to an embodiment of the present application;

[0074] Figure 5 A schematic block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0075] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0076] In addition, the term "and / or" herein merely describes an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects.

[0077] Figure 1 A flowchart of a method for automatically generating a document based on a template and a large model according to an embodiment of the present application. As shown in Figure 1 The method for automatically generating a document based on a template and a large model includes the following steps:

[0078] Step S1, a user inputs information and uploads a file to be reviewed through a browser interface;

[0079] Step S2, parsing the uploaded multi-modal file and extracting structured data;

[0080] Step S3, dynamically combining templates according to a decision tree to output a standardized template instance file;

[0081] Step S4, calling a large model to generate a project background and a file content summary;

[0082] Step S5, rendering the generated content and merging with attachments to form a complete meeting review file;

[0083] Step S6, performing intelligent checking for cross-document consistency and performing intelligent error correction;

[0084] Step S7, generating a final document with version trace information and storing it.

[0085] The step S1 specifically includes:

[0086] Receiving a dynamic template identifier selected by a user and an associated decision rule set through a browser interface, wherein the decision rule set is defined by a Boolean logic expression;

[0087] Receiving template variable values filled by a user and multi-modal files to be reviewed uploaded by the user, and the supported file types include PDF, DOCX, PPT and image scans.

[0088] The step S2 specifically comprises:

[0089] A file type routing mechanism is established to distribute files to corresponding parsing engines according to file extensions and content features;

[0090] The parsing engine performs hierarchical parsing, including applying a layout segmentation model to PDF files, implementing OOXML deep parsing on DOCX files, and using end-to-end OCR table reconstruction on image scans;

[0091] Structured data including text paragraphs, table data and image titles is extracted from the parsing results, and a unified JSON intermediate representation is constructed;

[0092] The loss function calculation formula of the layout segmentation model is:

[0093] L = 0.6L mask + 0.3L class + 0.1L box

[0094] Wherein, L represents the total loss value, L mask represents the mask segmentation loss, L class represents the classification loss, L box represents the bounding box loss;

[0095] The implementation of OOXML deep parsing specifically comprises:

[0096] Decompressing the DOCX file: decompressing the DOCX file as a ZIP compressed package to obtain its internal file structure.

[0097] Parsing the document structure: starting from document.xml, using an XML parser (such as DOM or SAX) to parse the nodes of the document. The document is composed of elements such as paragraphs, tables, pictures, etc. in order.

[0098] Processing paragraphs: for each paragraph, parse the text runs and text nodes therein.

[0099] Processing tables: for tables, parse each row and cell.

[0100] Processing pictures and embedded objects: by parsing the w: drawing node, obtain the embedded information of the picture

[0101] Parsing style inheritance: the style in DOCX is hierarchical inheritance. Specifically, direct format: the style attribute specified directly in the element has the highest priority. Paragraph style: the style referenced by the w: pStyle element has the second highest priority. Document default style: the default style defined in styles.xml has the lowest priority. The style inheritance relationship needs to be recursively parsed to calculate the actual style of each element.

[0102] Extract variable markers: Find variables in Docxtemplater format in text nodes. Record the name of each variable, the paragraph it is in, the text run, and style information to preserve styles during later replacement.

[0103] Build document object model: Organize parsed content, styles, tables, images, etc. into a structured document object.

[0104] Resource handling: Copy resources like images encountered during parsing to a designated location and record the resource's reference path in the document object.

[0105] Generate structured output: Output a structured data representation (e.g., JSON) of the document's content, structure, and styles, which can be used for subsequent template rendering, content extraction, etc.

[0106] The OCR table reconstruction specifically includes the following steps:

[0107] Input the original table image to generate candidate structures;

[0108] Calculate the cell alignment and empty cell rate respectively;

[0109] Construct a table reconstruction optimization function, the formula is:

[0110] Γ(T)=arg min[0.7·D rate (T,S)+0.3·Empty cell (S)]

[0111] Where Γ(T) represents the table reconstruction optimization function, T represents the original table region, S represents the candidate table structure, D rate represents the cell alignment, and Empty k represents the empty cell rate;

[0112] Select the candidate table structure that minimizes the table reconstruction optimization function as the optimal table structure.

[0113] The loss function is the core loss function of MaskR-CNN, used for document element segmentation:

[0114] Input: Document image, such as a scan, output: Structured document elements, such as text blocks / tables / pictures;

[0115] Optimization goal: Simultaneously improve the accuracy of the three categories: 60% weight to ensure cell-level table line detection, 30% weight to ensure correct identification of title / body type, and 10% weight to optimize region positioning.

[0116] Mask R-CNN includes three main sub-networks: backbone network, RPN network, and head network.

[0117] The step S3 specifically includes:

[0118] Based on the decision tree, the component selection is performed, and candidate components are recursively searched from a pre-built tree structure template library, wherein the candidate components include: core components including: title page, approval process table, risk analysis matrix; conditional components including: high-amount approval module, emergency plan chapter; nested sub-templates including: legal clause appendix, expert review opinion table;

[0119] For each candidate component, a conditional function matching is performed, specifically including: verifying whether the user input variable meets the component activation condition, checking the enterprise approval system requirements, confirming whether the pre-component has been activated, and outputting a component matching degree score;

[0120] The calculation formula of the component matching degree score is:

[0121]

[0122] Wherein, represents the matching degree score of the candidate component C k , f j represents the jth conditional function, v j represents the input variable value, p represents the number of conditional functions, and k represents the number of components.

[0123] When all conditions are completely met, i.e., all f j = 1:

[0124] When any condition is not met, i.e., there is f j = 0: That is, the component is immediately eliminated.

[0125] In the conditional function, each f j returns a value between 0 and 1, representing the degree of satisfaction of a single condition. When all conditions are completely met, the score is 1, and any condition not met will be eliminated.

[0126] Specific examples of the conditional function are:

[0127] 1. Project amount threshold function (financial approval component)

[0128] The user input value of the "project amount" is:

[0129] ≥ 1 million: complete match (1 point) → strong activation of senior approval table

[0130] 50-100 million: partial match (0.8 points) → weakly activate the basic approval form

[0131] <50 million: low match (0.3 points) skip this component;

[0132] 2. Risk level matching function (emergency plan component)

[0133] User fill in the "risk description" text:

[0134] Keyword detection: if the description contains any of the keywords "explosion", "leakage", "collapse", it gets a maximum of 1 point,

[0135] Semantic similarity: calculate the cosine similarity between the description and "high" (maximum 0.2 points);

[0136] 3. Legal clause dependency function (signature column component)

[0137] Both "legal clause appendix" and "approval process table" must be activated to insert the signature column, and the minimum value of the scores of the two (bucket effect) is taken,

[0138] Example: legal clause score = 0.9, approval process score = 0.6, final score = 0.6;

[0139] 4. Time sensitivity function (urgent identification component)

[0140] No urgency if the deadline is more than 7 days away, condition function value 1, weak reminder if the deadline is 3 days away, condition function value 0.7, strong warning if overdue, condition function value 0.33.

[0141] Function design features:

[0142] Threshold type (such as amount): stepwise segmentation function

[0143] Semantic type (such as risk): combination of keywords and vector similarity

[0144] Dependency type (such as signature column): associated with other component scores

[0145] Decay type (such as urgency): time sensitivity index decay

[0146] According to the matching score, execute hierarchical decision-making, where component matching score ≥ 0.8 is strongly activated, immediately insert the component and lock the position; 0.6 ≤ component matching score < 0.8 is weakly activated, add a placeholder and prompt manual confirmation; component matching score < 0.6 is not activated, skip this component; generate component activation log: record decision basis and variable traceability;

[0147] The style weight is assigned according to the component priority, wherein the component priority is divided into five levels, priority 1 component is a legal validity core component, including contract signing page, approval signature column and seal area; priority 2 component is a decision key content component, including risk analysis matrix, financial data table and resolution clause; priority 3 component is a business logic necessary component, including project timeline, technical parameter table and resource allocation diagram; priority 4 component is an auxiliary explanation component, including data source annotation, term explanation box and reference appendix; priority 5 component is a replaceable component, including decorative header and footer, background watermark and promotional slogan;

[0148] The three-layer style fusion is performed and a standardized template instance file is output, wherein the base layer applies enterprise global style specifications; the component layer mixes component self-brought styles according to weights; and the dynamic adjustment layer detects conflict attributes and assigns final values according to weights.

[0149] The component priority is shown in Table 1.

[0150] Table 1

[0151]

[0152] The step S4 specifically includes:

[0153] The knowledge-enhanced prompt words are constructed, including: retrieving relevant clauses, historical cases and industry specifications from the enterprise private knowledge base, generating semantic vector encoding, converting the variable values filled by the user into a structured feature vector, integrating the text and table data output by the multi-modal analysis, and generating an abstract vector;

[0154] The large language model with domain knowledge enhancement is called to process the prompt words, and project background and file content abstracts are generated;

[0155] The confidence score is added to the AI-generated content, and when the confidence score is lower than the threshold, the manual review flag is triggered.

[0156] The calculation formula of the confidence score is:

[0157] Confidence = 0.5·FactScore + 0.3·LogicScore + 0.2·RiskScore

[0158] Wherein, Confidence represents the confidence score, FactScore represents the fact accuracy score, i.e. the semantic similarity between the generated content and the enterprise knowledge base clauses, LogicScore represents the logic coherence score, which is used to represent the semantic coherence between the front and rear paragraphs, and RiskScore represents the risk compliance score, which matches 2000+ risk word libraries, and 0.2 points are deducted for each occurrence of a violation word.

[0159] The step S5 specifically comprises:

[0160] The dynamically combined template is rendered with variable values and AI-generated content using the Docxtemplater engine to generate an initial DOCX document.

[0161] The document merging component is called to append the user-uploaded files to the end of the initial document according to the preset sorting rules, forming a complete conference review document.

[0162] The step S6 specifically comprises:

[0163] A three-level scanning mechanism is used to scan document elements, detect and mark all format abnormalities deviating from standard values.

[0164] Based on the entity association engine, key data in multiple documents is extracted, a conflict detection matrix is constructed to identify numerical contradictions, term differences, and logical conflicts, and a heat map is used to visualize the distribution of conflict points.

[0165] Intelligent error correction is performed according to error types, where format errors call for template style resetting, content conflicts trigger manual confirmation processes, compliance issues are associated with the knowledge base to automatically replace outdated clauses and add desensitization processing.

[0166] A visual audit view containing a hierarchical problem list, a three-dimensional health score, and executable modification instructions is output, and a version bloodline map is recorded for traceability analysis.

[0167] The three-dimensional health score includes format, content, and compliance scores.

[0168] Three-level scanning is performed on document layout elements:

[0169] 1. Basic style verification:

[0170] Font properties (font type, size, color) are detected segment by segment, and paragraph formats (indentation, line spacing, paragraph spacing) are scanned against the enterprise VI standard library with an accuracy of 0.1 characters.

[0171] Check the position of the header and footer (within ±1mm of the border distance) and the page number format

[0172] 2. Complex element review:

[0173] Table style analysis: border line width (0.5±0.1pt), cell fill color (RGB tolerance <5) Image embedding detection: whether the size ratio meets the template requirements, and whether the resolution is ≥300dpi

[0174] Numbering system verification: whether multi-level title numbering is continuous, and whether bullet points are uniform

[0175] 3. Global specification matching:

[0176] Cover page element positioning: company logo size error ≤ 3%, position offset ≤ 5mm

[0177] Approval area verification: number of signature columns consistent with template requirements, date field format uniform

[0178] Table of contents dynamic verification: correspondence between title page number and actual content pages.

[0179] The hierarchical problem list specifically includes:

[0180] Urgent items (red): format errors > 3 places / content conflicts / high-risk compliance issues

[0181] Warning items (yellow): style deviation / inconsistent terminology / medium risk issues

[0182] Suggestion items (blue): optimization suggestions / non-mandatory modification items.

[0183] The step S7 specifically includes:

[0184] Constructing a document generation graph, recording the source of each data item, marking the content conversion path and timestamp;

[0185] Generating a difference visualization report, distinguishing AI-generated content from manually modified parts by color;

[0186] Storing the final document in cloud storage, generating a unique file ID and access link, and returning a downloadable file package to the user interface.

[0187] As shown in the following figure, it is the main part of the method of the present application: Figure 2

[0188] 1. Document template

[0189] Pre-written and saved as a.docx format file, serving as the basic framework for generating the target file. The template not only contains the main content of each part of the document, but also sets the text and paragraph style in detail, including font, font size, color, paragraph spacing, line spacing, etc. The main content uses marked variables, following the Docxtemplater rules, to facilitate subsequent replacement with actual content, laying the foundation for document format uniformity.

[0190] 2. User input information and uploaded files to be reviewed

[0191] The implementation system is developed based on B / S structure, and users access the operation page through the browser. After selecting the document template, the variable information required by the template needs to be filled in and the files to be reviewed need to be uploaded. The system server receives the template ID, variable content and file information to provide data support for subsequent processing.

[0192] ​3、Data processing and document generation program

[0193] As shown in the implementation system, data processing and document generation are mainly completed on the server side. The server is developed in Typescript language (compatible with Javascript language) and runs in nodejs environment. The document processing uses the Javascript version of the library. Figure 3

[0194] 3.1, Build document template parameters

[0195] According to the variables to be replaced in the template, the received user input variable content and uploaded file information are used to build json format data.

[0196] 3.2, Large model generates content

[0197] The language large model is used to generate the content of the project background and main content blocks.

[0198] 3.3, Render template file

[0199] According to the template ID to load file data, combined with the constructed parameter data and AI generated content, the Docxtemplater library rendering method is called to replace the variables with actual content, and the intermediate file T001.docx is generated.

[0200] 3.4, Paste attachment content

[0201] Iterate through the uploaded files, download them to the local computer, and call the docx-merger library merge method to add the files to the end of T001.docx as attachments, generating the final file F001.docx.

[0202] 4, Large model

[0203] The public cloud service interface provided by the large model manufacturer can also be used. The large model service interface can also be used for local deployment. The large model is mainly used to generate project background and file content summary in this scheme, fully utilizing the AI writing ability, and improving the professionalism and efficiency of document content generation.

[0204] 5, Generated conference review file document

[0205] The generated F001.docx file is uploaded to the file service of the system, and the file ID, file name and other information are returned to the page for user file preview and download operation.

[0206] ​It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited by the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.

[0207] The above is the introduction of the method embodiment, and the system embodiment is used to further illustrate the scheme of the present application.

[0208] Figure 4 The system structure diagram of the embodiment of the present application is shown in Figure 4 A system for automatically generating a document based on a template and a large model is used to implement a method for automatically generating a document based on a template and a large model, and includes:

[0209] A dynamic template management engine is used to store a tree structure template component library and a versioned decision rule set, and to perform conditional trigger type template combination and style conflict arbitration.

[0210] A multi-modal data fusion module is used to realize heterogeneous file parsing and structured data extraction through layout segmentation, table reconstruction, and entity linking technology.

[0211] An AI generation unit with knowledge enhancement is used to fuse an enterprise knowledge base and a risk matrix to generate a draft review opinion with confidence control.

[0212] A cross-document consistency checker is used to implement three-order cross-document intelligent checking of format specifications, content conflicts, and compliance clauses, and to perform minimal intervention correction on detected errors.

[0213] A visual version traceability interface is used to realize document full life cycle management through data bloodline graph recording content source and conversion path.

[0214] An intelligent evaluation and repair module is used to evaluate document health degree and perform repair on checking errors.

[0215] The intelligent evaluation and repair module includes:

[0216] A document health degree evaluation instrument panel is used to display three-dimensional health degree scores in real time and generate a list of improvement suggestions.

[0217] An automatic repair engine is used to perform minimal change repair on checking errors and to keep repair history records.

[0218] An intelligent recommendation module is used to recommend related clauses based on context and actively prompt missing information items.

[0219] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the foregoing method embodiment, which will not be described here.

[0220] In the technical solutions of the present application, the acquisition, storage and application of user personal information involved comply with relevant laws and regulations and do not violate public order and good customs.

[0221] According to the embodiments of the present application, the present application further provides an electronic device, a readable storage medium and a computer program product.

[0222] Figure 5 A schematic block diagram of an electronic device 500 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0223] The electronic device 500 includes a computing unit 501 that can perform various appropriate actions and processes according to a computer program stored in a ROM 502 or a computer program loaded from the storage unit 508 into a RAM 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An I / O interface 505 is also connected to the bus 504.

[0224] Various components in the electronic device 500 are connected to the I / O interface 505, including an input unit 506, such as a keyboard, a mouse, etc., an output unit 507, such as various types of displays, speakers, etc., a storage unit 508, such as a magnetic disk, an optical disk, etc., and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.

[0225] The computing unit 501 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit 501 performs various methods and processes described above, such as a method for automatically generating a document based on a template and a large model. For example, in some embodiments, a method for automatically generating a document based on a template and a large model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded onto the RAM 503 and executed by the computing unit 501, one or more steps of a method for automatically generating a document based on a template and a large model described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform a method for automatically generating a document based on a template and a large model by any other appropriate means, such as by means of firmware.

[0226] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0227] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0228] In the context of this application, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0229] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0230] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0231] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server can arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0232] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps described in this application can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in this application can be achieved, which is not limited herein.

[0233] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for automatically generating documents based on templates and large models, characterized in that, Includes the following steps: Step S1: The user enters information and uploads the document to be reviewed through the browser interface; Step S2: Parse the uploaded multimodal file and extract structured data; Step S3: Output standardized template instance files based on the dynamic combination template of the decision tree; Step S4: Use the large model to generate project background and document content summary; Step S5: Render the generated content and merge it with attachments to form a complete meeting review document; Step S6: Perform cross-document consistency intelligent verification and intelligent error correction; Step S7: Generate and store the final document with version traceability information.

2. The method according to claim 1, characterized in that, Step S1 specifically includes: The browser interface receives the dynamic template identifier selected by the user and the associated decision rule set, wherein the decision rule set is defined using Boolean logic expressions; It receives template variable values ​​filled in by users and uploaded multimodal documents for review. Supported file types include PDF, DOCX, PPT and image scans.

3. The method according to claim 2, characterized in that, Step S2 specifically includes: Establish a file type routing mechanism to distribute files to the corresponding parsing engines based on file extensions and content characteristics; The parsing engine performs hierarchical parsing, including: applying a layout segmentation model to PDF files, performing OOXML deep parsing on DOCX files, and using end-to-end OCR table reconstruction for scanned images; Extract structured data from the parsing results, including text paragraphs, table data, and image titles, and construct a unified JSON intermediate representation; The OCR table reconstruction specifically includes the following steps: Generate candidate structures from the input original table image; Calculate cell alignment and empty cell rate separately; Construct a table reconstruction optimization function, and calculate the formula as follows: Γ(T)=argmin[0.7·D cell (T,S)+0.3·Empty rate (S)] Where Γ(T) represents the table reconstruction optimization function, T represents the original table region, S represents the candidate table structure, and D... cell Empty indicates cell alignment. rate Indicates the percentage of empty cells; The candidate table structure output that minimizes the table reconstruction optimization function is selected as the optimal table structure.

4. The method according to claim 1, characterized in that, Step S3 specifically includes: Candidate components are recursively retrieved from a pre-built tree-structured template library, wherein the candidate components include: core components, including: title page, approval process table and risk analysis matrix; conditional components, including: high-amount approval module and emergency plan section; nested sub-templates, including: legal clause appendix and expert review opinion form; Perform conditional function matching for each candidate component, specifically including: verifying whether the user input variables meet the component activation conditions, checking the enterprise approval system requirements, confirming whether the preceding components have been activated, and outputting the component matching score. Based on the matching score, a tiered decision is made: a component matching score ≥ 0.8 is strongly activated, the component is immediately inserted and its position is locked; a component matching score ≤ 0.6 < 0.8 is weakly activated, a placeholder is added and manual confirmation is prompted; a component matching score < 0.6 is not activated and the component is skipped; a component activation log is generated to record the decision basis and variable tracing. Style weights are assigned based on component priority, which is divided into 5 priority levels. Priority 1 components are core components for legal validity, including the contract signing page, approval signature area, and official seal area; Priority 2 components are components for key decision-making content, including risk analysis matrix, financial data table, and resolution clauses; Priority 3 components are components necessary for business logic, including project timeline, technical parameter table, and resource allocation diagram; Priority 4 components are auxiliary explanatory components, including data source notes, terminology explanation box, and reference appendix; Priority 5 components are replaceable components, including decorative headers and footers, background watermarks, and promotional slogans. Perform three-layer style fusion and output a standardized template instance file. The base layer applies the enterprise's global style specification; the component layer mixes the component's built-in styles according to weights; and the dynamic adjustment layer detects conflicting attributes and assigns final values ​​according to weights.

5. The method according to claim 4, characterized in that, Step S4 specifically includes: Construct knowledge-enhanced prompts, including: retrieving relevant clauses, historical cases, and industry standards from the enterprise's private knowledge base, generating semantic vector codes, converting user-entered variable values ​​into structured feature vectors, integrating text and tabular data output from multimodal parsing, and generating summary vectors; The domain knowledge-enhanced large language model is invoked to process prompt words and generate project background and document content summaries. Add a confidence score to AI-generated content, and trigger a manual review flag when the confidence score is below a threshold.

6. The method according to claim 5, characterized in that, Step S5 specifically includes: The Docxtemplater engine is used to render dynamically combined templates and variable values, as well as the content generated from the large model, to produce the initial DOCX document. The document merging component is invoked to append user-uploaded files to the end of the initial document according to preset sorting rules, forming a complete meeting review document.

7. The method according to claim 6, characterized in that, Step S6 specifically includes: A three-level scanning mechanism is used to scan document elements, detect and mark all formatting anomalies that deviate from standard values; Based on the entity association engine, key data in multiple documents are extracted, a conflict detection matrix is ​​constructed to identify numerical contradictions, terminological differences and logical conflicts, and the distribution of contradiction points is visualized through heatmaps. Intelligent error correction is performed based on the error type. For example, formatting errors trigger template style reset, content conflicts trigger manual confirmation, and compliance issues are automatically replaced with expired clauses and desensitized processing by associating them with the knowledge base. The output includes a tiered list of issues, a three-dimensional health score, and a visual audit view with executable modification instructions, while simultaneously recording the version lineage chart for traceability analysis.

8. The method according to claim 7, characterized in that, Step S7 specifically includes: Build a document generation graph, record the source of each data item, and mark the content transformation path and timestamp; Generate a difference visualization report, using color to distinguish between AI-generated content and manually modified parts; Store the final document to cloud storage, generate a unique file ID and access link, and return a downloadable file package to the user interface.

9. A system for automatically generating documents based on templates and large models, used to implement the method for automatically generating documents based on templates and large models as described in any one of claims 1-8, characterized in that, include: A dynamic template management engine is used to store a tree-structured template component library and a set of versioned decision rules, and to arbitrate condition-triggered template combinations and style conflicts. The multimodal data fusion module is used to parse heterogeneous files and extract structured data through layout segmentation, table reconstruction, and entity linking technologies. AI generation unit with knowledge enhancement is used to integrate enterprise knowledge base and risk matrix to generate draft review opinions with confidence control; A cross-document consistency checker is used to implement three-level intelligent cross-document checks for format specifications, content conflicts, and compliance clauses, and to perform minimal intervention corrections for detected errors. A visual version traceability interface is used to record the source and transformation path of content through a data lineage graph, thereby achieving full lifecycle management of documents. The intelligent assessment and repair module is used to assess the health of documents and repair verification errors.

10. The system according to claim 9, characterized in that, The intelligent assessment and repair module includes: The document health assessment dashboard displays a three-dimensional health score in real time and generates a list of improvement suggestions. An automatic repair engine is used to perform minimal change repairs on checksum errors and retain a repair history. The intelligent recommendation module is used to recommend relevant terms based on context and proactively suggest missing information items.

Citation Information

Patent Citations

  • Methods, apparatus, equipment and storage media for automatic generation of meeting minutes

    CN110493019B

  • Automatic conference recording and abstract generating method for intelligent conference

    CN120340497A

Cited By

  • Automatic secret evaluation compliance document generation system and method based on strategic data management and control and full-link verification

    CN121212111A

  • Large model auxiliary decision-making method and system applied to natural resource informatization management

    CN121615755A

  • Large model information extraction and structure restoration system for long text document

    CN121638215A

  • A system for extracting large model information and restoring structure from long text documents

    CN121638215B

  • Intelligent configuration method, assembly and application of standard file

    CN121683746A