File generation method and device based on large model and RAG
By using large-scale models and RAG technology to perform deep learning and semantic analysis on tender documents, a tender document outline is generated and populated with materials. Combined with NLP review and format adjustment, this solves several technical challenges in tender document generation, achieving efficient, professional, and secure tender document generation and lowering the application threshold for small and medium-sized enterprises.
Patent Information
- Application Number
- CN202511321664.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-12-30
AI Technical Summary
Existing large-scale models and RAG technology have problems in bid document generation, such as insufficient coverage of long-tail information, search blind spots, content illusion, poor logical coherence, lack of professionalism, high compliance and security risks, poor human-machine collaboration, and high deployment costs, resulting in high application thresholds and difficulty in widespread implementation.
By using a large-scale LLM model to perform deep learning and semantic analysis on tender documents, core project information is obtained. Combined with industry knowledge graphs and template databases, a tender document outline is generated, materials are filled in and a draft is generated. NLP technology is used to review and revise the content, realize format adjustment and manual review, and finally output the final draft. The knowledge base is generated by integrating multiple source documents, supporting collaborative editing and real-time quality control.
It improved the professionalism and logical coherence of bidding documents, reduced compliance risks, enhanced human-machine collaboration efficiency, reduced optimization costs, lowered the application threshold for SMEs, and ensured the real-time effectiveness and comprehensiveness of documents.
Smart Images

Figure CN121234892A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent file parsing and generation technology, specifically to a file generation method and apparatus based on a large model and RAG. Background Technology
[0002] Based on large-scale modeling and intelligent agent technology, the intelligent generation application has acquired core capabilities related to document parsing, content generation, and compliance checks, and has been initially applied in scenarios such as bid document preparation. Specifically, the intelligent retrieval stage utilizes large-scale models and RAG technology to parse various historical data such as personnel, resumes, and performance records, constructing a vector knowledge base. The bid preparation intelligent agent can accurately locate knowledge fragments, significantly improving retrieval quality and efficiency. The intelligent generation stage relies on large-scale models to analyze customer needs, match reference knowledge, and automatically generate bid documents. Simultaneously, an online rich text editor assists manual drafting of proposals, shortening the proposal preparation cycle. The document optimization and proofreading stage uses large-scale models to polish, expand, or abbreviate the proposal content using professional language, helping to improve the professionalism and completeness of the proposal content and providing technical support for bidding work.
[0003] While large-scale modeling and intelligent agent technologies have shown potential in the aforementioned scenarios, several shortcomings remain. First, knowledge base construction relies heavily on the quality and timeliness of historical data; RAG technology's insufficient coverage of long-tail information can lead to retrieval blind spots, affecting the comprehensiveness of knowledge acquisition. Second, generated content carries the risk of "illusion," potentially producing false information and being sensitive to prompts, making it difficult to guarantee content professionalism and stylistic consistency. Third, long text generation suffers from poor logical coherence, prone to content repetition or contradictions, and has limited capabilities in complex reasoning and cross-domain knowledge integration. Fourth, personalization capabilities are weak, relying heavily on general models with low customization levels and lacking a continuous learning mechanism based on user feedback. Fifth, compliance and security risks are prominent, with potential for sensitive information leakage and a lack of authoritative review mechanisms, failing to fully replace human review. Sixth, the human-machine collaboration experience is poor; the integration of AI and the editor is low, and AI suggestions lack interpretability, impacting user trust and work efficiency. Seventh, model deployment costs are high, resource consumption is significant, and cross-domain generalization requires extensive optimization, raising the application threshold for small and medium-sized enterprises and hindering widespread technology adoption. Summary of the Invention
[0004] To address these issues, this invention provides a document generation method and apparatus based on a large model and RAG (Research Aggregator), resolving problems such as insufficient coverage of long-tail information in RAG and search blind spots. Through semantic analysis of the large model, it solves problems such as content "illusion" and difficulty in ensuring professionalism and logical coherence; through real-time quality control, compliance review, and collaborative editing, it addresses issues such as high compliance and security risks and poor human-machine collaboration. Simultaneously, it reduces optimization costs and addresses the problems of high model deployment costs and high application barriers for small and medium-sized enterprises.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a file generation method based on large models and RAG, comprising:
[0006] The tender document uploaded by the user is subjected to deep learning and semantic analysis using a large model LLM to obtain the core project information in the tender document; the business field and industry category to which the tender document belongs are determined based on the core project information.
[0007] Based on the core project information, business areas, and industry categories in the tender document, and referring to the industry knowledge graph and the established tender document template database, a tender document template matching the user's needs is selected; based on the tender document template and the core project information, a tender document outline is generated.
[0008] Based on the outline of the tender document, the required data is extracted from the set material library through the large model LLM and populated to generate the first draft of the tender document;
[0009] The initial draft of the bid is comprehensively reviewed using NLP technology to identify factors that could lead to rejection and to check for inconsistencies in the expression of key terms, generating problem suggestions and correction recommendations. Based on these suggestions, the editors revise the initial draft of the bid to generate a revised version.
[0010] The revised bid document will undergo format adjustments, layout optimization, and manual review to produce the final bid document.
[0011] As a preferred solution for document generation based on large models and RAG, it also includes: classifying and integrating heterogeneous documents from multiple sources such as Word, PDF, and images within the enterprise to generate a knowledge base, ensuring the documents are effective in real time.
[0012] As a preferred approach for generating documents based on large models and RAG, the initial draft of the tender document includes: a solution overview, qualification certificates, company profile, and past performance.
[0013] As a preferred solution for document generation based on large models and RAG, if several editors are collaborating during the process of generating the initial draft of the tender document, the "collaborative tender editing" mode is activated, allowing several editors to simultaneously edit the content of their respective chapters online.
[0014] As a preferred solution for document generation based on large models and RAG, during the comprehensive quality review of the initial draft of the tender document using the NLP technology, spelling errors, non-standard formats, and layout problems existing in the editing process of the initial draft of the tender document are captured and corrected in real time.
[0015] This invention also provides a file generation apparatus based on large models and RAG, and based on the above-mentioned file generation method based on large models and RAG, it includes:
[0016] The tender document parsing module is used to perform deep learning and semantic analysis on the tender document files uploaded by users through a large model LLM to obtain the core project information in the tender document files; and to determine the business field and industry category to which the tender document files belong based on the core project information.
[0017] The template filtering and catalog outline generation module is used to filter and obtain bid templates that match user needs based on the core project information, business areas, and industry categories in the tender document, with reference to the industry knowledge graph and the set bid template database; and to generate a bid outline based on the bid templates and the core project information.
[0018] The tender draft generation module is used to generate a tender draft by extracting the required data from the set material library and filling the data based on the tender outline and the large model LLM.
[0019] The quality review and correction module is used to review the quality of the initial draft of the bid using NLP technology, identify factors that would lead to rejection and the consistency of key terminology, and generate problem prompts and correction suggestions; based on the problem prompts and correction suggestions, the initial draft of the bid is revised to generate a revised version of the bid;
[0020] The final draft output module is used to adjust the format, optimize the layout, and manually review the revised bid document to output the final bid document.
[0021] As a preferred solution for a document generation device based on a large model and RAG, it also includes: a knowledge base management module, which is used to classify and integrate heterogeneous documents from multiple sources such as Word, PDF and images within the enterprise to generate a knowledge base and ensure that the documents are valid in real time.
[0022] As a preferred embodiment of a document generation device based on a large model and RAG, the initial draft of the tender document in the tender document generation module includes: a solution overview, qualification certificates, company profile, and past performance.
[0023] As a preferred solution for a document generation device based on a large model and RAG, in the tender draft generation module, if several editors are collaborating on the process of generating the tender draft, a "collaborative tender editing" mode is activated, allowing several editors to simultaneously edit the content of their respective chapters online.
[0024] As a preferred solution for a document generation device based on a large model and RAG, the quality review and correction module, during the process of conducting a comprehensive quality review of the initial draft of the tender document using the NLP technology, captures and corrects spelling errors, non-standard formats, and layout problems that exist in the editing process of the initial draft of the tender document in real time.
[0025] This invention has the following advantages: It uses a large-scale LLM model to perform deep learning and semantic analysis on user-uploaded tender documents to obtain core project information. Based on this core project information, it determines the business area and industry category of the tender document. Based on the core project information, business area, and industry category, and referring to an industry knowledge graph and a pre-defined tender template database, it selects a tender template that matches the user's needs. Based on the tender template and the core project information, it generates a tender outline. Based on the tender outline, it uses the large-scale LLM model to extract necessary data from a pre-defined material library for data filling, generating a draft tender. It uses NLP technology to conduct a comprehensive quality review of the draft tender, identifying rejection factors and consistency in key terminology, generating problem prompts and correction suggestions. Editors revise the draft tender based on the problem prompts and correction suggestions, generating a revised tender. The revised tender undergoes format adjustments, layout optimization, and manual review, resulting in the final tender. This invention also categorizes and integrates heterogeneous documents from multiple sources, including Word, PDF, and images, within an enterprise to generate a knowledge base, ensuring the documents remain valid in real time. By leveraging large-scale model technology and the RAG architecture, this invention constructs an automated workflow encompassing parsing, content generation, inspection, and knowledge base management modules. The knowledge base management module integrates and updates multi-source documents in real time, addressing the issues of insufficient long-tail information coverage and retrieval blind spots in RAG. Through large-scale model semantic analysis and multi-module collaboration, it solves the problems of content "illusion" and difficulty in ensuring professionalism and logical coherence. Real-time quality control, compliance review, collaborative editing, and version management address the issues of high compliance and security risks and poor human-machine collaboration, while reducing optimization costs and resolving the problems of high model deployment costs and high application barriers for SMEs. Attached Figure Description
[0026] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.
[0027] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.
[0028] Figure 1 This is a flowchart illustrating a file generation method based on a large model and RAG provided in Embodiment 1 of the present invention;
[0029] Figure 2 This is a schematic diagram illustrating the specific implementation process of a file generation method based on a large model and RAG provided in Embodiment 1 of the present invention;
[0030] Figure 3 This is a schematic diagram of the implementation process of a possible embodiment provided in Embodiment 1 of the present invention;
[0031] Figure 4 This is a schematic diagram of the architecture of a file generation device based on a large model and RAG provided in Embodiment 2 of the present invention. Detailed Implementation
[0032] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] Example 1
[0034] See Figure 1 and Figure 2 Embodiment 1 of the present invention provides a file generation method based on large models and RAG, including the following steps:
[0035] S1. Use a large-scale model (LLM) to perform deep learning and semantic analysis on the user-uploaded tender document file to obtain the core project information in the tender document file; determine the business area and industry category to which the tender document file belongs based on the core project information;
[0036] S2. Based on the core project information, business area, and industry category in the tender document, and referring to the industry knowledge graph and the established tender document template database, a tender document template matching the user's needs is selected; based on the tender document template and the core project information, a tender document outline is generated.
[0037] S3. Based on the outline of the tender document, extract the required data from the set material library through the large model LLM to fill in the data and generate the first draft of the tender document.
[0038] S4. Use NLP technology to conduct a quality review of the initial draft of the bid, identify factors that could lead to rejection and check for inconsistencies in the expression of key terms, and generate problem prompts and correction suggestions; based on the problem prompts and correction suggestions, revise the initial draft of the bid and generate a revised version of the bid;
[0039] S5. Adjust the format, optimize the layout, and manually review the revised bid document to output the final bid document.
[0040] In this embodiment, it also includes: S6, classifying and integrating heterogeneous documents from multiple sources such as Word, PDF, and images within the enterprise to generate a knowledge base, ensuring that the documents are valid in real time.
[0041] In this embodiment, in step S1, deep learning and semantic analysis are performed on the tender document uploaded by the user using a large model LLM to obtain the core project information in the tender document; the business field and industry category to which the tender document belongs are determined based on the core project information.
[0042] Specifically, the large-scale LLM model is used to perform deep learning and semantic analysis on user-uploaded tender documents. Leveraging the powerful natural language processing capabilities of LLM, the document text is parsed segment by segment to accurately identify and extract core project information, including but not limited to: project name, scope of bidding, technical requirements, construction period requirements, budget amount, bid deadline, and other key content. After obtaining the core project information, based on the deep learning results of LLM on the business characteristics and domain terminology of various industries, combined with the business scenario descriptions and industry-specific expressions reflected in the tender documents, the business domain to which the tender documents belong is automatically determined, such as construction engineering, information technology services, equipment procurement, etc., and the industry category, such as municipal industry, medical industry, education industry, etc., providing accurate information basis for subsequent template matching and content generation.
[0043] In this embodiment, in step S2, based on the core project information, business area, and industry category of the tender document, and referring to the industry knowledge graph and the set tender document template database, a tender document template matching the user's needs is selected; based on the tender document template and the core project information, a tender document outline is generated.
[0044] Specifically, based on the core project information of the tender document obtained in step S1, and the determined business area and industry category, the system calls upon the built-in industry knowledge graph and the set tender document template database. The industry knowledge graph covers business processes, common requirements, and standard specifications in various fields; the set tender document template database stores a massive number of standardized tender document templates adapted to different business areas and industry categories; through a large-scale model (LLM), the key information of the tender document is semantically matched and similarity calculated with the template features in the knowledge graph and template database to filter out tender document templates that highly match the user's bidding needs; after determining the matching template, the LLM combines the template's structural framework with the core project information extracted in step S1 to automatically plan the chapter settings of the tender document, clarify the core content direction of each chapter, and generate a logically clear and content-appropriate tender document outline. The outline includes chapter titles, a summary of the core content of each chapter, and the logical connections between chapters.
[0045] In this embodiment, in step S3, based on the outline of the tender document, the required data is extracted from the set material library through the large model LLM to fill the data and generate the first draft of the tender document;
[0046] Specifically, based on the outline of the tender document generated in step S2, the large-scale model (LLM) precisely extracts the required materials from a pre-defined material library according to the content requirements of each chapter of the outline. This material library includes various types of materials such as company qualification documents, past project performance cases, technical solution templates, and standard clause libraries, and these materials are categorized and labeled according to business areas and content types. The LLM extracts material content highly relevant to the themes of each chapter through semantic retrieval and content matching. For example, it extracts information such as the company's business license and relevant qualification certificates from the "Company Qualifications" chapter, and key pages of contracts and acceptance report summaries from similar projects from the "Past Performance" chapter. Subsequently, the LLM integrates the extracted materials with the core project information of the tender document obtained in step S1, automatically filling them into the corresponding chapters of the outline according to the format requirements and language style of the tender document template, generating a structurally complete and initially adapted draft tender document.
[0047] In this embodiment, in step S4, NLP technology is used to conduct a comprehensive quality review of the initial draft of the bid, check for factors that would lead to rejection and the consistency of key terminology, and generate problem prompts and correction suggestions; the editors revise the initial draft of the bid based on the problem prompts and correction suggestions to generate a revised version of the bid;
[0048] Specifically, a comprehensive quality audit is conducted on the initial draft of the tender document generated in step S3 using NLP technology (including core capabilities such as rule engines and pattern recognition).
[0049] The rules engine checks each disqualifying factor in the initial draft against a pre-set list of bid compliance standards and risk points for bid rejection (such as whether the legal representative's signature page is missing, whether the bid validity period meets the requirements, and whether the technical parameters respond to the bidding requirements). Pattern recognition technology performs a full document scan of key terms in the initial draft (such as project name, technical parameters, company abbreviation, etc.) to identify inconsistencies and contradictions, such as using "≥100Mbps" before and ">100Mbps" after the same technical parameter, or using "XX Technology Company" before and "XX Technology Company" after the company name. After the review is completed, the system automatically generates problem prompts and correction suggestions, such as: "The bid security payment certificate clause is missing here. It is recommended to add it to the 'Commercial Terms' section" and "The 'Project Duration' is inconsistent. It is recommended to unify it to '180 calendar days'." Based on the problem prompts and correction suggestions output by the system and their own professional judgment, the editors make targeted modifications, additions, and improvements to the problem content in the initial draft of the bid, and finally generate a revised version of the bid.
[0050] In this embodiment, in step S5, the revised bid document is formatted, optimized, and manually reviewed to output the final bid document.
[0051] Specifically, after obtaining the revised bid document, the document is first formatted and optimized. Based on the format requirements specified in the tender document, such as font type, font size, line spacing, margin parameters, and page number format, the overall document format is automatically standardized. Chapter headings are styled correctly, for example, first-level headings are bolded in Song typeface, size 2, and second-level headings are bolded in Song typeface, size 3. The layout of tables and images is adjusted to ensure a clean and professional appearance. After format optimization, professional personnel conduct a manual review of the revised bid document, focusing on checking for omissions, format deviations, and hidden issues not discovered during the initial machine review and correction process, such as the accuracy of technical descriptions and logical coherence. After manual review confirms there are no issues, the final bid document is output for the user to submit their bid.
[0052] In this embodiment, in step S6, heterogeneous documents from multiple sources, including Word, PDF, and images, within the enterprise are categorized and integrated to generate a knowledge base, ensuring that the documents are valid in real time.
[0053] Specifically, it categorizes and integrates heterogeneous documents from multiple sources, such as Word, PDF, and images, within the enterprise to generate a knowledge base, ensuring the documents are effective in real time, and providing accurate data retrieval support for the parsing, generation, and review stages of steps S1-S5.
[0054] In this embodiment, the system has a built-in version control system throughout the entire tender document generation process, allowing users to flexibly switch and manage between different versions of the document, ensuring that the content of each tender document strictly follows the prompt word specifications and industry standards, and realizing full lifecycle management of the document.
[0055] In one possible embodiment, a specific example of an environmental technology company participating in a "city wastewater treatment plant upgrade and renovation project" bidding is provided below:
[0056] The specific implementation steps are as follows: Figure 3 As shown:
[0057] T1. Verification of tender documents:
[0058] After performing deep learning and semantic analysis on the user-uploaded tender document file using a large-scale LLM model, the tender document parsing results are output:
[0059] Key information: The project has a treatment capacity of 50,000 tons / day, and the upgraded emission standard must meet the Class A standard of the "Urban Wastewater Treatment Plant Pollutant Discharge Standard". The construction period is 12 months, the budget is 65 million yuan, the business area is "environmental engineering", and the industry category is "municipal environmental protection industry".
[0060] Business personnel initiated a manual comparison: comparing the parsed results with the original tender document page by page, and found that the "Equipment Warranty Period Requirement" was parsed as "2 years," while the original document clearly stated "3 years," indicating a parsing error. The business personnel selected "Manual Correction," directly adjusting the field to "3 years" in the system and marking the correction record ("2024-09-12, Zhang San corrected: Equipment warranty period changed from 2 years to 3 years"), ensuring the accuracy of subsequently generated content.
[0061] T2. Table of Contents Generation:
[0062] Template Selection: Based on the revised core information, the system matches templates from the resource library. By comparing tags such as "sewage treatment plant renovation" and "environmental engineering bidding," three candidate templates are selected: "Municipal Sewage Treatment Project Bidding Document (Standard Version)," "Environmental Engineering EPC Project Bidding Document," and "Sewage Treatment Equipment Procurement and Installation Bidding Document." Business personnel review the template outline preview and find that the "Environmental Engineering EPC Project Bidding Document" includes suitable sections such as design scheme, equipment procurement, and construction organization, and select this template.
[0063] Automatic outline generation: The large model LLM combines the template framework with core information to generate a table of contents outline: The first-level table of contents includes project overview, technical solution, business response, and qualifications and performance; under "technical solution", a second-level table of contents is added: "Wastewater treatment process upgrade solution" and "Equipment selection and parameters"; in "business response", a sub-chapter on "warranty period commitment" is clearly defined, forming a structured outline that fits the project requirements.
[0064] T3. Material Integration and Filling:
[0065] Module Information Acquisition: The system searches the material library chapter by chapter according to the table of contents: the "Qualifications and Performance" chapter automatically matches scanned copies of qualification documents such as "Level II General Contracting for Municipal Public Works Construction" and "Level A Special Design for Environmental Engineering"; the "Technical Solutions" chapter extracts process flow diagrams, sludge treatment technical parameters and other information from similar historical projects (such as "Renovation of a Wastewater Treatment Plant in a Development Zone of a Certain City"); the "Equipment Selection" section calls the technical parameter tables of "MBR Membrane Modules" and "Aeration Systems" in the equipment library.
[0066] Qualification Material Retrieval: For the requirement of "Project Manager Qualification", the system queries the company's qualification database to obtain the information of personnel holding the "First-Class Construction Engineer (Municipal Public Works)" certificate (including name, certificate number, and social security payment certificate) and automatically fills it into the corresponding position.
[0067] Content generation and matching: LLM integrates the retrieved materials with the bidding requirements to generate the "Equipment Warranty Period Commitment" content: "We promise that the warranty period for all equipment in this project is 3 years (starting from the date of acceptance), which exceeds the industry standard by 1 year and is in line with and better than the requirements of the bidding documents," ensuring that the content is consistent with the revised information.
[0068] T4. Directory content traversal and generation:
[0069] The system generates content hierarchically according to the outline: first, it completes basic chapters such as "Project Overview" and "Business Response"; for complex chapters such as "Wastewater Treatment Process Upgrade Plan," it uses a "semi-automatic generation" method.
[0070] LLM first outputs a draft process based on membrane bioreactor (MBR), and technicians add "innovative points for sludge reduction treatment," with the system updating the content in real time. During the generation process, for the "Construction Schedule" section, the "12-month construction period segmentation plan" is retrieved from the "Project Schedule Template Library" and automatically replaced with key nodes for this project, such as "Complete equipment installation in the 3rd month."
[0071] T5. Manual review and adjustment:
[0072] The automatically generated content is pushed to the shared editing interface of the business and technical teams in real time: the technical lead discovers an error in the unit of "Equipment Operation Energy Consumption Estimation" ("kWh / day" is mistakenly written as "kWh / month"), and corrects it directly online; the business staff checks the "Bid Quotation Summary Table" and adds the "Spare Parts and Components Separate Costs" item, and the system records all modification traces. After the modification is completed, the team signs it through the system's "Collaboration Confirmation" function, triggering the next step of quality audit.
[0073] T6. Final Refinement and Output:
[0074] Iterative process: NLP technology screened out that "some equipment parameters did not have the testing basis marked". After the system pushed a correction prompt, the technicians added "based on the 'Technical Specification for Testing of Urban Wastewater Treatment Plants' (HJ 582-2010)", and the second review was approved.
[0075] Formatting and Beautification: The system adopts a unified format according to the bidding requirements: the cover adds the project name and company logo, the body text uses "SimSun font, size 5, 1.5 line spacing", the chart number format is "Figure XY" (X is the chapter number, Y is the chart number), and a table of contents with jump function is automatically generated.
[0076] Final Review: Joint review by business, technical, and legal personnel: The legal team confirms that the "Contract Terms Response" has no legal risks; the business team verifies that the "Payment Method" is consistent with the bidding requirements; and the technical team confirms that the process description has no logical flaws. After all personnel sign off, the system outputs a final PDF draft (with electronic signature), and simultaneously generates an editable version for archiving, completing the tender document preparation.
[0077] The application scenarios of this invention are as follows:
[0078] This invention is applicable to bidding processes in various industries, including construction engineering, environmental engineering, medical equipment procurement, and information technology services, especially for scenarios requiring rapid response to complex bidding requirements. For example, when a construction company participates in bidding for a municipal road renovation project, this invention can automatically parse core information such as the construction period, budget, and technical standards in the bidding documents, match suitable bid templates, generate a table of contents, extract qualification certificates and past performance data from the company's resource library to fill in the content, and then use NLP technology to review and screen for bid rejection risks. Finally, it quickly outputs compliant and professional bid documents, solving the problems of low efficiency, error-proneness, and difficulty in data integration in traditional bid preparation.
[0079] For procurement projects of government departments, universities, and public hospitals, such as the purchase of laboratory equipment for universities and the upgrading of information systems in public hospitals, this invention can help suppliers efficiently prepare procurement response documents. The system can accurately analyze the compliance requirements in the procurement documents, generate response documents that conform to the official format by referring to the knowledge graph of the procurement field, and automatically check key items such as "whether the procurement budget limit is met" and "whether proof of small and micro enterprises is provided," ensuring that the response documents fully comply with the procurement specifications and increasing the supplier's chances of being selected.
[0080] It can be extended to the compilation of various professional documents within an enterprise, such as project feasibility study reports, technical solution documents, and customer service plans. For example, when a technology company customizes a software development project for a client, based on the client's requirements document, the system can analyze the core requirements through a large model, generate a draft solution by referring to the company's technical solution template library, extract similar project architecture diagrams, technical parameters, and other materials from the technical material library to supplement the content, and then use NLP technology to review the logical coherence and terminology consistency of the solution, helping enterprises quickly output high-quality client solutions and reduce the workload of repetitive writing by technical personnel.
[0081] Addressing the pain points of SMEs—such as a lack of professional documentation teams and high model deployment costs—this invention lowers the application threshold for SMEs through a lightweight RAG architecture and modular design. For example, when small equipment sales companies participate in local government equipment procurement bids, they do not need to build complex knowledge bases themselves. They can rely on the system to integrate the company's existing qualifications, performance, and other basic information, and automatically generate bid documents and complete compliance verification using a large model. This achieves intelligent documentation at a lower cost, enhancing their competitiveness in bidding against large enterprises.
[0082] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.
[0083] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0084] Example 2
[0085] See Figure 4 Embodiment 2 of the present invention also provides a file generation device based on large models and RAG, comprising:
[0086] The tender document parsing module 001 is used to perform deep learning and semantic analysis on the tender document files uploaded by users through a large model LLM to obtain the core project information in the tender document files; and to determine the business field and industry category to which the tender document files belong based on the core project information.
[0087] The template filtering and catalog outline generation module 002 is used to filter and obtain bid templates that match user needs based on the core project information, business area, and industry category of the tender document, with reference to the industry knowledge graph and the set bid template database; and to generate a bid outline based on the bid template and the core project information.
[0088] The tender draft generation module 003 is used to generate a tender draft by extracting the required data from the set material library and filling the data based on the tender outline and the large model LLM.
[0089] The quality review and correction module 004 is used to review the quality of the initial draft of the bid using NLP technology, check for factors that would lead to rejection and consistency in the expression of key terms, and generate problem prompts and correction suggestions; based on the problem prompts and correction suggestions, the initial draft of the bid is revised to generate a revised version of the bid;
[0090] The final draft output module 005 is used to adjust the format, optimize the layout, and manually review the revised bid document to output the final draft of the bid document.
[0091] This embodiment also includes: a knowledge base management module 006, which is used to classify and integrate multi-source heterogeneous documents such as Word, PDF and images within the enterprise to generate a knowledge base and ensure that the documents are valid in real time.
[0092] In this embodiment, the initial draft of the tender document in the tender document generation module 003 includes: a solution overview, qualification certificates, company profile, and past performance.
[0093] In this embodiment, in the initial draft generation module 003, if several editors collaborate on editing during the initial draft generation process, the "collaborative editing" mode is activated, allowing several editors to simultaneously edit the content of their respective chapters online.
[0094] In this embodiment, the quality review and correction module 004, during the process of conducting a comprehensive quality review of the initial draft of the tender document using the NLP technology, captures and corrects spelling errors, non-standard formats, and layout problems that exist in the editing process of the initial draft of the tender document in real time.
[0095] It should be noted that the information interaction and execution process between the modules of the above system are based on the same concept as the method embodiment in Embodiment 1 of this application, and the resulting technical effects are the same as those in the method embodiment of this application. For details, please refer to the description in the method embodiment shown above in this application, and it will not be repeated here.
[0096] Example 3
[0097] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium storing program code for a file generation method based on a large model and RAG. The program code includes instructions for executing the file generation method based on a large model and RAG according to Embodiment 1 or any possible implementation thereof.
[0098] Computer-readable storage media can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives, SSDs).
[0099] Example 4
[0100] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;
[0101] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor can execute a file generation method based on a large model and RAG according to Embodiment 1 or any possible implementation thereof by calling the program instructions.
[0102] Specifically, a processor can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. This memory can be integrated into the processor or located outside the processor and exist independently.
[0103] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable system. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0104] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing systems. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Optionally, they can be implemented using program code executable by a computing system, thereby storing them in a storage system for execution by the computing system. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0105] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A large model and RAG-based file generation method, characterized by, Comprise: Through the large model LLM, the core project information in the tender document uploaded by the user is obtained by deep learning and semantic analysis; According to the core project information, the business field and industry category to which the tender document belongs are determined; Based on the core project information, the business field and the industry category of the tender document, the tender template matching the user's demand is obtained by screening the industry knowledge graph and the bid document template database; Based on the bid document template and the core project information, the bid document directory outline is generated; Based on the bid document directory outline, the required materials are extracted from the set material library by the large model LLM to fill in the materials, and the bid document draft is generated; Through NLP technology, the bid document draft is audited for quality, and the consistency of the key term expression and the factors for bid rejection are checked to generate problem prompts and correction opinions; According to the problem prompts and the correction opinions, the bid document draft is revised to generate a revised version of the bid document; The revised version of the bid document is adjusted in format, optimized in layout and manually reviewed to output the final version of the bid document.
2. The method of claim 1, wherein, Also include: Classify and integrate multi-source heterogeneous documents such as word, pdf and pictures in the enterprise to generate a knowledge base.
3. The method of claim 2, wherein, The bid document draft includes: solution overview, qualification certificate, company profile and past performance.
4. The method of claim 3, wherein, In the process of generating the bid document draft, if several editors collaborate to edit, the "collaborative bidding" mode is started, allowing several editors to edit the chapter content they are responsible for online at the same time.
5. The method of claim 4, wherein, In the process of comprehensive quality audit of the bid document draft by the NLP technology, spelling errors, non-standard formats and layout problems existing in the editing process of the bid document draft are captured and corrected in real time.
6. A large model and RAG based file generation device, adopting the large model and RAG based file generation method of any one of claims 1-5, characterized in that, Comprise: A tender document analysis module for deep learning and semantic analysis of the tender document uploaded by the user through the large model LLM to obtain the core project information in the tender document; According to the core project information, the business field and industry category to which the tender document belongs are determined; Template screening and directory outline generation module, for screening the bid document template matching the user's demand based on the core project information, the business field and the industry category of the tender document, and referring to the industry knowledge graph and the bid document template database; Based on the bid document template and the core project information, the bid document directory outline is generated; Bid document draft generation module, for generating the bid document draft based on the bid document directory outline by extracting the required materials from the set material library through the large model LLM to fill in the materials; Quality audit correction module, for quality auditing of the bid document draft by NLP technology, checking the consistency of key term expression and factors for bid rejection, and generating problem prompts and correction opinions; According to the problem prompts and the correction opinions, the bid document draft is revised to generate a revised version of the bid document; Bid document final output module, for adjusting the format, optimizing the layout and manually reviewing the revised version of the bid document to output the final version of the bid document.
7. The large model and RAG-based file generation apparatus according to claim 6, characterized by, Also include: The knowledge base management module is used for classifying and integrating multi-source heterogeneous documents such as word, pdf and pictures in an enterprise, and generating a knowledge base.
8. The large model and RAG-based file generation apparatus according to claim 7, characterized by, In the tender offer draft generation module, the tender offer draft includes: solution overview, qualification certificate, company profile and past performance.
9. The large model and RAG-based file generation apparatus according to claim 8, characterized by, In the tender offer draft generation module, in the process of generating the tender offer draft, if several editors perform collaborative editing, the "collaborative bidding" mode is started, so that several editors can edit the chapter content they are responsible for online at the same time.
10. The large model and RAG-based file generation apparatus according to claim 9, characterized by, In the quality audit correction module, in the process of comprehensively auditing the tender offer draft by the NLP technology, spelling errors, non-standard formats and typesetting problems existing in the editing process of the tender offer draft are captured and corrected in real time.
Citation Information
Patent Citations
Bid file generation method and system based on large model, terminal and storage medium
CN117764039A
Bid text generation method and system based on large language model and storage medium
CN118982007A
Bidding document writing management system and biding document writing management method
CN120046588A
Bidding book and bidding book automatic generation method based on AI all-in-one machine
CN120068833A
Bidding document generation method based on retrieval enhancement generation and large language model
CN120297242A
Cited By
Intelligent bidding document generation and waste bidding risk confrontation optimization method and system
CN121883137A
Method and system for intelligent tender generation and countermeasures optimization against risks of tender rejection
CN121883137B
Intelligent bidding document generation method and system based on bidding document topology modeling
CN122047181A