Energy document generation method and equipment based on RAG and large model fusion

By integrating RAG with a large model, and combining preprocessing with knowledge base verification, the problems of parameter errors and logical inconsistencies in traditional energy document generation are solved, enabling the efficient generation of high-quality, logically coherent energy documents.

CN120930634APending Publication Date: 2025-11-11BEIJING SHENGXINNUO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511113092.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional energy documentation relies on manual writing, which can lead to problems such as parameter errors, logical inconsistencies, and incomplete information. Large models also suffer from a lack of professional knowledge in the energy field, resulting in low document quality.

Method used

The method adopts a fusion approach based on RAG and large models. Input is obtained through a human-computer interaction interface, the original materials are preprocessed to generate text blocks, the parameters and logic are verified using an energy domain knowledge base, supplementary information is obtained from the knowledge base by combining keywords, and finally high-quality documents are generated.

Benefits of technology

Perform parameter and logic validation at the text block level to avoid error accumulation, enhance content integrity and authority, ensure document logical rigor and user compliance, and improve generation efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930634A_ABST
    Figure CN120930634A_ABST
Patent Text Reader

Abstract

The invention relates to the field of natural language processing and generative artificial intelligence, in particular to an energy document generation method and device based on RAG and large model fusion. The method comprises the following steps: preprocessing an original material of a user to generate a plurality of first text blocks; on the basis of a pre-constructed energy field knowledge base, checking and correcting the parameters and context logic of each first text block to obtain a second text block; inputting the energy theme, the first cue word and all the second text blocks into the fine-tuned large model to generate a preliminary document; extracting keywords from the preliminary document, and acquiring supplementary technical description and limiting conditions from the energy field knowledge base according to the extracted keywords; displaying the preliminary document, the supplementary technical description and the limiting conditions through a human-computer interaction interface; and inputting the preliminary document, the second cue word, the supplementary technical description and the limiting condition into the fine-tuned large model to generate a final document. According to the invention, the quality of the generated document is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of natural language processing and generative artificial intelligence, specifically to a method and device for generating energy documents based on the fusion of RAG and large models. Background Technology

[0002] In the energy industry, with the rapid development of information technology and the continuous expansion of the energy sector, the generation and management of energy documentation faces increasingly complex challenges. Energy documentation covers a wide range of aspects, from energy project planning and equipment operation and maintenance to the interpretation of energy policies. Its accuracy and professionalism are directly related to decision-making, project implementation, and safe operation in the energy industry.

[0003] Traditional energy documentation generation relies primarily on manual writing, which requires writers to possess deep energy expertise and consumes a significant amount of time and effort. Due to the limitations of manual writing, documents are prone to errors such as incorrect parameters, logical inconsistencies, and incomplete information, resulting in inconsistent document quality.

[0004] With the rise of artificial intelligence, natural language processing technology has been widely used in document generation. However, while large models have powerful capabilities in language understanding and generation, they lack a deep understanding of energy-related expertise, which may lead to errors in the generated documents and affect their quality. Summary of the Invention

[0005] To address the aforementioned problems in the prior art, this invention proposes an energy document generation method and device based on the fusion of RAG (Retrieval-Augmented Generation) and large models, thereby improving the quality of the generated documents.

[0006] In a first aspect, the present invention proposes an energy document generation method based on the fusion of RAG and large models, the method comprising: The system obtains the user's input of energy-related topics, raw materials, and initial prompts through a human-computer interaction interface. The original material is preprocessed to generate several first text blocks; Based on a pre-built knowledge base in the energy field, the parameters and context logic of each first text block are verified and corrected to obtain the second text block; Input the energy theme, the first prompt word, and all the second text blocks into the fine-tuned large model to generate a preliminary document; Keywords are extracted from the preliminary document, and supplementary technical descriptions and limitations are obtained from the energy field knowledge base based on the extracted keywords. The preliminary document, the supplementary technical description, and the limiting conditions are displayed through the human-computer interaction interface; The second prompt word input by the user is obtained through the human-computer interaction interface; The initial document, the second prompt, the supplementary technical description, and the constraints are input into the fine-tuned large model to generate the final document.

[0007] Preferably, the step of "preprocessing the original material to generate several first text blocks" includes: Read the text and tables from the original materials and interpret the images to generate a text file; The text file is processed by removing irrelevant characters and standardizing unit symbols. A natural language processing model is then used to identify and remove irrelevant content, resulting in a cleaned text file. The irrelevant content includes advertising text and copyright notices. The parameters, mathematical expressions, and logical connectors in the cleaned text file are annotated; The cleaned text file is divided into first text blocks. During the division, the integrity of parameters and mathematical expressions is maintained according to the annotations, and the same group of logical connectors are kept in the same text block. The parameters include: technical parameters, policy parameters, and equipment parameters.

[0008] Preferably, the step of "verifying and correcting the parameters and context logic of each first text block based on a pre-built energy domain knowledge base to obtain a second text block" includes: For each parameter in the first text block, a first verification result is obtained based on a pre-built knowledge base in the energy field. Perform context logic validation on each of the first text blocks to obtain the second validation result; For each of the first text blocks, the corresponding second text block is obtained by correcting it according to the first and second verification results.

[0009] Preferably, the step of "verifying the parameters in each of the first text blocks based on a pre-built knowledge base in the energy field to obtain a first verification result" includes: Extract parameters for each of the first text blocks; Based on the aforementioned energy field knowledge base, the validity of parameter names is verified, and the parameter values ​​are checked to see if they are within a reasonable range. Output the first verification result; The first verification result includes: a flag indicating whether the parameter name and parameter value verification passed / failed, and the correct parameter name and parameter value range.

[0010] Preferably, the step of "performing contextual logic verification on each of the first text blocks to obtain a second verification result" includes: For each of the first text blocks, verify whether the logic within the text block is reasonable; Verify the logical coherence across text blocks between the first text blocks; Output the second verification result; The second verification result includes: the logical verification result within each of the first text blocks and the logical verification result across text blocks.

[0011] Preferably, the step of "verifying the internal logic of each first text block as reasonable" includes: For each of the first text blocks, sentence component integrity is checked, and whether contradictory conjunctions are used correctly is identified; Based on the energy sector rule base, regular expressions are used to match the violation patterns in each of the first text blocks; Based on the energy sector rule base, verify whether the mathematical expression in each of the first text blocks is correct.

[0012] Preferably, the step of "verifying the logical coherence across text blocks between the first text blocks" includes: Track whether the referential relationships between pronouns and nouns in the first text block are clear and consistent across text blocks; Extract the time expression from the first text block to verify the time sequence rationality of different text blocks; Use causal reasoning models to detect hypothesis-test breakpoints; Calculate the topic similarity between adjacent text blocks, and indicate logical jumps when the score is low.

[0013] Preferably, the step of "correcting each first text block according to the corresponding first verification result and second verification result to obtain the corresponding second text block" includes: For each of the first text blocks, if there is a parameter type or parameter value verification failure, but no internal logic verification failure, then the parameter type or parameter value is corrected according to the energy field knowledge base. For each of the first text blocks, if there is an internal logic verification failure, manual correction is performed through the human-computer interaction interface; If a cross-text block logic check fails, manual correction can be performed through the human-computer interaction interface.

[0014] Preferably, the keywords include: equipment model, geographical location, and standards and specifications; The supplementary technical description includes: the extracted technical parameters, application scenarios and compatibility descriptions corresponding to the device model, and the local policies corresponding to the extracted geographical location; The constraints include: environmental restrictions, operation and maintenance requirements and security warnings corresponding to the extracted device model, protection requirements corresponding to the extracted geographical location and timeliness prompts corresponding to the extracted standard specifications.

[0015] In a second aspect, the present invention provides a computer-readable storage device storing a computer program that can be loaded by a processor and executed as described above.

[0016] In a second aspect, the present invention provides a computer-readable storage device storing a computer program that can be loaded by a processor and executed as described above.

[0017] The present invention has the following beneficial effects: The energy document generation method proposed in this invention, based on the fusion of RAG and large models, performs parameter / logic verification at the text block stage. Compared to existing technologies that perform overall verification after document generation, this avoids error accumulation. For example, verifying the rationality of "coolant temperature values" in nuclear power documents in advance can prevent derivation errors during subsequent document generation. Supplementing technical descriptions and limitations by retrieving them from the knowledge base based on keywords in the initial document enhances the completeness and authority of the content. Therefore, this invention moves knowledge verification from the document level to the text block level, constructing a dual quality assurance system of pre-generation verification and post-generation enhancement. Furthermore, allowing users to input prompts during both document generation stages ensures that the final generated document meets both professional standards and user needs, improving the quality and efficiency of document generation.

[0018] By separately verifying the logical rationality within text blocks and the logical coherence across text blocks, the rigor of document logic is fully guaranteed, effectively avoiding logical errors and content gaps, and laying a solid foundation for generating high-quality, logically coherent energy documents.

[0019] By retrieving supplementary technical specifications and limitations related to equipment model, geographical location, and standards and regulations, users can gain a comprehensive understanding of equipment performance, applicable scenarios, and local policies. At the same time, they can clarify environmental limitations, operation and maintenance requirements, and safety warnings, making the technical content described in the final document more accurate. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the main steps of an embodiment of the energy document generation method based on the fusion of RAG and large models in this invention. Detailed Implementation

[0021] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] It should be noted that in the description of this invention, the terms "first" and "second" are used merely for ease of description and do not indicate or imply the relative importance of the described devices, elements, or parameters, and therefore should not be construed as limiting the invention. Furthermore, the term "and / or" in this invention merely describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0024] Figure 1 This is a schematic diagram illustrating the main steps of an embodiment of the energy document generation method based on the fusion of RAG and large models in this invention. Figure 1 As shown, the method in this embodiment includes steps S10-S80: Step S10: Obtain the energy theme, original materials, and first prompt word input by the user through the human-computer interaction interface.

[0025] Among them, the energy theme is used to limit the scope and focus of the generated document, such as "offshore wind power operation and maintenance"; the first prompt word can include task instructions and format requirements, etc., to control the initial document generation framework and ensure that the content meets the user's basic needs, such as "generate a report based on the materials I uploaded, describing the operation and maintenance methods, procedures and precautions, etc."; the original materials are provided by the user and are the basic materials used to generate the document.

[0026] Step S20: Preprocess the original material to generate several first text blocks.

[0027] Specifically, step S20 may include steps S21-S24: Step S21: Read the text and tables in the original material and interpret the images to generate a text file.

[0028] Specifically, OCR (Optical Character Recognition) technology or document parsing tools are used to extract text and table content from the original materials, ensuring that the row and column relationships are preserved when structured data (such as Excel tables) is converted into plain text. For example, the phrase "Wind turbine model: SG-5.0-172, rated power: 5MW" in a table is parsed into coherent text.

[0029] Text descriptions of images are generated using multimodal models (such as GPT-4V), with a focus on extracting information related to the energy theme (such as technical parameter labels in equipment structure diagrams and data points in trend charts).

[0030] Step S22: Remove irrelevant characters from the text file and standardize unit symbols. Use a natural language processing model to identify and remove irrelevant content to obtain a cleaned text file.

[0031] Specifically, remove garbled text and special characters (such as HTML tags). Irrelevant characters such as non-text elements (e.g., headers and footers); standardized unit symbols (e.g., "kWh" is standardized to "kW·h") to avoid ambiguity in subsequent parsing; and the use of natural language processing models (e.g., BERT-CRF) to identify and remove irrelevant content such as advertising text and copyright notices, while retaining technical content. For example, "Contact us: 400-123-xxxx" in a text block was identified as an advertisement and deleted.

[0032] Step S23: Annotate the parameters, mathematical expressions, and logical connectors in the cleaned text file.

[0033] The parameters include technical parameters, policy parameters, and equipment parameters. Technical parameters may include efficiency (e.g., "η=95%"), capacity (e.g., "energy storage system: 100MWh"), and emissions (e.g., "CO2≤500g / kWh"). Equipment parameters may include equipment model (e.g., "inverter model: SUN2000-8KTL") and operating threshold (e.g., "operating temperature: -30°C~60°C"). Policy parameters may include subsidy amount (e.g., "0.3 yuan / kWh") and quota ratio (e.g., "renewable energy share ≥30%").

[0034] Step S24: Divide the cleaned text file into the first text block. During the division, maintain the integrity of parameters and mathematical expressions according to the annotations, and ensure that the same group of logical connectors are in the same text block.

[0035] Parameter integrity: If a parameter spans multiple sentences (e.g., “Rated power: 5MW, corresponding to international standard IEC 61400”), it will be forcibly merged into a single text block.

[0036] Logical coherence: The same set of logical connectors (such as "because...therefore...") and their related content are kept in the same block to avoid semantic breaks.

[0037] Step S30: Based on the pre-built energy domain knowledge base, verify and correct the parameters and context logic of each first text block to obtain the second text block.

[0038] Specifically, step S30 may include steps S31-S33: Step S31: Verify the parameters in each first text block based on a pre-built knowledge base in the energy field to obtain the first verification result. This step can specifically include steps S311-S313: Step S311: Extract parameters for each first text block.

[0039] Step S312: Based on the knowledge base in the energy field, verify the legality of the parameter name and check whether the parameter value is within a reasonable range.

[0040] For example, in step S311, “blade length 56 meters”, “rated wind speed 12 m / s” and “operating temperature -30°C to 50°C” were extracted from a certain first text block (note: the parameter name here is missing the word “degree”).

[0041] The verification process in step S312 is as follows: (1) Blade length 56 meters → Parameter name (blade length) is valid, and the value is within a reasonable range (20-100 meters) → Pass; (2) Rated wind speed 12m / s → Parameter name (rated wind speed) is valid and the value is within a reasonable range (8-15m / s) → Pass; (3) Operating temperature -30°C to 50°C → Parameter name (operating temperature) is invalid, range conforms to standard (-40°C to 60°C) → Failure, the correct parameter name and range are "operating temperature -30°C to 50°C" Step S313: Output the first verification result.

[0042] The first verification result includes: a flag indicating whether the parameter name and parameter value verification passed / failed, and the correct parameter name and parameter value range.

[0043] Step S32: Perform context logic verification on each first text block to obtain a second verification result. Specifically, this step may include steps S321-S323: Step S321: For each first text block, verify whether the logic inside the text block is reasonable.

[0044] Specifically, this step may include steps S3211-S3213: Step S3211: For each first text block, check the sentence component integrity and identify whether the contradictory conjunctions are used correctly.

[0045] For example, a first text block is: "A certain 5MW wind turbine has a cut-in wind speed of 3m / s and a rated wind speed of 12m / s. However, its power curve shows..." Power generation begins at 2 m / s, which contradicts the definition of cut-in wind speed. According to the formula P = 0.5 × ρ × A × v³ × Cp, when the air density ρ... When the wind speed is 1.225 kg / m³, the swept area A = 18600 m², and Cp = 0.42, the theoretical power at a wind speed of 3 m / s is calculated to be... 1.2MW. " (1) Sentence component integrity check: Check whether the subject, predicate, and key parameters are complete.

[0046] By example: The rated wind speed is 12 m / s. (Subject "rated wind speed" + predicate "is" + parameter "12m / s"); Example of failure: "Power generation began at a speed of 2 m / s." (Subject missing, "crew" needs to be added).

[0047] (2) Identification of contradictory conjunctions: Detect logical conflicts before and after contradictory conjunctions such as "however" and "but".

[0048] This example is contradictory: Previous text: Cut-in wind speed = 3 m / s (Power generation is only possible if the speed is ≥3m / s); The following text: "Electricity was generated at 2 m / s" → The logical contradiction indicates that the word "however" is used correctly in the text block.

[0049] Step S3212: Based on the energy sector rule base, use regular expressions to match the violation patterns in each first text block.

[0050] For example, Regular expression pattern: r"cut-in wind speed [is|is].*?but.*?[lower than|less than].*?cut-in wind speed" Matching example: "The cut-in wind speed is 3 m / s, but it generates electricity at 2 m / s." → Violation of regulations.

[0051] Step S3213: According to the rule base of the energy field, verify whether the mathematical expression in each first text block is correct.

[0052] For example, the rule base stores valid wind power formulas: allowed_formulas = ["P=0.5×ρ×A×v³×Cp","P=ρ*A*v³*Cp / 2"]; Here, the two formulas within the square brackets are the allowed equivalent forms.

[0053] Therefore, the formula P=0.5×ρ×A×v³×Cp in this example is valid.

[0054] Next, we'll verify the parameters: P = 0.5 × 1.225 × 18600 × (3)³ × 0.42 ≈ 1.2MW; The calculation result matches "1.2MW" in the text block → Pass.

[0055] If the text block says "the power is 5MW at 3m / s" → the values ​​are contradictory.

[0056] Step S322: Verify the logical coherence across text blocks between the first text blocks.

[0057] Suppose we have the following three consecutive first text blocks (Offshore Wind Power Operation and Maintenance Report): Text block A: "In October 2023, a typhoon passed through a wind farm, causing cracks to appear on the blade of wind turbine #15. The maintenance team immediately shut down the turbine and applied for spare parts replacement." Text block B: "Since the delivery cycle for imported blades is six months, the project team temporarily used 3D printing technology to repair the cracks, ensuring that the wind turbine could resume power generation in November." Text block C: "This incident prompted the wind farm to revise its emergency response plan, incorporating 3D printing repairs into its rapid response process and stockpiling localized spare parts." This step may specifically include steps S3221-S3224: Step S3221: Track whether the referential relationships between pronouns and nouns in the first text block are clear and consistent across text blocks.

[0058] For example, “#15 wind turbine” in text block A and “wind turbine” in text block B refer to the same thing and are unambiguous.

[0059] Error example: If text block C contains This incident prompted the wind farm to revise its emergency response plan. Change to This event prompted it Revise the emergency response plan It is necessary to clarify whether "it" refers to the wind farm, the team, or the wind turbine.

[0060] Step S3222: Extract the time expression from the first text block and verify the rationality of the timing of different text blocks.

[0061] For example: text block A (October typhoon) → text block B (November recovery) → text block C (subsequent contingency plan revision), the sequence is reasonable.

[0062] Step S3223: Use a causal inference model to detect the hypothesis-verification break.

[0063] The purpose of this step is to check whether the causal chain of technology decisions is complete.

[0064] For example: Text block A (leaf damage) → Text block B (3D printing repair), the cause and effect are reasonable (import delay → innovative solution).

[0065] Text block B (temporary fix) → Text block C (contingency plan revision), the cause and effect are reasonable (experience feedback to management process).

[0066] Case of breakage: If text block C reads "Therefore the project team purchased a new type of drone", it is not directly related to blade repair.

[0067] Step S3224: Calculate the topic similarity between adjacent text blocks, and prompt logical jump when the score is low.

[0068] For example: Text blocks A and B have a high score for topic similarity (wind turbine failure → emergency technology).

[0069] The similarity of themes between text blocks B and C (technical repair → management improvement) is medium (the keyword "contingency plan" needs to be connected).

[0070] Low-scoring example: If text block C is changed to "At the same time, the wind farm began to build an employee canteen", the topic will be completely irrelevant.

[0071] Step S323: Output the second verification result.

[0072] The second verification result includes: the logical verification result within each first text block and the logical verification result across text blocks.

[0073] Step S33: For each first text block, correct it according to the corresponding first and second verification results to obtain the corresponding second text block. This step may specifically include steps S331-S333: Step S331: For each first text block, if there is a parameter type or parameter value verification failure, but no internal logic verification failure, then correct the parameter type or parameter value according to the energy field knowledge base.

[0074] If the internal logic is incorrect, correcting the parameters is meaningless. Therefore, we only correct the parameters for text blocks with correct internal logic.

[0075] Step S332: For each first text block, if there is an internal logic verification failure, manual correction is performed through the human-computer interaction interface.

[0076] Errors requiring manual correction fall into two categories: those where both the internal logic and parameters are incorrect, or those where the parameters are correct but the internal logic is incorrect.

[0077] Step S333: If cross-text block logic verification fails, manual correction is performed through the human-computer interaction interface.

[0078] Cross-text block logical errors are relatively complex cases, and manual correction will yield better results.

[0079] Step S40: Input the energy theme, the first prompt word, and all the second text blocks into the fine-tuned large model to generate a preliminary document.

[0080] The large model is pre-tuned for the energy sector to adapt to energy-related document generation.

[0081] Step S50: Extract keywords from the preliminary document, and obtain supplementary technical descriptions and constraints from the energy knowledge base based on the extracted keywords.

[0082] The keywords include: equipment model, geographical location, and standards and specifications; supplementary technical descriptions include: the technical parameters, application scenarios, and compatibility descriptions corresponding to the extracted equipment model, as well as the local policies corresponding to the extracted geographical location; the limiting conditions include: the environmental restrictions, operation and maintenance requirements, and security warnings corresponding to the extracted equipment model, as well as the protection requirements corresponding to the extracted geographical location and the timeliness reminders corresponding to the extracted standards and specifications.

[0083] For example, if a device model keyword is extracted: XX-100KTL inverter, the following content is retrieved from the knowledge base: (1) Supplementary technical specifications: (1a) Technical parameters: Maximum input voltage: 1100V (MPPT voltage range 200V-1000V); European efficiency: 98.6% (CEC efficiency: 98.2%) Communication protocol: Supports IEC61850-7-420.

[0084] (1b) Application scenarios: This model is suitable for centralized photovoltaic power plants (1MW and above) and is not recommended for rooftop distributed projects (a micro inverter needs to be used instead).

[0085] (1c) Compatibility Notes: It needs to be used in conjunction with Company A's YY Intelligent Management System; third-party monitoring platforms require additional protocol converter configuration.

[0086] (2) Restrictions: (2a) Environmental limitations: Operating temperature range: -25℃ to 60℃. Power derating is 0.5% for every 100m increase in altitude above 2000 meters.

[0087] (2b) Operation and maintenance requirements: The cooling fan must be cleaned every 6 months (refer to Chapter 4 of the manufacturer's maintenance manual).

[0088] (2c) Safety warning: Exposed DC-side terminals must be fitted with protective covers (meeting NFPA 70E arc protection requirements).

[0089] Step S60: Display the preliminary document, supplementary technical specifications, and limitations through a human-computer interaction interface.

[0090] Users can view the initial document through the human-computer interaction interface and refer to the retrieved supplementary technical descriptions and limitations. They can also input their suggestions for modifying the document as a second prompt.

[0091] Step S70: Obtain the second prompt word input by the user through the human-computer interaction interface.

[0092] The second suggestion keyword can be used to adjust the level of detail, focus, or compliance requirements of the final document based on user feedback on the initial document. For example, it can tell you which content in the initial document of the large model needs to be adjusted, and which parts of the document need to be modified or supplemented according to the supplementary technical instructions and constraints in the search.

[0093] Step S80: Input the preliminary document, second prompt, supplementary technical description and constraints into the fine-tuned large model to generate the final document.

[0094] Although the steps in the above embodiments are described in the above order, those skilled in the art will understand that in order to achieve the effect of this embodiment, different steps do not need to be executed in such an order. They can be executed simultaneously (in parallel) or in a reverse order. These simple variations are all within the protection scope of this invention.

[0095] Based on the above method embodiments, the present invention also provides an embodiment of a computer-readable storage device storing a computer program that can be loaded by a processor and executed as described above.

[0096] The computer-readable storage device may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0097] Those skilled in the art will recognize that the method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the invention.

[0098] The technical solution of the present invention has now been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions resulting from these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. An energy document generation method based on the fusion of RAG and large models, characterized in that, The method includes: The system obtains the user's input of energy-related topics, raw materials, and initial prompts through a human-computer interaction interface. The original material is preprocessed to generate several first text blocks; Based on a pre-built knowledge base in the energy field, the parameters and context logic of each first text block are verified and corrected to obtain the second text block; Input the energy theme, the first prompt word, and all the second text blocks into the fine-tuned large model to generate a preliminary document; Keywords are extracted from the preliminary document, and supplementary technical descriptions and limitations are obtained from the energy field knowledge base based on the extracted keywords. The preliminary document, the supplementary technical description, and the limiting conditions are displayed through the human-computer interaction interface; The second prompt word input by the user is obtained through the human-computer interaction interface; The initial document, the second prompt, the supplementary technical description, and the constraints are input into the fine-tuned large model to generate the final document.

2. The energy document generation method based on RAG and large model fusion according to claim 1, characterized in that, The step of "preprocessing the original material to generate several first text blocks" includes: Read the text and tables from the original materials and interpret the images to generate a text file; The text file is processed by removing irrelevant characters and standardizing unit symbols. A natural language processing model is then used to identify and remove irrelevant content, resulting in a cleaned text file. The irrelevant content includes advertising text and copyright notices. The parameters, mathematical expressions, and logical connectors in the cleaned text file are annotated; The cleaned text file is divided into first text blocks. During the division, the integrity of parameters and mathematical expressions is maintained according to the annotations, and the same group of logical connectors are kept in the same text block. The parameters include: technical parameters, policy parameters, and equipment parameters.

3. The energy document generation method based on RAG and large model fusion according to claim 1, characterized in that, The steps of "verifying and correcting the parameters and context logic of each first text block based on a pre-built knowledge base in the energy field to obtain a second text block" include: For each parameter in the first text block, a first verification result is obtained based on a pre-built knowledge base in the energy field. Perform context logic validation on each of the first text blocks to obtain the second validation result; For each of the first text blocks, the corresponding second text block is obtained by correcting it according to the first and second verification results.

4. The energy document generation method based on RAG and large model fusion according to claim 3, characterized in that, The step of "validating the parameters in each of the first text blocks based on a pre-built knowledge base in the energy field to obtain a first validation result" includes: Extract parameters for each of the first text blocks; Based on the aforementioned energy field knowledge base, the validity of parameter names is verified, and the parameter values ​​are checked to see if they are within a reasonable range. Output the first verification result; The first verification result includes: a flag indicating whether the parameter name and parameter value verification passed / failed, and the correct parameter name and parameter value range.

5. The energy document generation method based on RAG and large model fusion according to claim 4, characterized in that, The step of "performing context logic validation on each of the first text blocks to obtain a second validation result" includes: For each of the first text blocks, verify whether the logic within the text block is reasonable; Verify the logical coherence across text blocks between the first text blocks; Output the second verification result; The second verification result includes: the logical verification result within each of the first text blocks and the logical verification result across text blocks.

6. The energy document generation method based on RAG and large model fusion according to claim 5, characterized in that, The step of "verifying the internal logic of each first text block" includes: For each of the first text blocks, sentence component integrity is checked, and whether contradictory conjunctions are used correctly is identified; Based on the energy sector rule base, regular expressions are used to match the violation patterns in each of the first text blocks; Based on the energy sector rule base, verify whether the mathematical expression in each of the first text blocks is correct.

7. The energy document generation method based on RAG and large model fusion according to claim 5, characterized in that, The steps for "verifying the logical coherence across text blocks between the first text blocks" include: Track whether the referential relationships between pronouns and nouns in the first text block are clear and consistent across text blocks; Extract the time expression from the first text block to verify the time sequence rationality of different text blocks; Use causal inference models to detect hypothesis-test breakpoints; Calculate the topic similarity between adjacent text blocks, and indicate logical jumps when the score is low.

8. The energy document generation method based on RAG and large model fusion according to claim 7, characterized in that, The step of "correcting each first text block according to the corresponding first and second verification results to obtain the corresponding second text block" includes: For each of the first text blocks, if there is a parameter type or parameter value verification failure, but no internal logic verification failure, then the parameter type or parameter value is corrected according to the energy field knowledge base. For each of the first text blocks, if there is an internal logic verification failure, manual correction is performed through the human-computer interaction interface; If a cross-text block logic check fails, manual correction can be performed through the human-computer interaction interface.

9. The energy document generation method based on RAG and large model fusion according to claim 1, characterized in that, The keywords include: equipment model, geographical location, and standards and specifications; The supplementary technical description includes: the extracted technical parameters, application scenarios and compatibility descriptions corresponding to the device model, and the local policies corresponding to the extracted geographical location; The constraints include: environmental restrictions, operation and maintenance requirements and security warnings corresponding to the extracted device model, protection requirements corresponding to the extracted geographical location and timeliness prompts corresponding to the extracted standard specifications.

10. A computer-readable storage device, characterized in that, The computer program is stored that can be loaded by a processor and execute the method as described in any one of claims 1-9.