Wind power project report automatic generation method, system and device in combination with multi-target reinforcement learning and large model technology and storage medium
By combining multi-objective reinforcement learning and large-model technology, the problem that existing technology is difficult to generate high-quality, professional and consistent wind power project reports is solved, and the language style consistency, logical clarity and data accuracy of wind power project reports are achieved.
Patent Information
- Application Number
- CN202510161510.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-27
AI Technical Summary
The existing automated text generation system is difficult to take into account the professionalism, style consistency and logical coherence of wind power project reports, and the generation quality is insufficient, making it difficult to achieve complex data analysis and field-specific accuracy.
A method combining multi-objective reinforcement learning and large-model technology is adopted to generate a wind power project report through the steps of agent initialization status, strategy execution, reward function design and calculation, and strategy update. This method optimizes style unity, semantic coherence, and text accuracy through multi-objective reinforcement learning, and dynamically adjusts reward weights to meet the needs of different reporting types.
The language style consistency, logical clarity and data accuracy of wind power project reports have been achieved, ensuring that the generated reports comply with the wind power industry standards and meet the formality and professional requirements of the reports.
Smart Images

Figure CN120218075A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of natural language processing, text generation, and reinforcement learning, and particularly relates to an automatic wind power project report generation method, system, device, and storage medium that combines multi-objective reinforcement learning and large model technology. Background Art
[0002] Wind power project reports are important tools for wind power project evaluation, planning, and decision-making, and are widely used in scenarios such as technical evaluation of wind power generation systems, economic benefit analysis, and environmental impact assessment. Usually, such reports need to cover complex technical content, detailed data analysis, and relevant policies and regulations, and report writing needs to take into account professionalism, style consistency, and logical coherence. However, existing automated text generation systems often have difficulty meeting the above requirements and have significant deficiencies in terms of generation quality, style control, and semantic coherence.
[0003] Current report generation algorithms mainly include: (1) template-driven method; (2) natural language generation method; (3) pre-trained large model generation method. The template-driven method uses predefined templates and fills in relevant information according to the input data to generate structured reports, but it is difficult to effectively expand to different generation requirements, the generated content lacks diversity, and it is difficult to achieve complex data analysis. The natural language generation method usually relies on rule bases and machine learning models and can generate relatively fluent text, but its effect is overly dependent on the design of rules and models, requires high professional knowledge, and it is difficult to stably output high-quality text. Although the pre-trained large model generation method can obtain better generation effects and has a relatively wide range of applications, it has insufficient understanding of professional knowledge in specific industries or fields, may generate factual errors or inaccurate information, and needs to be carefully verified. Summary of the Invention
[0004] The first object of the present invention is to provide an automatic wind power project report generation method that combines multi-objective reinforcement learning and large model technology for the problems mentioned above.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] An automatic wind power project report generation method that combines multi-objective reinforcement learning and large model technology includes the following steps:
[0007] S1: Input the basic information of the wind power project and the report generation instruction;
[0008] S2: The intelligent agent initializes the current state according to the input information;
[0009] S3: Policy execution;
[0010] S4: Reward function design and calculation;
[0011] S5: Policy update;
[0012] S6: Generate a wind power project report.
[0013] While adopting the above technical solution, the present invention can also adopt or combine the following technical solutions:
[0014] As a preferred technical solution of the present invention: in step S2, the input information includes the report structure target, the background information of the wind power project, and the text context.
[0015] As a preferred technical solution of the present invention: in step S3, the policy execution includes text generation, style adjustment, logic correction, and data reference.
[0016] As a preferred technical solution of the present invention: in step S4, the reward function includes a style unity reward function, a semantic coherence reward function, and a text accuracy reward function.
[0017] As a preferred technical solution of the present invention: step S5 is specifically to combine into a total reward value by means of weighted summation:
[0018] R_{\text{total}} = w_1\cdot R_{\text{style}} + w_2\cdot R_{\text{semantic}} + w_3\cdot R_{\text{accuracy}}
[0019] In the formula, R_{\text{style}}, R_{\text{semantic}}, and R_{\text{accuracy}} are the style unity, semantic coherence, and text accuracy rewards respectively, and w_{\text{1}}, w_{\text{2}}, w_{\text{3}} are the weight coefficients of each target.
[0020] As a preferred technical solution of the present invention: the weight coefficients are dynamically adjusted according to the type of report generated.
[0021] The second object of the present invention is to provide an automatic wind power project report generation system combining multi-objective reinforcement learning and large model technology.
[0022] To this end, the above object of the present invention is achieved by the following technical solutions:
[0023] An input module, which is used to input the basic information of the wind power project and the report generation instruction;
[0024] An initialization module, which is used to initialize the current state according to the input information;
[0025] A policy execution module, which is used to select appropriate text generation actions according to the current state of the model, gradually generate technical analysis content, and evaluate semantic coherence and style consistency after each paragraph;
[0026] A reward function design and calculation module, which is used to calculate the reward value for each generated content according to the reward function;
[0027] A policy update module, which is used to correct the model by adjusting actions when it is found that the style or logic does not match;
[0028] A wind power project report generation module, which is used to generate a wind power project report.
[0029] The third object of the present invention is to provide an electronic device.
[0030] To this end, the above object of the present invention is achieved by the following technical solutions:
[0031] An electronic device includes a memory and a processor. An executable program is stored in the memory, and the processor is configured to run the executable program to execute the steps of the wind power project report automatic generation method combining multi-objective reinforcement learning and large model technology as described above.
[0032] The fourth object of the present invention is to provide a non-volatile storage medium.
[0033] To this end, the above object of the present invention is achieved by the following technical solutions:
[0034] A non-volatile storage medium stores an executable program, and when the executable program is executed by a processor, it realizes the steps of the wind power project report automatic generation method combining multi-objective reinforcement learning and large model technology as described above.
[0035] The present invention provides a method, system, device, and storage medium for automatically generating a wind power project report by combining multi-objective reinforcement learning and large model technology, which has the following beneficial effects: Through multi-objective reinforcement learning, the present invention optimizes the style unity to ensure that the generated report has a consistent language style from beginning to end, meeting the formality requirements of the wind power project report; the model can optimize the logical coherence of the generated text through context information, avoiding problems such as inconsistent paragraph content or semantic jumps; with the help of the wind power industry knowledge base and data detection module, the model can ensure that the technical terms, data references, and conclusion derivations in the generated report comply with industry standards, avoiding false data or incorrect derivations; the model dynamically adjusts the weights of each reward objective according to the report type and target to ensure that the generated report is both professional and meets the actual needs. Brief Description of the Drawings
[0036] Figure 1 It is a flowchart of the method for automatically generating a wind power project report by combining multi-objective reinforcement learning and large model technology provided by the present invention. Detailed Embodiments
[0037] The present invention will be further described in detail with reference to the accompanying drawings and specific embodiments.
[0038] A method for automatically generating a wind power project report by combining multi-objective reinforcement learning and large model technology includes the following steps:
[0039] S1: Input the basic information of the wind power project and the report generation instruction;
[0040] S2: The agent initializes the current state according to the input information;
[0041] In step S2, the input information includes the report structure target, the background information of the wind power project, and the text context.
[0042] Text context state: It includes the generated report text content and the context information. The model adjusts the generation strategy based on the current generation progress, the relationship between the front and back texts, and the style information.
[0043] Wind power project background information state: It contains the input wind power project parameters (such as geographical location, wind conditions, power generation capacity, equipment type, etc.). These data help the model understand the project background so as to provide accurate technical analysis and conclusion derivations when generating the report.
[0044] Report structure target state: It defines the overall structure of the report, clarifies which modules are included (such as technical analysis, economic benefits, policy recommendations, etc.), the content that each module needs to cover, and its corresponding style requirements. This state helps the model switch between different report modules and generate professional language suitable for the content of each module.
[0045] S3: Strategy execution;
[0046] In step S3, strategy execution includes text generation, style adjustment, logic correction, and data reference.
[0047] Text generation action: Each step generates the next sentence or paragraph, and selects the appropriate language style and sentence structure based on the context. Especially in the wind power technology analysis part, the model will generate relevant technical evaluation and data derivation content based on the project background and industry knowledge.
[0048] Style adjustment action: When the model detects that the current generated paragraph deviates from the target style, it adjusts the generated language style to keep it consistent with the formality and professionalism of the wind power report. Reports in the wind power industry often require unified terminology, professional expressions, and consistent language structure, and style adjustment actions can effectively ensure this.
[0049] Logical correction action: During the generation process, the model will make corrections based on semantic detection when semantic jumps or logical breaks are found, and generate logically connected content to ensure the coherence of the report content, especially in the reasoning and data deduction parts across chapters.
[0050] Data citation action: When technical data needs to be cited, the model will extract relevant wind power project data from the knowledge base and generate appropriate citation sentences. This action ensures that the technical data, efficiency analysis, economic benefits and other contents in the report meet the standards and requirements of the wind power industry.
[0051] S4: Reward function design and calculation;
[0052] The reward function is the core of multi-objective reinforcement learning, guiding the model to optimize the generation behavior. This paper defines three main reward functions, which are used to optimize style consistency, semantic coherence and text accuracy respectively:
[0053] 1) Style consistency reward function:
[0054] Metrics: Measures the match between the generated text and the target style, ensuring that the entire report is consistent in terms of voice, wording, and sentence structure. The match is calculated using n-gram matching and voice analysis tools.
[0055] Reward: A positive reward is given when the style of the generated text is consistent with the target style, otherwise a negative penalty is given.
[0056] 2) Semantic coherence reward function:
[0057] Indicators: Calculate the semantic similarity and logical cohesion between the current paragraph and the previous and next paragraphs to ensure the logical progression of the content between paragraphs. Calculate the semantic similarity between paragraphs through a contextual embedding model (such as BERT).
[0058] Reward: If the generated paragraph can reasonably connect the content before and after, the reward is +1; otherwise, if the content is incoherent due to logical jumps or breaks, the reward is deducted.
[0059] 3) Text accuracy reward function:
[0060] Metrics: Detect whether the technical terms, data references, and logical derivations in the generated text are accurate. Use the wind power industry knowledge base and data detection module to ensure that the cited data and inference conclusions conform to the wind power industry standards.
[0061] Reward: If the generated text correctly cites the technical parameters, power generation efficiency and other data of the wind power project, the reward is +3; if the citation is incorrect or the inference logic does not match, the reward is deducted.
[0062] S5: Policy update;
[0063] Combine them into a total reward value by weighted summation:
[0064] R_{\text{total}} = w_1\cdot R_{\text{style}} + w_2\cdot R_{\text{semantic}} + w_3\cdot R_{\text{accuracy}}
[0065] Wherein, R_{\text{style}}, R_{\text{semantic}} and R_{\text{accuracy}} are the rewards for style unity, semantic coherence and text accuracy respectively, and w_{\text{1}}, w_{\text{2}}, w_{\text{3}} are the weight coefficients of each objective.
[0066] The weight coefficients can be dynamically adjusted according to the type of generated report. For example, in a technical report, the weight w_{\text{3}} of text accuracy can be relatively high, while in a business report, the weight w_{\text{1}} of style unity may be higher.
[0067] S6: Generate a wind power project report.
[0068] The present invention adopts the Proximal Policy Optimization algorithm (PPO) based on policy gradient, and this algorithm shows stable optimization effects when dealing with multi-objective tasks.
[0069] Specifically, the above-mentioned method for automatically generating a wind power project report by combining multi-objective reinforcement learning and large model technology is implemented in the following way:
[0070] Example 1: Generation of Technical Analysis Report for Wind Power Project
[0071] S1: Input the basic information of the wind power project (such as project name, geographical location, installed capacity, etc.) and the report generation instruction (such as generating a technical analysis report) entered by the user;
[0072] S2: The intelligent agent initializes the current state according to the input information, including the report structure, project information, and initial text context;
[0073] S3: Policy execution. The model selects appropriate text generation actions according to the current state, gradually generates the technical analysis content, and evaluates the semantic coherence and style consistency after each paragraph.
[0074] S4: Reward function design and calculation: For each generated paragraph, the model calculates the reward value according to the reward function;
[0075] S5: Policy update. If it is found that the style or logic does not match, the model will correct it by adjusting the actions;
[0076] S6: Finally, output a technical analysis report for the wind power project with consistent language style, clear logic, and accurate data.
[0077] Example 2: Generation of Economic Benefit Analysis Report for Wind Power Project
[0078] S1: Input the financial data of the wind power project (such as investment cost, operation cost, return period) and the report generation instruction (such as generating an economic benefit analysis report) entered by the user;
[0079] S2: The intelligent agent initializes the current state according to the input information;
[0080] S3: Policy execution. The model generates analysis content, such as return on investment, operation cost prediction, etc., and evaluates the quality of the generated content according to the data accuracy and semantic coherence.
[0081] S4: Reward function design and calculation: For each generated paragraph, the model calculates the reward value according to the reward function;
[0082] S5: Policy update. If it is found that the style or logic does not match, the model will correct it by adjusting the actions;
[0083] S6: Finally, output an economic benefit analysis report for the wind power project that meets the financial data and economic evaluation criteria.
[0084] An automatic report generation system for wind power projects that combines multi-objective reinforcement learning and large model technology. The system includes the following modules:
[0085] Input module. The input module is used to input the basic information of the wind power project and the report generation instruction;
[0086] An initialization module, which is used to initialize the current state according to the input information;
[0087] A policy execution module, which is used to select a suitable text generation action for the model according to the current state, gradually generate technical analysis content, and evaluate semantic coherence and style consistency after each paragraph;
[0088] A reward function design and calculation module, which is used to calculate the reward value for each generated paragraph according to the reward function;
[0089] A policy update module, which is used to correct the model by adjusting actions when it is found that the style or logic does not match;
[0090] A wind power project report generation module, which is used to generate a wind power project report.
[0091] The present invention also provides an electronic device, including a processor and a memory for storing processor-executable instructions. Among them, when the processor is set to execute the executable instructions, the method steps of automatically generating a wind power project report by combining multi-objective reinforcement learning and large model technology as described above are realized.
[0092] The present invention also provides a non-volatile storage medium, in which an executable program is stored. When the executable program is executed by a processor, the method steps of automatically generating a wind power project report by combining multi-objective reinforcement learning and large model technology as described above are realized.
[0093] The above specific embodiments are used to explain the present invention, and are only the preferred embodiments of the present invention, rather than limiting the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A method for automatically generating wind power project reports combining multi-objective reinforcement learning and large model technology, characterized in that: The steps include: S1: Input basic information of wind power project and report generation instructions; S2: The agent initializes the current state based on the input information; S3: Strategy execution; S4: Reward function design and calculation; S5: strategy update; S6: Generate wind power project report.
2. The method for automatically generating wind power project reports combining multi-objective reinforcement learning and large model technology according to claim 1 is characterized in that: In step S2, the input information includes report structure objectives, wind power project background information, and text context.
3. The method for automatically generating wind power project reports combining multi-objective reinforcement learning and large model technology according to claim 1 is characterized in that: In step S3, strategy execution includes text generation, style adjustment, logic correction, and data reference.
4. The method for automatically generating wind power project reports combining multi-objective reinforcement learning and large model technology according to claim 1 is characterized in that: In step S4, the reward function includes a style consistency reward function, a semantic coherence reward function, and a text accuracy reward function.
5. The method for automatically generating wind power project reports combining multi-objective reinforcement learning and large model technology according to claim 1 is characterized in that: Step S5 specifically combines the total reward value into a weighted sum: R_{\text{total}} =w_1\cdotR_{\text{style}}+w_2\cdotR_{\text{semantic}}+w_3\cdotR_{\text{accuracy}} Where R_{\text{style}}, R_{\text{semantic}} and R_{\text{accuracy}} are the rewards for style consistency, semantic coherence and text accuracy respectively, and w_{\text{1}}, w_{\text{2}}, w_{\text{3}} are the weight coefficients of each objective.
6. The method for automatically generating wind power project reports combining multi-objective reinforcement learning and large model technology according to claim 5 is characterized in that: The weighting factor is adjusted dynamically based on the type of report generated.
7. A wind power project report automatic generation system combining multi-objective reinforcement learning and large model technology, characterized by: The system includes the following modules: An input module, the input module is used to input basic information of the wind power project and report generation instructions; An initialization module, the initialization module is used to initialize the current state according to input information; A strategy execution module, which is used to select appropriate text generation actions for the model according to the current state, gradually generate technical analysis content, and evaluate semantic coherence and style consistency after each paragraph; A reward function design and calculation module, which is used to calculate a reward value for each generated content according to a reward function; A strategy update module, wherein when a style or logic is found to be inconsistent, the model will make corrections by adjusting actions; A wind power project report generation module, wherein the wind power project report generation module is used to generate a wind power project report.
8. An electronic device, comprising a memory and a processor, characterized in that: An executable program is stored in the memory, and the processor is configured to run the executable program to execute the steps of the method for automatically generating a wind power project report combining multi-objective reinforcement learning and large model technology as claimed in any one of claims 1 to 6.
9. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores an executable program, and when the executable program is executed by the processor, the method steps for automatically generating a wind power project report combining multi-objective reinforcement learning and large model technology as claimed in any one of claims 1 to 6 are implemented.