New energy station technology complete dispatching report compiling method based on ORC and new energy AI large model

By automating the processing of technical due diligence documents for new energy sites through the LangChain task chain and the new energy AI big model, the problems of duplicate information search and data consistency are solved, efficient and accurate report preparation and review are achieved, labor costs are reduced, and the overall benefits of the report are improved.

CN120688512APending Publication Date: 2025-09-23BEIJING RETEC NEW ENERGY TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510824077.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the process of compiling technical due diligence reports for new energy stations, there are problems such as time-consuming repeated data searches, poor data consistency, and inconsistent information, which leads to heavy workload, low efficiency, and difficulty in ensuring accuracy.

Method used

LangChain is used to build task chains, combining OCR recognition, semantic analysis and new energy AI big models to automatically process documents related to technical due diligence of new energy stations, generate standardized electronic text through OCR recognition, use the new energy AI big model to extract core fields and compare and verify them with external knowledge bases, and generate and review reports.

Benefits of technology

It improves the efficiency of report preparation and review, reduces labor costs, ensures data consistency and traceability, and enhances the professionalism and reliability of reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688512A_ABST
    Figure CN120688512A_ABST
Patent Text Reader

Abstract

The invention discloses a new energy station technology exhaustion report compiling method based on ORC and a new energy AI large model. The new energy station technology exhaustion report compiling method comprises the following steps of 1, data collection, wherein new energy station technology exhaustion related documents are collected; 2, OCR recognition: carrying out character recognition through a multi-engine OCR recognition module, and generating a standardized electronic text; 3, semantic analysis and keyword extraction: carrying out semantic analysis on the recognized text by utilizing a new energy AI large model, and extracting core fields related to wind power plant or photovoltaic power station technology exhaustion; introducing an external new energy knowledge base through a retrieval enhancement generation framework, and performing backtracking verification and confidence enhancement on the extracted core field; step 4, report filling: based on the core field, automatically generating a first draft of a technical exhaustion report according to a preset chapter template; and 5, report auditing: performing data consistency verification on the core field in the first draft through Hash matching and a semantic similarity algorithm, and generating an audited technical exhaustion report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of technical report writing, and specifically provides a method for compiling a technical due diligence report for a new energy station based on ORC and a new energy AI big model. Background Art

[0002] Wind and photovoltaic power plants, as exemplars of green development, play a central role in promoting sustainable development. Their operating mechanism is based on converting abundant natural wind and solar energy into clean electricity, a process that has minimal environmental impact and demonstrates a high degree of eco-friendliness. Given the vast reserves of wind and solar energy, particularly the enormous potential of wind energy, countries around the world are increasingly prioritizing their development and utilization. This global trend paves a broad path for future advancements in wind and photovoltaic power generation technologies, signaling that these two renewable energy sources will assume an increasingly important role in the energy system, leading the clean energy revolution. Due diligence is a critical step in the development of new energy projects, assessing project feasibility, compliance, and investment value. Traditional due diligence reports often rely on manual analysis of large volumes of paper or unstructured documents, which is time-consuming, labor-intensive, and prone to omissions. While some document processing systems now include optical character recognition (OCR) capabilities, they still cannot meet the high requirements of the new energy industry for terminology, data analysis, and report standardization. Therefore, a systematic approach that can automatically recognize, understand, and generate professional reports is urgently needed.

[0003] like Figure 1 As shown, the current approach to compiling technical due diligence reports for new energy power plants typically involves extensive manual effort, with each report being compiled in sections. First, the project's technical specifications, design documents, and operating manuals are thoroughly reviewed to ensure compliance with industry standards. Next, on-site equipment inspections are conducted, meticulously examining the physical condition and operational status of key components such as wind turbines, photovoltaic panels, and inverters. Performance testing and data analysis are then used to assess the equipment's actual performance against design expectations. Furthermore, maintenance records are reviewed and operational data analyzed to assess equipment reliability and system efficiency. Furthermore, grid access agreements are reviewed to ensure power output quality and grid compatibility. Safety management procedures are reviewed to ensure safe and compliant plant operations. Technical risks are identified and assessed, and risk mitigation measures are developed. Finally, the results of all technical due diligence are synthesized into a detailed report, providing targeted improvement recommendations and implementation plans to support plant optimization and investment decisions. The entire process aims to comprehensively assess the technical status and potential risks of the new energy power plant, ensuring its efficient, safe, and compliant operation. The technical due diligence report is compiled by integrating these various components. After the first draft of the report is completed, the next step is to review the important information in the report and ensure that the internal auditors of the technical due diligence report review the accuracy of important data.

[0004] However, when compiling a report, the relevant data collected must be searched for the core data supporting the technical due diligence report. The audit report also requires further searching for information in the relevant data. This multiple data search process is time-consuming and requires the report compiler and initial reviewer to review the collected relevant data documents multiple times, and to find the exact location of the specific fields in the documents multiple times. This is time-consuming, and multiple searches may result in inconsistencies in the specific fields referenced in the relevant data documents supporting the report. Therefore, the existing technology has some shortcomings in the preparation of technical due diligence reports for new energy stations, as follows.

[0005] First, because existing technical due diligence reports for new energy stations are compiled manually and organized into chapters, compiling the reports requires searching and extracting key data from a large amount of construction unit information, supervision materials, design materials, production and operation materials, and compliance documents. This process is time-consuming and labor-intensive. During the review phase, auditors need to repeatedly review original materials to verify the accuracy of the data in the reports, further increasing the workload and time costs.

[0006] Secondly, there is the issue of information consistency. Multiple data queries may lead to inconsistencies in the supporting information cited in the report. Especially when information is exchanged between different people, there may be misunderstandings or omissions, which will affect the accuracy and credibility of the report.

[0007] Again, there is the problem of difficulty in data tracking. When the data mentioned in the report needs to be traced back to the specific location of the original document, it is very difficult to find it without an effective indexing or marking system, which affects the transparency and verifiability of the report. Summary of the Invention

[0008] In view of this, the purpose of the present invention is to provide a method for compiling a technical due diligence report for a new energy station based on ORC and a new energy AI big model, which has the advantages of low manpower investment, good data consistency and data traceability, and can improve the speed of writing technical due diligence reports and user experience.

[0009] In order to achieve the above object, the present invention provides the following technical solutions: A method for compiling technical due diligence reports for new energy stations based on ORC and a large new energy AI model uses LangChain to build a task chain, sequentially automating the task scheduling and scheduling of OCR recognition, semantic analysis, keyword extraction, and report filling modules. The method includes the following steps: Step 1: Data Collection: Collect documents related to the technical due diligence of new energy stations, including construction unit information, supervision information, design phase information, operation information, compliance information, and other information; Step 2: OCR Recognition: Use a multi-engine OCR recognition module to perform text recognition and layout structure restoration on the relevant documents of the new energy station technical due diligence to generate standardized electronic text; Step 3: Semantic Analysis and Keyword Extraction: Utilize the new energy AI big model to perform semantic analysis on the recognized text and extract core fields related to technical due diligence of wind farms or photovoltaic power stations, including wind speed time series data, wind turbine layout optimization parameters, annual average radiation, and inclined surface radiation. An external new energy knowledge base is introduced through a search-enhanced generation architecture to perform retrospective verification and confidence enhancement on the extracted core fields. Step 4: Report filling: Based on the core fields, automatically generate the draft of the technical due diligence report according to the preset chapter template; Step 5: Report review: Use hash matching and semantic similarity algorithms to verify the data consistency of the core fields in the draft, and generate a reviewed technical due diligence report.

[0010] Furthermore, in step 2, the multi-engine OCR recognition module uses a character recognition algorithm based on the CRNN architecture, combined with a preprocessing algorithm to perform denoising and text segmentation on the image.

[0011] The CRNN architecture is based on a convolutional neural network and a recurrent neural network in combination with a CTC mechanism, including: the convolutional neural network learns the high-level semantic representation of the input image and extracts a feature map with rich spatial information; the feature map is input into the recurrent neural network to model the temporal dependency between features, thereby capturing the contextual information in the text sequence; the CTC mechanism is introduced in the transcription stage to map the feature sequence learned by the model into the final character sequence output.

[0012] Furthermore, the preprocessing algorithm includes image denoising and text segmentation technology to improve character recognition efficiency.

[0013] Furthermore, in step three, the new energy AI big model is a customized big language model trained based on knowledge in the wind power and photovoltaic fields. The fields are classified and archived through the BERT and Transformer models combined with semantic analysis technology, and the context semantic model is used for context association to handle multiple citation problems.

[0014] Furthermore, in step three, the new energy AI big model uses the general big language model as the basic model, and expands or reconstructs the big language model structure according to the application scenario characteristics of the new energy industry, including: introducing a time series processing module to support tasks including power load forecasting and wind power forecasting; adding a graph neural network module to model the grid structure and equipment topology relationship; and integrating a multimodal processing module to support image recognition and voice interaction.

[0015] Furthermore, in step three, the external new energy knowledge base includes a project case library, a standard parameter library and an industry specification library, and the retrieval enhancement generation architecture is used to compare and verify the recognized text with the external industry knowledge base to achieve accuracy enhancement and extraction calibration of core fields.

[0016] Furthermore, in step five, the data consistency verification includes: performing hash matching on the extracted core fields and the design values ​​stored in the database, and detecting abnormal data using a semantic similarity algorithm.

[0017] Furthermore, the LangChain task chain adopts a directed acyclic graph structure, which sequentially calls the document segmentation module for OCR recognition, the embedding generation module for semantic analysis, the vector retrieval module for keyword extraction, and the report filling module to generate a technical due diligence report.

[0018] The beneficial effects of the present invention are: In order to solve the problems of repeated data search and data consistency encountered in the preparation and review of the technical due diligence report of new energy stations, the present invention is based on the method for preparing the technical due diligence report of new energy stations based on ORC and the new energy AI big model. Through the data collection step, the technical due diligence related documents of new energy stations including construction unit information, supervision information, design stage information, operation information, compliance information and other types of information are collected, and a centralized database containing all technical specifications, design documents, operation manuals, maintenance records, performance test data, safety management procedures, etc. is constructed; the technical due diligence related documents of new energy stations are converted into standard ization of electronic texts, so that the database should have a powerful search function, support keyword search, full-text retrieval and metadata tag classification, so as to quickly locate the required information; use the new energy AI big model to perform semantic analysis on the recognized text, automatically extract key data from the document, and compare it with other data in the external new energy knowledge base to ensure consistency, and generate a difference report to display inconsistent data for quick correction; when generating the first draft of the technical due diligence report, the core fields in the first draft are verified for data consistency through hash matching and semantic similarity algorithms, which can automatically detect logical errors, data missing and format problems in the report, and improve the review efficiency. In summary, the method for compiling a technical due diligence report for a new energy station based on ORC and the new energy AI big model of the present invention will be more efficient and accurate in the preparation and review process of the technical due diligence report, while also reducing labor costs and error probabilities, improving the professionalism and reliability of the report, and will significantly enhance the comprehensive benefits of the technical due diligence of new energy stations, laying a solid foundation for the long-term success of the project. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to make the purpose, technical solutions and beneficial effects of the present invention more clear, the present invention provides the following drawings for illustration: Figure 1 A flowchart of the method for compiling technical due diligence reports for existing new energy stations; Figure 2 The figure is a flowchart of the method for compiling a technical due diligence report on a new energy station based on ORC and the new energy AI big model of the present invention. DETAILED DESCRIPTION

[0020] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.

[0021] like Figure 2 As shown, this embodiment is based on the method for compiling a technical due diligence report for new energy stations using ORC and the new energy AI big model. LangChain is used to build a task chain, and the OCR recognition, semantic analysis, keyword extraction, and report filling modules are automatically scheduled and arranged, including the following steps.

[0022] Step 1: Data collection: Collect documents related to the technical due diligence of new energy stations. The categories of documents related to the technical due diligence of new energy stations include construction unit information, supervision information, design stage information, operation information, compliance information and other information. Specifically: (1) Construction unit information includes: Project establishment documents: including project approval documents, feasibility study report, preliminary design approval, etc., to ensure the legality and feasibility of the project.

[0023] Contracts and Agreements: Review all contracts and agreements signed with contractors, suppliers, and power grid companies, including EPC general contracting contracts, equipment procurement contracts, grid connection agreements, etc.

[0024] Construction process documents: construction logs, change orders, meeting minutes, project acceptance records, etc., reflecting the detailed progress and decision-making process during project construction.

[0025] (2) Supervision materials include: Supervision planning and implementation details: clarify the scope, content, objectives and methods of supervision work.

[0026] Quality control documents: including raw material inspection reports, hidden project acceptance records, equipment commissioning reports, etc., to ensure project quality.

[0027] Progress control information: construction schedule, progress deviation analysis, progress adjustment plan, etc., to monitor project progress.

[0028] Cost control information: budget, final accounts, change visas, payment applications, etc., to control project costs.

[0029] Safety and environmental supervision materials: production safety inspection records, implementation of environmental protection measures, etc., to ensure that the project meets safety and environmental protection requirements.

[0030] (3) Design stage information includes: General layout drawing, electrical wiring diagram, civil structure drawing, equipment selection instructions, etc., to confirm the rationality of the design.

[0031] Design Change Notification: Track the approval process and implementation of design changes to ensure that changes comply with specifications and requirements.

[0032] Performance calculation and analysis reports: such as wind resource assessment reports, photovoltaic radiation calculations, system efficiency analysis, etc., to evaluate the effectiveness of the design.

[0033] (4) Operational information includes: Design drawings and instructions: Equipment operation and maintenance manual: guide daily operation and maintenance to ensure the normal operation of the equipment.

[0034] Operation logs and data records: including power generation, fault records, maintenance activities, etc., for performance analysis and fault diagnosis.

[0035] Regular inspection and test reports: such as annual performance tests, safety inspection reports, etc., to ensure long-term stable operation of the equipment.

[0036] Energy management and efficiency reporting: Analyze energy consumption, system efficiency, and energy conservation and emission reduction effects.

[0037] (5) Compliance information includes: Environmental impact assessment report and approval: Confirm that the project meets environmental protection requirements and obtain environmental impact assessment approval.

[0038] Safety assessment report and approval: prove that the project safety measures meet the standards and obtain safety assessment approval.

[0039] Land use and planning permits: Ensure that the project land is legal and complies with urban planning.

[0040] Grid connection and dispatch agreement: A grid connection agreement signed with the grid company to ensure that power transmission meets the grid requirements.

[0041] Step 2: OCR recognition: Use the multi-engine OCR recognition module to perform text recognition and layout structure restoration on documents related to the technical due diligence of new energy stations to generate standardized electronic text.

[0042] In this embodiment, the multi-engine OCR recognition module uses a character recognition algorithm based on the CRNN architecture, combined with a preprocessing algorithm to perform image denoising and text segmentation. Specifically, the CRNN architecture is based on convolutional neural networks and recurrent neural networks, combined with a CTC mechanism. The CRNN-based character recognition algorithm is particularly suitable for character recognition tasks in complex document formats, such as wind resource assessment reports, equipment operation logs, and other specialized documents in the new energy industry.

[0043] Specifically, the CRNN (Convolutional Recurrent Neural Network) architecture is based on a combination of convolutional neural networks and recurrent neural networks, coupled with a connectionist temporal classification (CTC) mechanism. The CRNN learns high-level semantic representations of the input image and extracts feature maps rich in spatial information. The feature maps are then fed into the recurrent neural network to model the temporal dependencies between features, thereby capturing contextual information within the text sequence. Finally, the CTC mechanism is introduced during the transcription phase to map the learned feature sequence into the final character sequence output. The CTC (Connectionist Temporal Classification) mechanism eliminates the need for precise character segmentation in the input image, significantly improving recognition efficiency and accuracy.

[0044] In the recognition of documents with complex formats, CRNN demonstrates excellent applicability and technical advantages. To address the multi-scale character issues that may appear in documents, convolutional neural networks can effectively extract discriminative features for characters of different sizes, thereby ensuring consistent recognition results. For consecutive lines of text, there is a clear sequential relationship between characters, and the RNN structure can model this sequence information, further improving overall recognition accuracy. Because actual documents often suffer from problems such as character adhesion, blurriness, or poor printing quality, traditional methods struggle to effectively segment text. CRNN, combined with the CTC mechanism, can directly recognize entire paragraphs or lines of text, avoiding the tedious pre-segmentation process and improving the robustness and practicality of the system. This architecture has a strong tolerance for interference factors such as rotation, scaling, and deformation in images, making it more suitable for handling text recognition tasks in various complex layouts and low-quality scanned images in real scenarios. Therefore, the CRNN-based OCR recognition module has significant technical advantages in dealing with documents with complex formats.

[0045] In this embodiment, the preprocessing algorithm includes image denoising and text segmentation techniques to improve character recognition efficiency. For example, when processing old or blurry scans, image processing techniques are used to remove noise and ensure clear and readable text. Through machine learning, recognition accuracy can be gradually improved.

[0046] Step 3: Semantic analysis and keyword extraction: Use the new energy AI big model to perform semantic analysis on the recognized text and extract the core fields related to the technical due diligence of wind farms or photovoltaic power stations, including wind speed time series data, wind turbine layout optimization parameters, annual average radiation and inclined surface radiation. An external new energy knowledge base is introduced through the retrieval enhancement generation architecture to perform retrospective verification and confidence improvement on the extracted core fields. In this embodiment, the external new energy knowledge base includes a project case library, a standard parameter library, and an industry specification library. The retrieval enhancement generation architecture is used to compare and verify the recognized text with the external industry knowledge base to achieve precision enhancement and extraction calibration of the core fields.

[0047] In this embodiment, the new energy AI big model is a customized large language model trained based on knowledge in the wind power and photovoltaic fields. It uses BERT and Transformer models combined with semantic analysis technology to classify and archive fields, and uses a contextual semantic model to perform context association to handle multiple citations. The new energy AI big model can accurately extract key information (such as equipment model, operating parameters, maintenance records, etc.) to ensure that the extracted information matches the chapter requirements. In the semantic analysis driven by the new energy AI big model, the recognized text is input into the large language model (LLM) trained based on knowledge in the wind power and photovoltaic fields to achieve semantic understanding tasks such as industry terminology parsing, parameter understanding, and causal relationship judgment.

[0048] In this embodiment, the new energy AI big model uses a general big language model as its foundational model, leveraging its powerful contextual modeling capabilities and multi-task learning mechanism as the core of its language understanding and generation. Based on this foundation, the big language model structure is expanded or restructured according to the application scenarios of the new energy industry. This includes: introducing a time series processing module to support tasks such as power load forecasting and wind power forecasting; adding a graph neural network module to model power grid structure and device topology; and integrating a multimodal processing module to support image recognition and voice interaction.

[0049] Building a large new energy AI model based on an existing large language model is an innovative approach to deeply integrate general artificial intelligence technology with the new energy industry. This method fully utilizes the capabilities of existing large-scale language models in natural language understanding, knowledge representation, and reasoning generation, and combines the data resources, business logic, and technical requirements unique to the new energy field. Through model fine-tuning, knowledge injection, task adaptation, and system integration, a large vertical model for the new energy industry is constructed, thereby realizing intelligent support for key links such as design, operation, maintenance, prediction, and optimization in the new energy industry chain. In terms of data preparation and knowledge injection, this embodiment emphasizes the construction of a high-quality new energy industry corpus and professional knowledge map. By processing, annotating, and structuring technical documents related to new energy, a training corpus with industry characteristics is formed, which is used for continuous pre-training or instruction fine-tuning of the basic model. During the model training and optimization process, a multi-stage training strategy is adopted, including domain-adaptive pre-training, task-oriented fine-tuning, and reinforcement learning optimization.

[0050] In this embodiment, a contextual semantic model is constructed to handle the problem of multiple references of the same field in multiple documents, thereby reducing the risk of information inconsistency. This method is particularly suitable for document types that require long-term tracking and comparative analysis, such as equipment operation logs and wind resource assessment reports.

[0051] Specifically, in actual application scenarios, especially in areas involving multi-document processing (such as new energy project management, legal document archiving, technical specification integration, etc.), the same field information may be cited multiple times in multiple documents, but their expression methods, semantic orientations or contextual meanings may be different, resulting in inconsistent or even contradictory information. In order to effectively solve this problem, the semantic model with context understanding capabilities constructed in this embodiment becomes the key. Through deep learning and natural language processing technology, this model can accurately identify the semantic background of fields in different documents and achieve consistent parsing and unified expression. The method for constructing a contextual semantic model mainly includes the following core steps: data collection and annotation, context modeling architecture design, multi-document semantic alignment mechanism construction, training strategy formulation, and deployment optimization application.

[0052] In this embodiment, the RAG (Retrieval-Augmented Generation) architecture enables keyword calibration and knowledge enhancement. The specific process is as follows: When a user enters a query or text to be processed, the retriever first encodes it into a vector representation and searches the knowledge base for semantically similar document fragments or knowledge items. This retrieved information is then fed into the generator along with the original input, serving as additional context in the final text generation process. The recognized text is then compared and verified against an external industry knowledge base, achieving precision enhancement and extraction calibration for core fields such as wind speed data, light intensity, installed capacity, and economic indicators.

[0053] Specifically, RAG (Retrieval-Augmented Generation) is a deep learning architecture that combines information retrieval and text generation techniques. It aims to improve the accuracy and interpretability of natural language processing models in open-domain question answering, knowledge-intensive tasks, and specialized applications. By organically integrating an external knowledge base with a generative model, this architecture enables the model to dynamically retrieve and utilize relevant background information during text generation, thereby avoiding the knowledge limitations and illusions inherent in relying solely on internal training data. The RAG architecture effectively implements keyword calibration and knowledge augmentation in scenarios such as post-OCR text understanding, new energy industry document analysis, and field consistency verification. The RAG architecture consists of two core components: a retriever and a generator. The retriever is responsible for rapidly locating contextual information relevant to the current input from a large-scale external knowledge base. The generator, based on the input text and retrieved knowledge content, jointly generates semantically accurate and logically coherent output.

[0054] The specific implementation methods of semantic analysis and keyword extraction are further explained below with reference to specific examples.

[0055] 1. Taking the “wind resource assessment report” in a wind farm project as an example, the specific implementation methods of semantic analysis and keyword extraction are further explained.

[0056] (1) Identify the extraction and field examples as follows: Wind speed time series data: Hourly wind speed data, used to assess the long-term stability of wind resources.

[0057] Wind turbine layout optimization parameters: including the spacing between wind turbines, wind energy utilization coefficient, etc., to ensure a reasonable layout design.

[0058] (2) Reasons for extracting the example fields: Wind speed time series data: used to verify the deviation between performance calculations in the design phase and actual operation, providing data support for investors.

[0059] Layout optimization parameters: ensure the scientific nature of technical due diligence and the accuracy of conclusions.

[0060] 2. Taking the “Photovoltaic Radiation Calculation Report” in a photovoltaic power station project as an example, the specific implementation methods of semantic analysis and keyword extraction are further explained.

[0061] (1) Identify the extraction and field examples as follows: Average annual radiation: calculated based on historical meteorological data of the project location, used to assess the power generation potential of photovoltaic power stations.

[0062] Inclined surface radiation: The effective radiation calculated based on the installation angle of the photovoltaic panel directly affects the design efficiency of the power station.

[0063] (2) Reasons for extracting the example fields: Average annual radiation: provides basic data for the selection and capacity design of photovoltaic power generation systems in the design phase.

[0064] Inclined surface radiation: ensures that the design phase and actual power generation efficiency are consistent, reducing investment risks.

[0065] Through OCR technology, these fields are efficiently identified and extracted during the report preparation stage and compared with the design values ​​in the database to ensure data consistency and traceability. At the same time, it significantly reduces labor costs and error rates, providing strong technical support for technical evaluation and investment decisions in the new energy industry.

[0066] Step 4: Report filling: Based on the core fields, the first draft of the technical due diligence report is automatically generated according to the preset chapter template.

[0067] Step 5: Report Review: Use hash matching and semantic similarity algorithms to verify the data consistency of the core fields in the draft, and generate a reviewed technical due diligence report.

[0068] In this embodiment, data consistency verification includes: hash matching the extracted core fields with the design values ​​stored in the database, and using a semantic similarity algorithm to detect abnormal data, ensuring the consistency and accuracy of the data and improving the accuracy and credibility of the due diligence report. When reviewing the first draft, the internal auditor uses ORC technology to identify and query the core fields of the collected new energy station technical due diligence related documents according to different categories. After identifying the core fields, the location of the corresponding fields in the file is recorded. The auditor can quickly review whether the identified field information is correctly displayed in the technical due diligence report, and quickly complete the review to form a reviewed technical due diligence report.

[0069] Specifically, this embodiment establishes a real-time monitoring system based on hash matching and semantic similarity algorithms to automatically detect abnormal data and issue alarms to ensure stable operation of the system.

[0070] In this embodiment, the LangChain task chain utilizes a directed acyclic graph (DAG) structure, sequentially invoking the document segmentation module for OCR recognition, the embedding generation module for semantic analysis, the vector retrieval module for keyword extraction, and the report filling module to generate the technical due diligence report. Specifically, given the diverse types of new energy sites (e.g., wind power, photovoltaics, energy storage), and the varying project phases (e.g., project approval, commissioning, and technical upgrades), the task chain should be highly flexible, allowing users to freely combine task nodes, adjust execution order, and replace submodules based on different scenarios. Furthermore, the system should provide a complete execution log and intermediate result output, enabling technicians to easily trace the basis and logic of each processing step, enhancing system transparency and credibility.

[0071] The above embodiments are merely preferred embodiments for the purpose of fully illustrating the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are within the scope of protection of the present invention. The scope of protection of the present invention shall be subject to the claims.

Claims

1. A method for compiling a technical due diligence report for a new energy station based on ORC and a new energy AI big model, characterized by: LangChain is used to build a task chain, which automatically arranges and schedules the OCR recognition, semantic analysis, keyword extraction, and report filling modules in sequence. The steps include: Step 1: Data Collection: Collect documents related to the technical due diligence of new energy stations, including construction unit information, supervision information, design phase information, operation information, compliance information, and other information; Step 2: OCR Recognition: Use a multi-engine OCR recognition module to perform text recognition and layout structure restoration on the relevant documents of the new energy station technical due diligence to generate standardized electronic text; Step 3: Semantic Analysis and Keyword Extraction: Utilize the new energy AI big model to perform semantic analysis on the recognized text and extract core fields related to technical due diligence of wind farms or photovoltaic power stations, including wind speed time series data, wind turbine layout optimization parameters, annual average radiation, and inclined surface radiation. An external new energy knowledge base is introduced through a search-enhanced generation architecture to perform retrospective verification and confidence enhancement on the extracted core fields. Step 4: Report filling: Based on the core fields, automatically generate the draft of the technical due diligence report according to the preset chapter template; Step 5: Report review: Use hash matching and semantic similarity algorithms to verify the data consistency of the core fields in the draft, and generate a reviewed technical due diligence report.

2. The method for compiling a technical due diligence report for a new energy station based on ORC and a new energy AI big model according to claim 1 is characterized by: In the step 2, the multi-engine OCR recognition module adopts a character recognition algorithm based on the CRNN architecture and combines it with a preprocessing algorithm to perform denoising and text segmentation on the image.

3. The method for compiling a technical due diligence report for a new energy station based on ORC and a new energy AI big model according to claim 3 is characterized by: The CRNN architecture is based on a convolutional neural network and a recurrent neural network in combination with a CTC mechanism, including: the convolutional neural network learns the high-level semantic representation of the input image and extracts a feature map with rich spatial information; the feature map is input into the recurrent neural network to model the temporal dependency between features, thereby capturing the contextual information in the text sequence; the CTC mechanism is introduced in the transcription stage to map the feature sequence learned by the model into the final character sequence output.

4. The method for compiling a technical due diligence report for a new energy station based on ORC and a new energy AI big model according to claim 3 is characterized by: The pre-processing algorithm includes image denoising and text segmentation technology to improve the efficiency of character recognition.

5. The method for compiling a technical due diligence report for a new energy station based on ORC and a new energy AI big model according to claim 1 is characterized by: In step three, the new energy AI big model is a customized big language model trained based on knowledge in the wind power and photovoltaic fields. It classifies and archives fields through the BERT and Transformer models combined with semantic analysis technology, and performs context association through the context semantic model to handle multiple citation problems.

6. The method for compiling a technical due diligence report for a new energy station based on ORC and a new energy AI big model according to claim 5 is characterized by: In step three, the new energy AI big model uses the general big language model as the basic model, and expands or reconstructs the big language model structure according to the application scenario characteristics of the new energy industry, including: introducing a time series processing module to support tasks including power load forecasting and wind power forecasting; adding a graph neural network module to model the grid structure and equipment topology relationship; and integrating a multimodal processing module to support image recognition and voice interaction.

7. The method for compiling a technical due diligence report for a new energy station based on ORC and a new energy AI big model according to claim 1 is characterized by: In step three, the external new energy knowledge base includes a project case library, a standard parameter library and an industry specification library. The retrieval enhancement generation architecture is used to compare and verify the recognized text with the external industry knowledge base to achieve accuracy enhancement and extraction calibration of core fields.

8. The method for compiling a technical due diligence report for a new energy station based on ORC and a new energy AI big model according to claim 1 is characterized by: In the step 5, the data consistency verification includes: performing hash matching on the extracted core fields and the design values ​​stored in the database, and detecting abnormal data using a semantic similarity algorithm.

9. The method for compiling a technical due diligence report for a new energy station based on ORC and a new energy AI big model according to claim 1 is characterized by: The LangChain task chain adopts a directed acyclic graph structure, which sequentially calls the document segmentation module for OCR recognition, the embedding generation module for semantic analysis, the vector retrieval module for keyword extraction, and the report filling module to generate a technical due diligence report.

Citation Information

Cited By

  • File identification processing system based on artificial intelligence model and RAG

    CN120894793A