New energy equipment health management intelligent question-answering system based on large model and knowledge graph
The intelligent question-and-answer system for health management of new energy equipment, which utilizes large models and knowledge graphs, solves the challenges of data integration and fault diagnosis in the health management of new energy equipment. It achieves efficient integration of multi-source data and intelligent fault diagnosis, thereby improving the level of intelligence in equipment management and operational efficiency.
Patent Information
- Application Number
- CN202511042582.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-18
AI Technical Summary
Health management of new energy equipment faces data processing challenges. Traditional methods are difficult to effectively integrate multi-source heterogeneous data, resulting in low efficiency in fault diagnosis and knowledge acquisition, failing to meet real-time and intelligent requirements, and existing systems have poor adaptability and scalability, making it impossible to accurately identify complex faults.
The intelligent question-and-answer system for health management of new energy equipment, which adopts large-scale models and knowledge graphs, includes modules for data acquisition and preprocessing, deep fusion of multi-source heterogeneous data, knowledge graph construction, large-scale model training, question-and-answer interaction, fault prediction and health assessment, visualization interaction and decision support, and system optimization, to achieve efficient integration and analysis of multi-source data and accurate fault diagnosis.
It realizes intelligent health management of new energy equipment, can quickly and accurately provide fault diagnosis and prediction, improve operation and maintenance efficiency, reduce operation and maintenance costs, adapt to changes in equipment operating status, provide personalized services, and improve equipment reliability and safety.
Smart Images

Figure CN120973894A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of new energy technology, specifically a new energy equipment health management intelligent question and answer system based on large models and knowledge graphs. BACKGROUND
[0002] With the continuous growth of global demand for clean energy, the new energy industry is showing a rapid development trend. The widespread application of new energy equipment is of great significance for reducing carbon emissions and alleviating the energy crisis, but it also brings great challenges to the health management of equipment.
[0003] New energy equipment is complex in structure and contains numerous precision components. For example, a wind turbine is composed of multiple key components such as blades, hubs, gearboxes, and generators, and the operating state of each component affects the overall performance. At the same time, the data generated during equipment operation is diverse, including real-time data collected by sensors (such as temperature, vibration, and rotation speed), historical data recorded by monitoring systems, technical data in equipment maintenance manuals, and meteorological data from external environments. These data formats are diverse, including structured data and a large amount of unstructured data, making it difficult for traditional data processing methods to effectively integrate and analyze them.
[0004] New energy equipment often operates in harsh environments, such as offshore wind power facing high salt fog and strong winds, and photovoltaic power stations distributed in complex lighting areas, which increases the risk of equipment failure. Equipment failure modes are complex and diverse, and different failures may be related and influence each other. Currently, fault diagnosis mainly relies on human experience and simple threshold judgment, making it difficult to accurately identify early signs of failure and effectively predict potential failures, resulting in prolonged equipment downtime and increased maintenance costs.
[0005] The knowledge system in the new energy field is vast and constantly updated, and equipment maintenance personnel have difficulty quickly obtaining accurate and comprehensive solutions when encountering problems. Traditional document-based knowledge storage and retrieval methods are inefficient and cannot meet the real-time and intelligent needs. In addition, there is a lack of effective correlation between knowledge from different sources, making it difficult to form a systematic knowledge network, limiting the application effect of knowledge in equipment health management.
[0006] In view of the above problems, the prior art has various forms of solutions, such as a system based on sensor monitoring and threshold alarm. This system monitors the operating parameters of the equipment in real time by installing sensors at key parts of the equipment, and sets corresponding threshold values for alarm. When the monitored parameters exceed the threshold values, the system will issue an alarm to indicate that the equipment may have an abnormality. Although this system can find simple abnormal conditions of the equipment to some extent, its limitations are also obvious. It can only detect whether the parameters exceed the threshold values, and cannot perform in-depth analysis and diagnosis on complex faults. Moreover, the setting of threshold values often lacks flexibility and is difficult to adapt to changes in different equipment under different operating conditions, which may result in false alarms or missed alarms.
[0007] There is also a fault diagnosis method based on an expert system. This method uses expert experience to build a rule base and matches the equipment fault phenomena with the rules in the rule base to diagnose faults. However, the acquisition and updating of expert experience is a difficult problem. With the continuous development of new energy technology, new fault modes and problems are constantly emerging, and the existing rule base is difficult to cover all fault scenarios. In addition, this system has poor adaptability and scalability, and cannot make accurate diagnoses once it encounters a situation not covered by the rule base.
[0008] In addition, there is also a traditional question and answer system. When dealing with new energy equipment health management problems, it is mostly based on keyword matching technology. When the user raises a question, the system searches the pre-stored text library for content containing relevant keywords and returns it as an answer to the user. This method cannot understand the semantics and intent of the user's question, and the answer content often lacks pertinence and depth. It is difficult to provide effective answers to complex problems and cannot meet the needs of operation and maintenance personnel for professional knowledge. SUMMARY
[0009] The technical problem to be solved by the present application is to provide a new energy equipment health management intelligent question and answer system based on a large model and a knowledge graph to overcome the deficiencies of the prior art in new energy equipment data processing, fault diagnosis and prediction, knowledge acquisition and application. The system aims to realize efficient integration and analysis of multi-source heterogeneous data, accurate fault diagnosis and prediction, and fast and accurate knowledge retrieval and question and answer, to provide timely and effective technical support for new energy equipment operation and maintenance personnel, improve the health management level of new energy equipment, reduce operation and maintenance costs, and promote the intelligent development of the new energy industry.
[0010] To solve the above technical problems, embodiments of the present application provide the following technical solutions: a new energy equipment health management intelligent question and answer system based on a large model and a knowledge graph, comprising a data acquisition and preprocessing module, a multi-source heterogeneous data deep fusion module, a knowledge graph construction module, a large model training module, a question and answer interaction module, a fault prediction and health assessment module, a visual interaction and decision support module, and a system optimization module.
[0011] The data acquisition and preprocessing module is configured to acquire multi-source data of the new energy equipment and perform preprocessing such as cleaning and denoising.
[0012] The multi-source heterogeneous data deep fusion module is configured to perform standardization, feature fusion and semantic alignment on the preprocessed data.
[0013] The knowledge graph construction module is configured to extract entities and relationships from data and documents, and construct and update a knowledge graph.
[0014] The large model training module is configured to fine-tune a pre-trained language model using labeled data.
[0015] The question and answer interaction module is configured to understand user questions, retrieve knowledge in the knowledge graph and reason to generate answers and feed back to the user.
[0016] The fault prediction and health assessment module is configured to extract device operation data features, construct a prediction model and assess the health status of the device.
[0017] The visualization interaction and decision support module is configured to visually display data, support user interaction and generate decision recommendation reports.
[0018] The system optimization module is configured to optimize the system based on user feedback and system performance monitoring results.
[0019] Further, the multi-source data acquired by the data acquisition and preprocessing module includes sensor data, monitoring system data, weather data, and data in device maintenance manuals and technical documents. During data cleaning, the 3σ principle is used to remove outliers, and linear interpolation is used to fill in missing data.
[0020] Further, the multi-source heterogeneous data deep fusion module includes a data standardization submodule, a feature fusion submodule, and a semantic alignment submodule.
[0021] The data standardization submodule performs format unification and dimension conversion on different types of data.
[0022] The feature fusion submodule uses principal component analysis and deep autoencoders for feature extraction and fusion.
[0023] The semantic alignment submodule uses knowledge graph entity mapping technology to align related entities in different data sources.
[0024] Further, the knowledge graph construction module includes a knowledge extraction submodule, a knowledge fusion submodule, and a knowledge update and expansion submodule.
[0025] The knowledge extraction submodule identifies entities and relationships from preprocessed data and documents;
[0026] The knowledge fusion submodule integrates the extracted knowledge, eliminating duplicate and conflicting information;
[0027] The knowledge update and expansion submodule regularly updates the knowledge graph according to new data and technological developments.
[0028] Further, the large model training module collects and cleanses text data related to new energy equipment and performs labeling, selects a suitable pre-trained language model for fine-tuning, and during the fine-tuning process, the entity and relationship information in the knowledge graph is integrated into the training in the form of embedding vectors. Through the model evaluation and optimization submodule, the evaluation indicators of accuracy, recall rate, F1 value, and mean absolute error are used to evaluate and optimize the model.
[0029] Further, the question and answer interaction module includes a question understanding submodule, a knowledge retrieval and reasoning submodule, an answer generation and feedback submodule:
[0030] The question understanding submodule performs word segmentation, part-of-speech tagging, syntactic analysis, semantic understanding, and disambiguation on user questions;
[0031] The knowledge retrieval and reasoning submodule retrieves related knowledge in the knowledge graph and analyzes and reasons in combination with the reasoning ability of the large model and real-time data;
[0032] The answer generation and feedback submodule presents the reasoning results in natural language form and supports multiple output forms, while recording user questions and answers.
[0033] Further, the fault prediction and health assessment module includes a data feature extraction submodule, a prediction model construction submodule, and a health state assessment submodule, including:
[0034] The data feature extraction submodule uses Fourier transform and wavelet transform methods to extract time domain and frequency domain features from equipment operation data;
[0035] The prediction model construction submodule is based on deep learning algorithms and combines historical fault data and current equipment operation features to train a fault prediction model;
[0036] The health state assessment submodule determines the evaluation index weight using the analytic hierarchy process based on the health standards in the knowledge graph, scores and grades the equipment health state through fuzzy comprehensive evaluation, and sends a warning signal when the score is below the threshold.
[0037] Further, the visual interaction and decision support module includes a visualization display submodule, a user interaction submodule, and a decision suggestion generation submodule, wherein:
[0038] The visualization display submodule presents data in various graphical ways such as geographic information system, not limited to maps, line graphs, and column graphs;
[0039] The user interaction submodule supports users to customize query conditions through the interface, and real-time acquisition of relevant data statistics and analysis results;
[0040] The decision suggestion generation submodule generates decision suggestion reports including but not limited to fault cause analysis, maintenance measures, time arrangement, and resource demand based on large model inference and knowledge graph rules for device fault prediction and health assessment results.
[0041] Further, the system optimization module includes a user feedback analysis submodule, a system performance monitoring submodule, and a knowledge updating and model retraining submodule, wherein:
[0042] The user feedback analysis submodule collects feedback information of users on system answers and decision suggestions, analyzes the problems of the module and the link and feeds back to the relevant module;
[0043] The system performance monitoring submodule monitors the performance indicators of the system in real time, not limited to response time, throughput, and accuracy, and takes corresponding optimization measures when the indicators are abnormal;
[0044] The knowledge updating and model retraining submodule updates the knowledge graph periodically and re-trains the large model.
[0045] Further, the knowledge retrieval and reasoning submodule captures the context semantic information in the user's question in real time, constructs a dynamic reasoning path related to the device running state; the answer generation and feedback submodule generates a multi-dimensional explanatory answer containing the fault evolution path and historical similar case comparison for complex fault scenarios, and supports the presentation of knowledge association in the form of interactive graph node expansion, realizes the deep semantic analysis and logical deduction of user's question.
[0046] Further, the prediction model construction submodule adopts a multi-task learning framework to simultaneously train device fault prediction, residual life prediction, and performance degradation trend prediction tasks, share the bottom feature extraction network, and output prediction results by task; the health state evaluation submodule combines the failure mode knowledge of device components in the knowledge graph, constructs a dynamic evaluation model based on Bayesian network, updates the posterior probability of component health state by real-time input of multi-source monitoring data, and uses reinforcement learning algorithm to adaptively adjust the index weight of analytic hierarchy process, realizes the probabilistic evaluation and dynamic threshold early warning of device health state.
[0047] The beneficial effects of the above technical solutions of the present application are as follows:
[0048] 1、The present application deeply integrates large models with knowledge graph technology, combines the characteristics of new energy equipment health management field, realizes the whole process innovation from data collection, knowledge construction to intelligent question and answer and system optimization. The large model provides strong semantic understanding and reasoning ability, the knowledge graph provides structured domain knowledge support, and the two cooperate with each other, so that the system can process complex natural language problems and provide accurate and comprehensive answers, which is a new application mode in the field of new energy equipment health management.
[0049] 2、The knowledge updating and expanding submodule in the knowledge graph construction module of the present application can track the development of new energy technology and the change of equipment operation state in real time, and update the knowledge in the knowledge graph in time. This dynamic updating mechanism ensures the timeliness and integrity of the knowledge graph, so that the system can adapt to the changing actual demand, and has significant advantages compared with the traditional static knowledge storage mode.
[0050] 3、Through the user feedback analysis submodule and the system performance monitoring submodule, the system can adaptively optimize according to the user's use habit, feedback opinion and system running state. The large model training module re-trains according to new data, and the knowledge graph construction module updates knowledge, so as to provide more personalized and accurate services for users, and improve user experience and practicability of the system.
[0051] 4、Based on big data analysis and machine learning algorithm, the system can predict potential faults of new energy equipment. Through real-time monitoring and analysis of equipment operation data, combined with fault mode and reason knowledge in the knowledge graph, the abnormal state of the equipment can be found in advance, and corresponding maintenance suggestions and decision support can be provided, so as to help users reduce equipment failure rate and improve equipment reliability and safety.
[0052] 5、The multi-source heterogeneous data deep fusion module of the present application adopts advanced data standardization, feature fusion and semantic alignment technology, effectively solves the fusion problem of multi-source heterogeneous data of new energy equipment. It can deeply fuse data from different data sources, different formats and different semantics, extract more valuable information, provide a solid data foundation for subsequent analysis and application, and enhance the system's processing capacity and analysis precision for complex data.
[0053] 6、The visual interaction and decision support module of the application greatly improves the efficiency of users obtaining information and making decisions through intuitive visual display, convenient user interaction design, and scientific decision suggestion generation based on large models and knowledge graphs. Complex knowledge graphs and device data are presented in a graphical manner, making it easy for users to understand and analyze; interactive results are provided in real time according to user-defined queries to meet the diverse needs of users; the generated structured decision recommendation report provides direct and usable decision basis for users, and has important innovative significance in the new energy device management decision-making process. BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 The principle block diagram of the large model and knowledge graph-based new energy device health management intelligent question and answer system of the application;
[0055] Figure 2 The principle block diagram of the data acquisition and preprocessing module of the large model and knowledge graph-based new energy device health management intelligent question and answer system of the application;
[0056] Figure 3 The principle block diagram of the multi-source heterogeneous data deep fusion module of the large model and knowledge graph-based new energy device health management intelligent question and answer system of the application;
[0057] Figure 4 The principle block diagram of the knowledge graph construction module of the large model and knowledge graph-based new energy device health management intelligent question and answer system of the application;
[0058] Figure 5 The principle block diagram of the large model training module of the large model and knowledge graph-based new energy device health management intelligent question and answer system of the application;
[0059] Figure 6 The principle block diagram of the question and answer interaction module of the large model and knowledge graph-based new energy device health management intelligent question and answer system of the application;
[0060] Figure 7 The principle block diagram of the fault prediction and health assessment module of the large model and knowledge graph-based new energy device health management intelligent question and answer system of the application;
[0061] Figure 8 The principle block diagram of the visual interaction and decision support module of the large model and knowledge graph-based new energy device health management intelligent question and answer system of the application;
[0062] Figure 9 The principle block diagram of the system optimization module of the large model and knowledge graph-based new energy device health management intelligent question and answer system of the application. DETAILED DESCRIPTION
[0063] To make the technical problems, technical solutions and advantages to be solved by the present application clearer, specific embodiments will be described in detail below with reference to the drawings.
[0064] Embodiment 1
[0065] As shown in Figure 1 , the present application discloses a new energy equipment health management intelligent question and answer system based on large model and knowledge graph, which mainly consists of eight core modules such as data acquisition and preprocessing module 101, multi-source heterogeneous data deep fusion module 102, knowledge graph construction module 103, large model training module 104, question and answer interaction module 105, fault prediction and health assessment module 106, visualization interaction and decision support module 107 and system optimization module 108. Each module cooperates with each other to realize the functions of the system. The specific architecture is shown in Figure 1 .
[0066] As shown in Figure 2 , the principle diagram of the data acquisition and preprocessing module 101 is shown. This module is responsible for collecting various data generated during the operation of new energy equipment, and cleaning, converting and integrating the data to provide high-quality data support for subsequent modules. It specifically includes the following sub-modules:
[0067] Multi-source data acquisition sub-module 1011: Through various data acquisition interfaces, real-time acquisition of different types of data is realized. For sensor data, industrial Ethernet, 4G / 5G and other communication technologies are used to collect physical quantities such as temperature, pressure and vibration of the equipment at high frequency; for monitoring system data, database connection technology (such as JDBC, ODBC) is used to obtain historical operation records, alarm information and other information of the equipment; for external environmental data (such as weather data), data interface with the weather department or third-party weather data platform is used for data docking. In addition, it also supports extracting key information from unstructured text such as equipment maintenance manual and technical documents, and converting them into structured data using OCR technology and text analysis algorithm.
[0068] Data cleaning and denoising sub-module 1012: The collected data often contains noise, error values and missing values, which affects the data quality. This sub-module uses various data cleaning algorithms, such as abnormal value detection based on statistical methods (such as 3σ principle) to identify and remove data points that deviate significantly from the normal range; uses data interpolation algorithms (such as linear interpolation, spline interpolation) to fill in missing values; uses filtering algorithms (such as Kalman filtering, median filtering) to smooth the noisy data, improving the accuracy and reliability of the data.
[0069] As shown in Figure 3As shown, it is a principle block diagram of the multi-source heterogeneous data deep fusion module 102. The module uses advanced fusion technology to eliminate data dimension differences and semantic ambiguity for multi-source heterogeneous data related to new energy equipment, and provides high-quality fusion data for subsequent analysis and application. It specifically includes the following sub-modules:
[0070] Data standardization submodule 1021: standardize multiple types of data such as Internet of Things sensor data, meteorological data, and power grid data. The specific way of standardization is: for numerical data, use normalization (such as Min-Max normalization, Z-Score normalization) or standardization transformation to make it have a unified scale and distribution; for categorical data, perform encoding processing (such as One-Hot encoding, LabelEncoding). Unified data format and dimension, eliminate data scale difference through normalization and standardization, so that data from different sources are comparable. For example, convert wind speed data (unit: m / s), light intensity data (unit: W / m 2 ) and voltage data (unit: V) to the interval [0, 1] for subsequent fusion calculation.
[0071] Feature fusion submodule 1022: use deep autoencoder (DAE) to fuse data features of different dimensions. DAE can more effectively fuse different types of data features by automatically learning the feature representation of the data to generate high-dimensional comprehensive feature vectors. DAE consists of an encoder and a decoder, and realizes nonlinear feature fusion through multiple layers of neural networks:
[0072] Encoder: compress high-dimensional input features (such as the concatenation vector of sensor features + meteorological features) through multiple hidden layers (such as using ReLU activation function) to map to low-dimensional hidden space (encoding layer) to learn the core feature representation of the data. For example, compress 1000-dimensional original features to 200-dimensional encoding vectors.
[0073] Decoder: starting from the low-dimensional encoding vector, reconstruct the original input features through the multiple-layer network symmetric to the encoder, with the goal of minimizing the reconstruction error.
[0074] The training process is: taking "input features -> encoding -> reconstructing input" as the training target, optimizing the network parameters (weights and biases) through the back propagation algorithm, so that the decoder output is as close to the original input as possible. During the training process, the DAE automatically learns the nonlinear relationship between different features (such as the coupling relationship between vibration features and wind speed), without the need for manual design of fusion rules. After training, the decoder is discarded and only the encoder part is retained. The new multi-source data is input into the encoder, and the low-dimensional encoding vector output by the encoder is the high-dimensional integrated feature vector after fusion. This vector integrates multi-dimensional information such as equipment operating state and environmental impact, and can be directly used as the input of the fault prediction model (such as LSTM, CNN) or health assessment module.
[0075] The semantic alignment submodule 1023: using knowledge graph entity mapping and natural language processing technology, eliminating semantic ambiguity of multi-source data. Align the data representing the same concept in different data sources. First, extract domain entities. For new energy equipment fields (such as photovoltaic and wind power), define regular expressions for entity names, fault types, etc. to quickly extract entities from text. Then use BERT and other NLP models to fine-tune on new energy field corpus (operation and maintenance logs, fault reports) to identify complex entities (such as extracting "photovoltaic module string" from "photovoltaic module string 3 hot spot"). Represent semantics through [CLS] vector or entity word vector. Then normalize entities by calculating the similarity between entity names and standard names using edit distance, such as "photovoltaic panel" and "photovoltaic module" with an edit distance of 2, and setting a threshold (such as ≤3) to determine that they are synonyms. Use Sentence-Transformers to generate semantic vectors for entity names and calculate cosine similarity with standard entity vectors in the knowledge graph. Finally, define standard entities (such as "photovoltaic module" as the root node, associated with "string" and "cell" sub-entities) and relationships (such as "contains" and "fault association"), and record industry standards (such as GB / T photovoltaic terminology) and equipment manual knowledge to form a domain entity library. Compare entity attributes (such as device code format and parameter range). If the core attributes of entities from multiple data sources (such as "rated power 250W" and "string voltage 380V") match, they are considered to be the same concept. Use the hierarchical structure of the knowledge graph (such as "photovoltaic module string" is a sub-entity of "photovoltaic module") to align low-level entities to high-level standard entities through relationship paths.
[0076] As shown in Figure 4 The knowledge graph construction module 103 is a principle block diagram. The knowledge graph is the core knowledge representation method of the system, which describes the concepts, entities, attributes and their relationships of new energy equipment in a structured form, providing knowledge support for intelligent question answering and fault diagnosis. This module includes the following submodules:
[0077] Knowledge extraction sub-module 1031: extract new energy equipment related entities, relationships and attributes from pre-processed data. Use named entity recognition (NER) technology to identify entities such as device name, component name, fault type, maintenance measures, etc.; through relationship extraction algorithm (such as rule-based extraction, machine learning-based extraction), the association between entities is mined, such as "device-component" relationship (wind turbine-gearbox), "fault-cause" relationship (gearbox failure-lubrication deficiency); At the same time, extract the attribute information of the entity, such as the model, rated power of the equipment, material, service life of the component, etc.
[0078] Knowledge fusion sub-module 1032, the specific implementation method is as follows: for multi-source input (such as equipment manual PDF, IoT real-time data, operation and maintenance Excel), use NLP+rule engine to parse knowledge:
[0079] (1) Text knowledge (fault report): BERT extracts entities ("photovoltaic components" "inverter"), relationships ("occurrence" fault), and attributes ("over-temperature 120℃");
[0080] (2) Structured data (sensor CSV): Through Schema mapping (such as "col_01"→"component number") to convert into triples (component A, temperature, 120℃) that can be recognized by the knowledge graph.
[0081] Convert unstructured / semi-structured knowledge into a unified triple format (entity-relation-entity / attribute). Then, eliminate duplicates through redundant knowledge detection, and convert knowledge fragments (such as fault handling processes) into subgraphs through subgraph hash matching, calculate hash values, such as through the hash combination of nodes / relationships, and determine duplicates if the hash is the same; through attribute similarity clustering, the entity attributes (such as "photovoltaic component power") are clustered and detected for repeated distribution, such as the parameters of "450W component" in multi-source knowledge are highly overlapped.
[0082] For repeated subgraphs, keep the most complete knowledge version, such as preferentially keeping knowledge fragments containing "fault repair steps"; supplement the source mark to mark that the knowledge comes from "power station A operation and maintenance log"; for attribute-repeated entities, merge into a unified attribute set, such as "component name" takes the longest description and "power" takes the average of multiple sources. Reorganize the fused knowledge according to the knowledge graph Schema, ensure that only one standard entity is retained for the same device / component, and the numerical / text attributes are consistent after merging / correction.
[0083] Knowledge Update and Expansion Sub-module 1033: Using technologies such as web crawlers, API interfaces, and message queues (such as Kafka), data is collected in real-time or at regular intervals from multiple channels including news websites, social media, industry forums, academic databases, device IoT platforms, and operation and maintenance systems. It can access structured data (such as database and API interface data), semi-structured data (such as HTML web pages, XML / JSON documents), and unstructured data (such as text, images, and voices). For example, through Kafka, millisecond-level data stream processing can be achieved to quickly capture changes in device status data.
[0084] After the data is collected, text cleaning is performed to remove HTML tags, special characters, redundant spaces, and unify the encoding format to UTF-8, etc. A stop word list is customized using a domain dictionary to remove meaningless words such as "de" (的) and "shi" (是). A language model (such as BERT) is used to detect and correct spelling mistakes or grammar problems to provide high-quality text for subsequent knowledge extraction. Named Entity Recognition (NER) and Relation Extraction (RE) technologies are used to extract key information from the text to form triples.
[0085] As Figure 5 shown, it is the principle block diagram of the large model training module 104. The large model undertakes the important tasks of semantic understanding and reasoning in this system. Through learning from a large amount of text data, the large model can understand the semantics of user questions and combine with the knowledge graph for reasoning to generate accurate answers. This module includes the following sub-modules:
[0086] Data Preparation Sub-module 1041: Collect and organize a large amount of text data related to the health management of new energy equipment, including equipment manuals, technical papers, fault cases, operation and maintenance records, etc. Data collection uses technologies such as file parsing, database APIs, and web crawlers, and multi-source data is aggregated through the ETL process; cleaning relies on regular expressions, NLP tools, statistical and knowledge graph methods to filter noise and handle missing values; annotation constructs a label system, assisted by supervised learning tools, text classification algorithms, and few-shot learning with the help of humans, and entity relationship association is achieved using NER and relation extraction; preprocessing transforms the text into vectors through word / sentence embedding and tokenization tools, packages it according to the model requirements, and integrates supervision information into the knowledge. This solution combines the NLP technology stack with the knowledge graph to solve the problems of multi-source data consistency, semanticization, and adaptability, and creates a high-quality labeled dataset for large model training, making the original text into learnable supervision signals.
[0087] Model Selection and Fine-Tuning Submodule 1042: This module selects suitable pre-trained language models (such as the GPT series, BERT, etc.) as the base model and fine-tunes them according to the characteristics and needs of the new energy equipment health management field. First, based on the scenario's requirements for language understanding and knowledge association, it selects pre-trained models such as the GPT series and BERT as the foundation, leveraging their general language knowledge reserves, starting from the model architecture (such as Transformer layers, attention mechanisms) and pre-training tasks (mask prediction, etc.). Next, the processed domain-labeled data is converted according to the model input format and input. Using the backpropagation algorithm, based on the loss between prediction and label, the gradient descent optimizer updates the parameters of each layer of the model (such as multi-head attention weights) to adapt to the domain semantic patterns. Simultaneously, for knowledge graph entities and relationships, the TransE knowledge embedding method is adopted. This involves concatenating knowledge vectors at the input layer or introducing a knowledge attention mechanism in the intermediate layer, integrating structured knowledge into the training, supplementing the domain knowledge logic, and guiding the model to accurately handle tasks such as fault diagnosis and question answering.
[0088] The Model Evaluation and Optimization submodule 1043 uses multiple evaluation metrics (such as accuracy, recall, and F1 score) to evaluate the fine-tuned model and analyze its performance on different types of questions. Based on the evaluation results, the model is optimized, such as adjusting the model structure, hyperparameter settings, and increasing the amount of training data. Furthermore, through manual review and user feedback mechanisms, problems and shortcomings in the model are identified, further optimizing model performance and improving the accuracy and reliability of question answering.
[0089] like Figure 6 The diagram shown is a block diagram of the question-and-answer interaction module 105. This module serves as the interface between the user and the system, responsible for receiving user questions, processing them using the large model and knowledge graph, and providing the answer back to the user. It specifically includes the following sub-modules:
[0090] The Problem Understanding Submodule 1051 utilizes core Natural Language Processing (NLP) technologies. Word segmentation breaks down the problem into lexical units, clarifying basic semantic granularities; part-of-speech tagging assigns grammatical attributes to words (e.g., nouns identifying equipment components), aiding in the identification of key entities; syntactic analysis constructs sentence grammatical structures (e.g., subject-verb-object relationships), clarifying the problem's logic. This is preprocessed and transformed into a computer-understandable structured representation (e.g., token sequences, syntax trees). Simultaneously, it connects to a knowledge graph, matching problem vocabulary with concepts and entities in the graph (e.g., equipment names, fault types), and combining entity relationships to eliminate semantic ambiguity. For example, "battery fault" clarifies whether it refers to an energy storage battery or a photovoltaic battery. Accurately grasping the problem type (fault diagnosis / operation consultation, etc.), involved equipment components, and core intent lays a solid semantic understanding foundation for subsequent question-and-answer interactions, enabling the system to accurately meet user needs. The specific natural language preprocessing techniques are as follows:
[0091] Word segmentation: Adopting word segmentation algorithms based on statistics (such as N-gram model) or deep learning (such as BERT-WordPiece) to cut the user question into lexical units (such as "photovoltaic inverter overload how to do" into "photovoltaic", "inverter", "overload", "how to do"), and make the semantic basic particles clear, laying the foundation for subsequent analysis.
[0092] Part-of-speech tagging: Relying on pre-trained language models (such as LSTM-CRF), combined with domain dictionaries, to tag the words with parts of speech (noun "inverter", verb "overload"), to assist in identifying key entities such as devices and actions, and to clarify the attributes of the question components.
[0093] Syntactic analysis: Using dependency parsing or constituency parsing algorithms (such as StanfordParser) to construct sentence syntax structures (such as "photovoltaic inverter" as the subject and "overload" as the predicate), to comb the logical relationship between words, and to extract the core logic of the question (device abnormal state and appeal).
[0094] Knowledge graph semantic association technology is as follows:
[0095] Entity matching: Through named entity recognition (NER, such as BiLSTM-CRF model) to locate the entities in the question such as devices, components, and faults (identify "photovoltaic inverter" entity from "photovoltaic inverter overload"), and then compare with the knowledge graph entity library, based on string matching and semantic similarity (such as Word2Vec vector distance) to achieve accurate mapping, and associate the entity attributes and relationships in the graph.
[0096] Semantic disambiguation: For polysemous words (such as "battery" can refer to energy storage battery / photovoltaic component battery), combined with the context relationship in the knowledge graph (such as "photovoltaic system" associated with "photovoltaic battery" entity), using graph traversal and subgraph matching algorithms, to filter the semantic interpretation that fits the new energy device health management scenario, eliminate ambiguity, and make the question clear (such as "battery failure" in the photovoltaic scenario).
[0097] Intention recognition: Fusion of preprocessed syntax structure and knowledge graph semantics, based on classification model (such as TextCNN) or rule matching, to judge the question type (fault diagnosis type "how to do", information query type "what is"), to extract the involved device components (photovoltaic inverter) and core appeal (solve overload), and to accurately understand the user's intention, to output clear semantic representation for subsequent question and answer interaction (such as entity: photovoltaic inverter, state: overload, intention: fault solving).
[0098] The Knowledge Retrieval and Reasoning Submodule 1052 is designed to solve health management problems related to new energy equipment. Guided by the results of problem understanding, knowledge retrieval utilizes graph database query languages (such as Cypher) to associate entities, relationships, and attributes in the knowledge graph, coupled with inverted index optimization, to quickly locate answers to simple questions. For complex questions, it integrates symbolic reasoning and large models, breaking down subtasks, constructing hints using knowledge graph logical rules, and deducing from the large model's thought chain. Reinforcement learning is then used to optimize reasoning strategies, improving efficiency and accuracy. Finally, a knowledge fusion algorithm transforms structured knowledge into natural language answers. Through this "retrieval-reasoning-integration" process, it adapts to problems of varying complexity, accurately extracts the value of the knowledge graph, and delivers clear and effective equipment health management solutions to users. It achieves collaboration between the knowledge graph and the large model, efficiently supporting scenarios such as fault diagnosis and maintenance decision-making for new energy equipment.
[0099] The answer generation and feedback submodule 1053: Based on the results of knowledge retrieval and reasoning, and relying on natural language generation technologies (such as the GPT series models) and algorithms such as bundle search, it transforms structured knowledge into natural language. Considering user background and intent, it uses user profiles, rule engines, or large model conditions to generate answers, adjusting the language style and depth to suit different users. Complex questions are output with structured explanations through text segmentation and logical connection; simple questions receive concise responses. Multimodal generation integrates visualization libraries (Matplotlib) and image models (StableDiffusion) to output charts and diagrams, enhancing readability.
[0100] like Figure 7 The diagram shown is a principle block diagram of the fault prediction and health assessment module 106. This module, based on data features and knowledge graph fault patterns, realizes quantitative assessment of equipment health status and fault prediction. Specifically, it includes the following sub-modules:
[0101] Data Feature Extraction Submodule 1061: For sensor data and historical maintenance records of new energy equipment, time-domain and frequency-domain features are extracted using methods such as Fourier transform and wavelet transform to construct equipment feature vectors. For example, Fourier transform is used to convert vibration signals from the time domain to the frequency domain, analyze their frequency components, and extract fault-related feature frequencies; wavelet transform is used to perform multi-scale decomposition of the signal to obtain energy distribution characteristics in different frequency bands, which is used to more accurately identify equipment faults.
[0102] The prediction model construction submodule 1062: Based on deep learning, it trains a fault prediction model using historical equipment fault data and dynamically adjusts the model parameters. It employs deep learning models such as Recurrent Neural Networks (RNNs), Long Short-Term Memory Networks (LSTMs), and Convolutional Neural Networks (CNNs) to capture the temporal and spatial features of equipment operation data, achieving accurate prediction of equipment faults. Simultaneously, through ensemble learning methods, it fuses the prediction results of multiple models to improve the reliability and stability of the prediction.
[0103] Health Status Assessment Submodule 1063: Based on equipment health standards and feature vectors in the knowledge graph, this module uses the Analytic Hierarchy Process (AHP) or fuzzy comprehensive evaluation method to score and classify the health status of key equipment components. It establishes an equipment health assessment system containing multiple evaluation indicators, determines the weights based on the importance of each indicator, and uses fuzzy comprehensive evaluation to comprehensively assess the equipment's health status. The system classifies equipment health levels into different categories such as healthy, sub-healthy, pre-failure, and failure, providing a basis for equipment maintenance decisions.
[0104] like Figure 8 The diagram shown is a block diagram of the visualization interaction and decision support module 107. This module provides an intuitive visualization interface and scientific decision-making suggestions. It specifically includes the following sub-modules:
[0105] The visualization submodule 1071 graphically presents the knowledge graph structure, health assessment results, and fault prediction trends. It uses graph visualization tools (such as Neo4jBrowser and Gephi) to display the knowledge graph, presenting the relationships between entities such as equipment, components, faults, and causes in an intuitive graphical way. It uses charts (such as bar charts, line charts, and pie charts) to display equipment health assessment results, such as the distribution of health scores for each component and the overall health trend of the equipment over time. It uses visualization libraries (such as Echarts and D3.js) to draw fault prediction trend charts, showing the probability of equipment failure over a future period. Furthermore, it can use Geographic Information System (GIS) technology to display the distribution of new energy equipment on a map and mark the health status of the equipment in real time, facilitating macro-level monitoring and management by administrators.
[0106] User Interaction Submodule 1072: This module allows users to customize query conditions through the interface and obtain real-time interactive results such as device cluster health statistics and geographic distribution analysis. It provides rich interactive components, including drop-down menus, radio buttons, checkboxes, and date pickers, allowing users to flexibly filter and query based on various conditions such as device type, time range, geographical location, and health status. For example, a user can choose to query "all wind turbines in a certain region with a health status of 'sub-healthy' in the past month." The system will retrieve and analyze the data in real time based on the user's selection and display the results in intuitive charts or tables. Furthermore, users can zoom, pan, and click on the visual interface for easier viewing and analysis of data details.
[0107] The decision recommendation generation submodule 1073, based on large-scale model reasoning and knowledge graph rules, generates structured decision recommendation reports for maintenance resource allocation, equipment upgrades, or replacements based on equipment failure prediction and health assessment results. When the system detects potential equipment failure risks or a decline in health status, it formulates a reasonable maintenance plan based on the severity of the failure, the importance of the equipment, and the availability of maintenance resources, including maintenance time, personnel arrangements, and a list of required spare parts. For situations requiring equipment upgrades or replacements, it provides users with detailed equipment selection recommendations and return on investment analysis, comprehensively considering factors such as equipment performance, cost, and technological development trends. The decision recommendation report is presented in a structured document format, including a problem description, analysis process, recommended measures, and expected results, helping users quickly understand the situation and make informed decisions.
[0108] like Figure 9 The diagram shown is a principle block diagram of the system optimization module 108. To ensure system performance and service quality and continuously adapt to the changing needs of new energy equipment health management, this module continuously optimizes the system. Specifically, it includes the following sub-modules:
[0109] User Feedback Analysis Submodule 1081: Collects user feedback on system responses, including satisfaction ratings and error correction suggestions. It uses natural language processing (NLP) technology to perform sentiment analysis and semantic understanding on user feedback, determining its positive or negative tendencies and extracting key issues and suggestions. It then correlates user feedback with various system modules to pinpoint the specific module and stage where the problem lies. For example, if a user's feedback is inaccurate, analysis can determine whether the problem stems from a problem in the large model training module or a data error in the knowledge graph construction module. Based on the analysis results, the issue is fed back to the relevant modules for targeted optimization and improvement.
[0110] System performance monitoring submodule 1082: Real-time monitoring of system performance indicators such as response time, throughput, accuracy, etc. By setting performance monitoring points at key nodes and interfaces of the system, performance data during system operation is collected and real-time analysis and visualization is performed. When abnormal system performance indicators are found, such as long response time, decreased throughput, etc., the performance analysis tool (such as flame chart, call stack analysis) is used to analyze the system bottleneck in depth to determine whether it is caused by insufficient hardware resources, low algorithm efficiency or network delay, etc. According to the analysis result, corresponding optimization measures are taken, such as optimizing database query statements, adjusting model parameters, increasing hardware resources, etc., to improve the overall performance of the system.
[0111] Knowledge updating and retraining submodule 1083: With the development of new energy technologies and the accumulation of equipment operation data, the knowledge in the knowledge graph and the training data of the large model need to be updated continuously. This submodule regularly updates the knowledge graph, acquires new equipment data, industry dynamics, research results, etc. through the data acquisition module, uses the knowledge extraction and fusion technology of the knowledge graph construction module to integrate new knowledge into the knowledge graph, and corrects the wrong information. At the same time, new text data is collected to retrain the large model, so that the model can learn the latest knowledge and language expression. During the retraining process, the incremental learning method is used to avoid repeated training of existing data and improve training efficiency.
[0112] In the specific work of the above-mentioned various modules of the present application, first, data acquisition and processing is performed, and multi-source data of new energy equipment is acquired by means of the data acquisition and preprocessing module, and is processed by cleaning, denoising and standardization. Second, data fusion analysis is performed, and the preprocessed data is standardized, fusion features and semantic alignment are performed, and data differences are eliminated by means of the multi-source heterogeneous data deep fusion module. Third, the graph construction and updating is performed, and the graph is constructed by extracting entity relationships from data and documents by means of the knowledge graph construction module. Fourth, the model training and optimization is performed, and the pre-trained model is fine-tuned by means of the large model training module using labeled text and graph information, and the model performance is optimized by means of evaluation indicators; the question and answer interaction module receives user questions, combines the training model and the knowledge graph, and generates answers and feedback through understanding and reasoning; the fault prediction and health assessment module extracts features and evaluates the state based on equipment data and rules, and gives early warning when the score is low; the visualization interaction and decision support module visualizes the results, supports user interaction, and generates decision recommendation reports; the system optimization module collects feedback and monitors performance, optimizes the modules accordingly, and realizes continuous upgrading and improvement of the system.
[0113] Among them, the knowledge retrieval and reasoning submodule constructs a dynamic reasoning path related to the running state of the device by capturing the context semantic information in the user's question in real time; the answer generation and feedback submodule generates a multi-dimensional explanatory answer containing the fault evolution path and historical similar case comparison for complex fault scenarios, and supports the presentation of knowledge association in the form of interactive graph node expansion, realizing the deep semantic analysis and logical deduction of user questions. The prediction model construction submodule adopts a multi-task learning framework to simultaneously train device fault prediction, residual life prediction, and performance degradation trend prediction tasks, sharing the underlying feature extraction network and outputting prediction results by task; the health state assessment submodule combines the failure mode knowledge of device components in the knowledge graph to construct a dynamic assessment model based on Bayesian networks, updates the posterior probability of component health state by real-time input of multi-source monitoring data, and uses reinforcement learning algorithm to adaptively adjust the index weight of analytic hierarchy process, realizing the probabilistic assessment and dynamic threshold warning of device health state.
[0114] Example 2
[0115] Wind turbine fault diagnosis and maintenance scenario
[0116] Data acquisition and preprocessing: In a certain wind farm, through the sensors installed on the key components of the wind turbine (such as blades, gearboxes, generators, etc.), the running data of the device is collected in real time, including vibration data, temperature data, speed data, etc. At the same time, the historical operation records and alarm information in the monitoring system, as well as the local meteorological data (such as wind speed, wind direction, air temperature, etc.) are collected. The collected data is transmitted to the data acquisition and preprocessing module, first for data cleaning to remove outliers and noise points, for example, using the 3σ principle to detect and eliminate abnormal large values in the vibration data. Then, data standardization processing is performed, and temperature, speed, etc. Data is normalized to have a unified scale. Finally, through feature engineering, new features such as the frequency spectrum features of vibration data and the rate of change of temperature are calculated, providing more rich data support for subsequent analysis.
[0117] The collected sensor data, weather data, and monitoring system data are transmitted to a multi-source heterogeneous data deep fusion module. In the data standardization submodule, different types of data are unified in format and dimensionally converted to make them comparable. For example, wind speed data and temperature data are normalized. Then, in the feature fusion submodule, principal component analysis (PCA) is used to reduce the dimensionality of the data, extract the main features, and fuse them with the features learned by the deep autoencoder (DAE), generating a comprehensive feature vector. In the semantic alignment submodule, the knowledge graph entity mapping technology is used to align related entities in different data sources to ensure the consistency of data semantics. For example, the "gearbox temperature" in the sensor data is associated with the "gearbox-temperature" entity in the knowledge graph.
[0118] The knowledge extraction submodule is used to extract entities and relationships from the preprocessed data and equipment maintenance manuals and technical documents. For example, entities such as "wind turbine", "blade", and "gearbox" are identified, as well as relationships such as "blade-connection-hub", "gearbox-contains-gear", and "wind turbine failure-reason-blade icing". The extracted knowledge is fused to eliminate duplicate and conflicting information, and a wind turbine health management knowledge graph is constructed. As the equipment operates and new data is collected, the knowledge update and expansion submodule regularly updates the knowledge graph, such as adding newly discovered fault modes or improved maintenance measures.
[0119] A large amount of text data related to wind turbines is collected, including equipment manuals, fault cases, technical papers, etc., which are cleaned and labeled to generate a training data set. The labeled content includes question classification, answer labeling, and association labeling with knowledge graph entities and relationships. A suitable pre-trained language model (such as BERT) is selected and fine-tuned on the training data set. During fine-tuning, the entity and relationship information in the knowledge graph is incorporated into the model training in the form of embedding vectors to enhance the model's understanding of wind turbine domain knowledge. After training, the model evaluation and optimization submodule is used to evaluate the model using accuracy, recall, and F1 value evaluation indicators, and adjust the model parameters based on the evaluation results to improve the model performance.
[0120] When the maintenance personnel encounter abnormal vibration problems of wind turbines on site, through the question and answer interaction module, input the question "What is the possible cause of abnormal vibration of a certain type of wind turbine and how to solve it?" The question understanding submodule first analyzes the question and identifies key information such as "a certain type of wind turbine" and "abnormal vibration", and determines the problem type as fault diagnosis and solution suggestion. Then, the knowledge retrieval and reasoning submodule retrieves knowledge related to the abnormal vibration of the wind turbine in the knowledge graph, and finds possible causes including blade imbalance, gearbox failure, bearing wear, etc. At the same time, the large model combines the information in the knowledge graph to reason and analyze the possibility of different causes. For example, if the local wind speed changes greatly recently, and the vibration frequency is related to the blade rotation frequency, the possibility of blade imbalance is greater. According to the reasoning result, the answer generation and feedback submodule generates a detailed answer, including possible causes, further checking methods (such as using a laser alignment instrument to check blade balance, checking gearbox oil quality, etc.) and corresponding solutions (such as adjusting blade counterweight, replacing worn gears, etc.), and feeds back to the maintenance personnel in a clear and easy-to-understand text form.
[0121] Using the fault prediction and health assessment module, the health of the key components of the wind turbine is evaluated and the fault is predicted. In the data feature extraction submodule, time domain and frequency domain analysis is performed on the vibration data to extract features such as vibration amplitude, frequency components, etc.; trend analysis is performed on the temperature data to extract features such as temperature change rate, etc. In the prediction model construction submodule, a long short-term memory network (LSTM) model is used to predict the gearbox oil temperature to detect potential risks of abnormal temperature rise in advance. In the health state evaluation submodule, according to the device health standards set in the knowledge graph, combined with the extracted feature vector, the weights of each evaluation index are determined by the analytic hierarchy process (AHP), and then the overall health state of the wind turbine is scored and graded by the fuzzy comprehensive evaluation method. For example, if the evaluation score is 80 or above, it is determined to be in a healthy state; if the score is between 60 and 80, it is in a sub-healthy state; if the score is less than 60, there is a risk of failure.
[0122] Visual interaction and decision support: Through the visual interaction and decision support module, the operation and maintenance personnel can intuitively view the health status and fault prediction results of the wind turbine. In the visualization display submodule, the health score distribution of each component of the wind turbine is displayed in the form of a chart, such as the health status of the blade, gearbox, generator and other components in the form of a column chart, which facilitates the operation and maintenance personnel to quickly understand the overall condition of the equipment. In the user interaction submodule, the operation and maintenance personnel can input query conditions through the interface, such as selecting a specific wind turbine, time range, etc., to obtain real-time running data statistics and health status trend of the equipment within the specified time period. When the system predicts that the equipment has potential fault risk, the decision suggestion generation submodule will generate a corresponding maintenance decision suggestion report according to the severity and possible impact of the fault. For example, it is recommended to arrange inspection and maintenance of the gearbox within the next 3 days, and provide the required spare parts list and maintenance personnel arrangement suggestion.
[0123] After the operation and maintenance personnel use the system to solve the problem, they give feedback on the system's answers, evaluating the accuracy and usefulness of the answers. The user feedback analysis submodule collects this feedback information and finds that some users think the description of the steps in the answer is not detailed enough. This problem is fed back to the large model training module and the knowledge graph construction module. The large model training module increases the learning of detailed operation step description in subsequent retraining; the knowledge graph construction module supplements and perfects the knowledge related to the inspection method, making it more detailed and accurate. At the same time, the system performance monitoring submodule monitors the response time and throughput of the system in real time to ensure that the system can still respond quickly to user problems under high concurrency. If it is found that the system response time is too long, optimization methods such as optimizing database query statements and increasing server resources are used to optimize.
[0124] Example 3
[0125] Photovoltaic power plant equipment performance evaluation and optimization scenario
[0126] In a certain photovoltaic power plant, through sensors installed on photovoltaic modules, inverters, combiner boxes and other equipment, the operation data of the photovoltaic power plant are collected, including light intensity, temperature, current, voltage, etc. At the same time, historical data such as construction planning, equipment parameters of the photovoltaic power plant, and local weather forecast data are collected. The data collection and preprocessing module cleans the collected data to remove erroneous data caused by sensor failure or communication interference. For missing light intensity data, linear interpolation method is used for filling. Then the data is standardized to unify the current and voltage data of different equipment to the same scale range. Through feature engineering, the performance indicators such as the photoelectric conversion efficiency of the photovoltaic module and the efficiency of the inverter are calculated to provide basic data for subsequent equipment performance evaluation.
[0127] The operation data, meteorological data and historical data of the photovoltaic power station are transmitted to the multi-source heterogeneous data deep fusion module. In the data standardization submodule, the light intensity, temperature and other data are normalized to have a unified dimension. In the feature fusion submodule, the deep autoencoder (DAE) is used to fuse different types of data features and extract a feature vector that can comprehensively reflect the operation state of the photovoltaic power station. For example, the light intensity, temperature and current, voltage characteristics of the photovoltaic module are fused to obtain a comprehensive feature vector containing environmental factors and equipment operation state. In the semantic alignment submodule, through the entity mapping of the knowledge graph, the related information about the photovoltaic module, inverter and other equipment in different data sources is uniformly associated, and the semantic ambiguity is eliminated.
[0128] Entities and relationships are extracted from preprocessed data and related technical documents to build a photovoltaic power station knowledge graph. The entities such as "photovoltaic module", "inverter", "junction box" and the relationships such as "photovoltaic module-connection-junction box", "junction box-connection-inverter", "performance degradation of photovoltaic module-reason-dust obstruction" are identified. The knowledge fusion submodule integrates the extracted knowledge to ensure the consistency and accuracy of the knowledge graph. With the operation of the photovoltaic power station and the development of technology, the knowledge update and expansion submodule continuously updates the knowledge graph, such as adding the performance parameters and maintenance knowledge of new photovoltaic modules.
[0129] A large amount of text data related to photovoltaic power station equipment performance evaluation and optimization is collected, including industry reports, research papers, operation and maintenance experience sharing, etc. These data are labeled, including question types (such as performance evaluation, fault diagnosis, optimization suggestions, etc.) and answer information. A suitable pre-trained language model (such as GPT-3-like model) is selected and fine-tuned on the labeled dataset. During fine-tuning, structured knowledge from the knowledge graph is combined to enable the model to better understand professional knowledge in the photovoltaic power station field. After training, the model is evaluated and optimized through the model evaluation and optimization submodule, using accuracy, mean absolute error and other evaluation indicators to evaluate the model and optimize the model based on the evaluation results to improve the accuracy and reliability of the model in answering photovoltaic power station questions.
[0130] When the photovoltaic power station manager wants to evaluate the performance of photovoltaic components in a certain area, the question "How to evaluate the performance and optimize the photovoltaic components in a certain area whose recent power generation efficiency has decreased?" is input through the question and answer interaction module. The question understanding submodule first performs word segmentation, part-of-speech tagging, and syntactic analysis on the question, identifies key information such as "photovoltaic components in a certain area", "power generation efficiency has decreased", "evaluate performance", and "optimize", and determines that the question type is performance evaluation and optimization suggestion. Then, combined with the concepts and entities in the knowledge graph, the question is understood and disambiguated semantically, and it is clear that the user is concerned about the performance evaluation and optimization measures after the power generation efficiency of photovoltaic components in a specific area has decreased.
[0131] The knowledge retrieval and reasoning submodule retrieves knowledge related to photovoltaic component performance evaluation and optimization in the knowledge graph based on the results of question understanding. First, the entity information related to the photovoltaic components in the area is obtained, such as component model, installation location, service life, and possible reasons for the decrease in power generation efficiency, such as dust obstruction, component aging, local shadow, and line loss. Then, using the reasoning ability of the large model, combined with the real-time collected operation data of the photovoltaic components in the area (such as light intensity, temperature, current, and voltage) and meteorological data (such as cloud coverage), the various possible reasons are analyzed and reasoned. For example, if the current light intensity is normal, but the component temperature is high, and there is no obvious dust accumulation in the area recently, the large model will further analyze the possibility of component aging and local shadow. Through the mining and comparison of historical data, it is found that some components in the area have been in service for a long time, and there are new constructions around the area that may cause local shadow, so it is judged that component aging and local shadow are the main reasons for the decrease in power generation efficiency.
[0132] The answer generation and feedback submodule generates natural language form answers based on the results of knowledge retrieval and reasoning and feeds them back to the user. For this question, the answer content may include: "The efficiency of photovoltaic components in a certain area may be reduced due to component aging and local shading. To evaluate its performance, it is recommended to first use professional photovoltaic component detection equipment, such as a photovoltaic IV curve tester, to detect the electrical performance of the components and obtain parameters such as open-circuit voltage, short-circuit current, and maximum power, and compare them with the initial parameters and standard parameters of the components to evaluate the aging degree of the components. At the same time, use a drone equipped with a high-definition camera to patrol the photovoltaic components in the area and check for local shading and determine the source and scope of the shading. For optimization measures, for severely aged components, it is recommended to replace them in a timely manner; for local shading problems, you can adjust the installation angle of the components or remove the shading objects to reduce the impact of shading on power generation efficiency. In addition, regular cleaning and maintenance of photovoltaic components can also help improve power generation efficiency." The answer is presented in a clear and easy-to-understand text form and supports multiple output formats, such as providing related detection equipment usage tutorial video links, diagrammatic instructions for implementing optimization measures, etc., to enhance the visual effect and readability of the answer. At the same time, the system records this user's question and answer for subsequent system optimization and user behavior analysis.
[0133] During the operation of the photovoltaic power station, the fault prediction and health assessment module continuously monitors and evaluates key equipment such as photovoltaic components and inverters. The data feature extraction submodule extracts time and frequency domain features from the collected equipment operation data using methods such as Fourier transform and wavelet transform. For example, Fourier transform is used to analyze the frequency components of the current data of photovoltaic components to obtain features related to the performance of the components; wavelet transform is used to decompose the voltage data of the inverter at multiple scales to extract energy distribution features in different frequency bands. The prediction model construction submodule trains a fault prediction model based on deep learning algorithms such as convolutional neural networks (CNN) and recurrent neural networks (RNN), combining historical fault data and current equipment operation features. For example, a CNN model is used to analyze image data of photovoltaic components (such as temperature distribution images of components obtained by a thermal imager) to predict whether there is a potential hot spot fault in the components; an RNN model is used to analyze the time series of inverter operation data to predict the fault trend of the inverter. The health status assessment submodule determines the weights of each evaluation index based on the health standards of photovoltaic equipment in the knowledge graph and the extracted feature vectors, and then uses the fuzzy comprehensive evaluation method to score and grade the health status of photovoltaic components and inverters. For example, the health status of photovoltaic components is divided into four levels: "good", "general", "poor", and "severe fault". When the health score of the components is below a certain threshold, the system sends an early warning signal to prompt the maintenance personnel to maintain in a timely manner.
[0134] The photovoltaic power station manager intuitively understands the operation status and equipment performance of the photovoltaic power station through visual interaction and decision support module. The visual display submodule presents data in various graphical ways, such as displaying the distribution of photovoltaic components in each area of the photovoltaic power station through a geographic information system (GIS) map, and marking the health status of the components with different colors; using a line chart to show the trend of the power generation efficiency of photovoltaic components in a certain area over time; using a column chart to compare the conversion efficiency of different inverters. The user interaction submodule supports the manager to customize query conditions through the interface, such as selecting specific time period, area, equipment type, etc., to obtain real-time operation data statistics and analysis results. For example, the manager can query "the power generation efficiency fluctuation of photovoltaic components in area A in the past week and the corresponding meteorological data", and the system will quickly generate the corresponding chart and data report. When the system detects that the photovoltaic power station has performance problems or potential failures, the decision suggestion generation submodule generates detailed decision suggestion reports based on large model reasoning and knowledge graph rules. The report content includes fault cause analysis, recommended maintenance measures, maintenance time arrangement, required manpower and material resources, etc. For example, if the system predicts that a certain inverter has a high risk of failure in the next week, the decision suggestion report may suggest that preventive maintenance be performed on the inverter during the low electricity consumption period of the weekend, replace some vulnerable components, and arrange professional technicians to operate on site, while providing a list of spare parts and tools required for maintenance.
[0135] The photovoltaic power station management personnel feed back the answers and decision suggestions provided by the system during the use of the system. The user feedback analysis submodule collects these feedback information, such as the management personnel thinks that the accuracy of some fault prediction needs to be improved, or some optimization suggestions have difficulties in actual operation, etc. Through the analysis of the feedback information, the module and link where the problem lies are determined, such as the training data of the fault prediction model is not comprehensive enough, or the description of some optimization measures in the knowledge graph is not detailed enough. These problems are fed back to the related modules, such as the large model training module, the knowledge graph construction module and the fault prediction and health assessment module, etc. The large model training module adjusts the training strategy according to the new feedback data, increases the related training data, and re-trains the model to improve the accuracy of fault prediction; the knowledge graph construction module updates and perfects the related knowledge, and supplements the detailed operation steps and actual application cases of the optimization measures; the fault prediction and health assessment module optimizes the feature extraction method and the prediction model structure, and improves the monitoring and prediction ability of the system to the equipment state of the photovoltaic power station. At the same time, the system performance monitoring submodule monitors the running performance indicators of the system in real time, such as response time, throughput, accuracy, etc. If it is found that the system response time is too long, optimization is carried out through optimizing the database query statement, adjusting the server resource configuration, etc.; if the accuracy decreases, the reason is analyzed in depth, and the model is re-evaluated and optimized to ensure that the system can continuously and stably provide high-quality services for the operation and management of the photovoltaic power station.
[0136] In summary, the present application realizes the comprehensive collection, cleaning, conversion and integration of various data of new energy equipment through innovative data collection and preprocessing technology and multi-source heterogeneous data deep fusion module, providing accurate and comprehensive data support for equipment health assessment. With the help of the knowledge graph construction module and the large model training module, combined with deep learning and machine learning algorithms, accurate diagnosis and early prediction of new energy equipment faults are realized. It can identify potential fault risks of equipment at an early stage, effectively reduce equipment failure rate, reduce downtime and reduce maintenance cost. Using the powerful semantic understanding and reasoning ability of the large model, and the structured knowledge storage and association query function of the knowledge graph, intelligent question answering for user problems is realized. It can accurately understand the user's problem intention and provide detailed, accurate and targeted solutions. Through the fault prediction and health assessment module and the visual interaction and decision support module, scientific basis is provided for the whole life cycle management of new energy equipment. Including decision support for equipment selection, maintenance plan development, fault handling scheme selection, etc. to help users make more reasonable and accurate decisions and improve the overall operation efficiency of new energy equipment. A continuous learning and optimization mechanism is established, and the knowledge graph and the large model are updated through the analysis of user feedback and system performance monitoring data by the system optimization module, so that the system can adapt to the development of new energy technology and the changes of equipment operation environment, and continuously improve the performance and service quality of the system.
[0137] The above describes the preferred embodiments of the present application. It should be noted that, for those skilled in the art, several improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as falling within the scope of the present application.
Claims
1. A smart question-and-answer system for health management of new energy equipment based on large-scale models and knowledge graphs, including a data acquisition and preprocessing module, a deep fusion module for multi-source heterogeneous data, a knowledge graph construction module, a large-scale model training module, a question-and-answer interaction module, a fault prediction and health assessment module, a visualization interaction and decision support module, and a system optimization module; The data acquisition and preprocessing module is used to acquire multi-source data from new energy equipment and perform preprocessing such as cleaning and noise reduction. The multi-source heterogeneous data deep fusion module is used to standardize, fuse features, and align semantics of the preprocessed data; The knowledge graph construction module is used to extract entities and relationships from data and documents, and to construct and update the knowledge graph. The large model training module is used to fine-tune the pre-trained language model using labeled data; The question-and-answer interaction module is used to understand user questions, retrieve knowledge from the knowledge graph, reason, and generate answers to be fed back to the user. The fault prediction and health assessment module is used to extract equipment operation data features, build a prediction model, and assess the health status of the equipment. The visualization interaction and decision support module is used to visualize data, support user interaction, and generate decision suggestion reports. The system optimization module is used to optimize the system based on user feedback and system performance monitoring results.
2. The intelligent question-and-answer system for health management of new energy equipment based on a large model and knowledge graph as described in claim 1, characterized in that, The data acquisition and preprocessing module collects multi-source data including sensor data, monitoring system data, meteorological data, and data from equipment maintenance manuals and technical documents. During data cleaning, outliers are removed using the 3σ principle, and missing data is filled in using linear interpolation.
3. The intelligent question-and-answer system for health management of new energy equipment based on a large model and knowledge graph as described in claim 1, characterized in that, The multi-source heterogeneous data deep fusion module includes a data standardization submodule, a feature fusion submodule, and a semantic alignment submodule. The data standardization submodule performs format unification and dimensional conversion on different types of data; The feature fusion submodule uses principal component analysis and deep autoencoder for feature extraction and fusion. The semantic alignment submodule uses knowledge graph entity mapping technology to align related entities in different data sources.
4. The intelligent question-and-answer system for health management of new energy equipment based on a large model and knowledge graph as described in claim 1, characterized in that, The knowledge graph construction module includes a knowledge extraction submodule, a knowledge fusion submodule, and a knowledge update and expansion submodule. The knowledge extraction submodule identifies entities and relationships from preprocessed data and documents; The knowledge fusion submodule integrates the extracted knowledge and eliminates duplicate and conflicting information; The knowledge update and expansion submodule updates the knowledge graph periodically based on new data and technological developments.
5. The intelligent question-and-answer system for health management of new energy equipment based on a large model and knowledge graph as described in claim 1, characterized in that, The large model training module collects and cleans text data related to new energy equipment, selects a suitable pre-trained language model for fine-tuning, and incorporates entity and relation information from the knowledge graph into the training in the form of embedding vectors during the fine-tuning process. The model evaluation and optimization submodule evaluates and optimizes the model using evaluation metrics such as accuracy, recall, F1 score, and mean absolute error.
6. The intelligent question-and-answer system for health management of new energy equipment based on a large model and knowledge graph as described in claim 1, characterized in that, The question-and-answer interaction module includes a question understanding submodule, a knowledge retrieval and reasoning submodule, and an answer generation and feedback submodule. The question understanding submodule performs word segmentation, part-of-speech tagging, syntactic analysis, semantic understanding, and disambiguation on user questions. The knowledge retrieval and reasoning submodule retrieves relevant knowledge from the knowledge graph and performs analysis and reasoning by combining the reasoning capabilities of the large model and real-time data. The answer generation and feedback submodule presents the reasoning results in natural language and supports multiple output formats, while also recording user questions and answers.
7. The intelligent question-and-answer system for health management of new energy equipment based on a large model and knowledge graph as described in claim 1, characterized in that, The fault prediction and health assessment module includes a data feature extraction submodule, a prediction model construction submodule, and a health status assessment submodule, including: The data feature extraction submodule uses Fourier transform and wavelet transform methods to extract time-domain and frequency-domain features from the equipment operation data; The prediction model construction submodule is based on deep learning algorithms and trains the fault prediction model by combining historical fault data and current equipment operating characteristics. The health status assessment submodule uses the health standards in the knowledge graph, employs the analytic hierarchy process to determine the weights of the assessment indicators, uses the fuzzy comprehensive evaluation method to score and classify the health status of the equipment, and issues an early warning signal when the score is below the threshold.
8. The intelligent question-and-answer system for health management of new energy equipment based on a large model and knowledge graph as described in claim 1, characterized in that, The visualization interaction and decision support module includes a visualization display submodule, a user interaction submodule, and a decision suggestion generation submodule, wherein: The visualization submodule presents data in various graphical formats, including but not limited to maps, line charts, and bar charts, using geographic information systems. The user interaction submodule allows users to customize query conditions through the interface and obtain relevant data statistics and analysis results in real time. The decision suggestion generation submodule, based on large model reasoning and knowledge graph rules, generates decision suggestion reports, including but not limited to fault cause analysis, maintenance measures, time arrangements, and resource requirements, based on equipment fault prediction and health assessment results.
9. The intelligent question-and-answer system for health management of new energy equipment based on a large model and knowledge graph as described in claim 1, characterized in that, The system optimization module includes a user feedback analysis submodule, a system performance monitoring submodule, and a knowledge update and model retraining submodule, wherein: The user feedback analysis submodule collects user feedback on the system's answers and decision suggestions, analyzes the modules and processes where problems occur, and feeds back to the relevant modules. The system performance monitoring submodule monitors the system's performance indicators in real time, including but not limited to response time, throughput, and accuracy, and takes corresponding optimization measures when the indicators are abnormal. The knowledge update and model retraining submodule periodically updates the knowledge graph and retrains the large model.
10. The intelligent question-and-answer system for health management of new energy equipment based on a large model and knowledge graph as described in claim 6, characterized in that, The knowledge retrieval and reasoning submodule captures contextual semantic information in user questions in real time and constructs dynamic reasoning paths related to the device's operating status. The answer generation and feedback submodule generates multi-dimensional explanatory answers for complex fault scenarios, including fault evolution paths and comparisons with similar historical cases. It also supports presenting knowledge relationships in the form of interactive graph nodes, enabling in-depth semantic analysis and logical deduction of user questions.
11. The intelligent question-and-answer system for health management of new energy equipment based on a large model and knowledge graph as described in claim 7, characterized in that, The prediction model construction submodule adopts a multi-task learning framework, simultaneously training equipment fault prediction, remaining life prediction, and performance degradation trend prediction tasks, sharing the underlying feature extraction network and outputting prediction results for each task; the health status assessment submodule combines the failure mode knowledge of equipment components in the knowledge graph to construct a dynamic assessment model based on a Bayesian network, updating the posterior probability of component health status by real-time input of multi-source monitoring data, and using reinforcement learning algorithms to adaptively adjust the index weights of the analytic hierarchy process, thereby achieving probabilistic assessment and dynamic threshold early warning of equipment health status.
Citation Information
Cited By
Electrochemical energy storage system fault identification and risk management system
CN121258214A
Photovoltaic panel cleaning robot fault detection method and system
CN121625229A
A method and system for full life cycle health management of a heating furnace installation
CN122366988A