Systems and methods for smart grid data analysis and management

The PowerGPT system addresses grid management challenges by providing predictive maintenance and interactive dashboards for smart grid data analysis, enhancing grid efficiency and reliability through proactive issue detection and optimization.

WO2025264302A1PCT designated stage Publication Date: 2025-12-26FLORIDA ATLANTIC UNIVERSITY

Patent Information

Application Number
PCT/US2025/025123
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-17
Filing Date
2025-04-17
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Modern grid operators face challenges in data processing, fault detection, asset optimization, and real-time decision-making due to the integration of renewable energy sources and the wear and tear of power transformers, leading to inefficiencies, increased downtime, and reduced reliability across sectors.

Method used

A generative pre-trained Transformer (GPT)-based system, PowerGPT, for smart grid data analysis and management, featuring predictive maintenance, interactive data dashboards, and a conversational chatbot for proactive system management, integrating wind power forecasting and power plant asset mapping.

Benefits of technology

Enhances grid efficiency and reliability by identifying potential issues before they cause disruptions, optimizing energy distribution, and ensuring grid stability through real-time data processing and proactive maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025025123_26122025_PF_FP_ABST
    Figure US2025025123_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are systems, methods, and a software system that can be used for prognostic management of energy systems including smart grid power transformers. In one embodiment, a smart grid data management system is provided. An exemplary system can comprise a Transformer-based fault classification model and a natural language processor (e.g., NLP model), where the classification model is configured to classify faults in the smart grid system and the natural language processor or model is configured to output responses and information using the classified faults
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket No.11605-066WO1 FAU 202408 SYSTEMS AND METHODS FOR SMART GRID DATA ANALYSIS AND MANAGEMENT CROSS-REFERENCE TO RELATED APPLICATIONS [1] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 660,675, titled “POWERGPT: SYSTEMS AND METHODS FOR SMART GRID DATA ANALYSIS AND MANAGEMENT”, filed on June 17, 2024, the content of which is incorporated by reference herein in its entirety. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT [2] This invention was made with government support under grant number CMMI- 2145571 and CNS-1950400 awarded by the National Science Foundation. The government has certain rights in the invention. BACKGROUND [3] With new emerging technologies connecting to the grid every day, the variety of energy consumption has put a diverse strain on electric power systems. For instance, the integration of renewable energy sources, such as solar and wind, introduces variability and intermittency, requiring advanced energy storage and grid management strategies to maintain stability. Similarly, the transmission of appropriate electrical energy from current distribution systems relies on the proper maintenance of power transformers. These devices regulate voltage through electromagnetic induction but dissipate energy as heat, exposing them to thermal stress and potential malfunctions. Over time, wear and tear lead to transient disturbances and mechanical failures, compromising both the transformer’s integrity and the stability of its electrical output. Modern grid operators face mounting challenges in data processing, fault detection, asset optimization, and real-time decision-making. Without advanced analytical tools, operators struggle to effectively monitor system-wide performance, predict potential failures, visualize complex relationships between components, and optimize resource allocation. This reactive approach to grid management leads to inefficiencies, increased downtime, higher maintenance costs, and reduced reliability across industrial, commercial, and residential sectors. Attorney Docket No.11605-066WO1 FAU 202408 [4] As such, there is a need for improved systems that can holistically manage the electrical grid. SUMMARY [5] Disclosed herein are systems, methods, and a software system for proactive data management of energy systems, including smart grid power transformers, wind farms, and power plant asset mapping, featuring an interactive data dashboard enhanced with natural language processing (NLP). In one example, a generative pre-trained Transformer (GPT)- based system for smart grid data analysis and management with power transformer condition monitoring capabilities is provided. [6] In one embodiment, a system including a proposed cutting-edge AI assistant, PowerGPT, transforms big data analysis and management in Smart Grid systems. This innovation delivers proactive system management through predictive maintenance and optimization capabilities that identify potential issues before they cause disruptions. Its streamlined grid control offers a user-friendly interface with visual diagrams and an intuitive chatbot that makes complex data more easily accessible and interpretable to operators. As an all-in-one platform, PowerGPT integrates wind power forecasting, power plant and assets mapping, fault diagnosis, and interactive data dashboards to provide comprehensive grid insights. Most importantly, the system is easy to scale, seamlessly adapting to various grid sizes and infrastructures with a modular design that allows for effortless integration of new functions as grid technologies continue to evolve. Embodiments of the present disclosure provide an AI- based personalized assistant system designed to streamline smart grid data maintenance and analysis. It enables intuitive visualization and early issue diagnosis within grids through a user- friendly natural language interface, which interacts with multiple backend functions on a cloud server. In some examples, by harnessing the power of deep learning and artificial intelligence, the embodiments described herein leverage established fault patterns within critical power grid components, such as transformers, to train a sophisticated deep neural network model to realize prognostic health management (PHM). This model can be seamlessly integrated into a backend Application Programming Interface (API). [7] In some implementations, through an intuitive Graphical User Interface (GUI), users of any expertise level can submit performance metrics for analysis by the model, receiving prompt data visualization and diagnostic insights in return. Moreover, the proposed GUI is designed to enhance user engagement, featuring a conversational chatbot interface that Attorney Docket No.11605-066WO1 FAU 202408 supplements its findings with relevant additional information. While akin to ChatGPT in functionality, the proposed product uniquely provides insights into power grid data and specializes in interpreting and diagnosing faults within power transformers—a capability currently unparalleled in existing offerings. Energy / power related entities (e.g., companies, utilities) can apply the embodiments described here for data management, such as prognostic health management, to operate and maintain industrial assets. [8] In some implementations, a power monitoring system is provided. The power monitoring system includes: at least one processor; and a memory having instructions thereon, wherein the instructions when executed by the at least one processor, cause the at least one processor to: analyze, using a classification model, energy system data to detect one or more fault classification outputs; responsive to receiving a direct user input, determine, using a second model, one or more user input parameters; determine (e.g., select, output, optimize), based at least in part on one or more fault characteristics associated with the one or more fault classification outputs and the one or more user input parameters, at least one of a plurality of models; and generate, using at least one of a plurality of models, one or more predictive outputs based at least in part on the one or more fault characteristics and the one or more user input parameters. [9] In some implementations, the second model is a generative Transformer-based model.

[0010] In some implementations, generating one or more predictive outputs includes directing the energy system data to at least one of the plurality of models.

[0011] In some implementations, an interactive power grid map, GridWatch, enables users to request visual representations of power generation facilities based on location and energy type. The system processes the request by querying a database of power plants, filtering results based on the specified criteria, and displaying an interactive map where facilities are color- coded according to their fuel type. If no location is specified, the system defaults to showing all power plants. This allows PowerGPT to dynamically generate and execute queries.

[0012] In some implementations, a comprehensive overview of a power grid landscape, the Data Dashboard, provides insights into power generation, consumption, and emissions data. The dashboard integrates an interactive Geographic Information System (GIS) map, which visualizes energy infrastructure, including power plants and transmission lines. Users can explore regional energy statistics, tracking production trends across different fuel sources and understanding how various sectors contribute to overall energy demand and emissions. Attorney Docket No.11605-066WO1 FAU 202408

[0013] In some implementations, the dashboard aggregates data from multiple sources, such as the Public Service Commission and the U.S. Energy Information Administration, ensuring up-to-date energy market insights. Users can access real-time pricing data, sector- based energy consumption breakdowns, and environmental impact assessments. The dashboard supports decision-making for policymakers, researchers, and industry professionals by presenting key metrics related to energy efficiency, renewable integration, and grid stability.

[0014] In some implementations, wind power prediction utilizes machine learning models to estimate energy generation based on wind conditions. The predictive model is using a tree- based learning algorithm that refines its predictions by making iterative improvements at each decision point. The system integrates with the National Weather Service API to retrieve real- time wind speed data based on user-specified locations. If no location is provided, the system defaults to using the user's IP-based location.

[0015] In some implementations, the system generates predictive visualizations in response to user queries regarding wind generation prediction. The system generates two graphs: one showing the hourly power generation forecast for the next 24 hours with confidence intervals, and another displaying cumulative power generation over the same period.

[0016] In some implementations, the one or more predictive outputs and / or the one or more user input parameters include at least one of a location, a fault type, an energy system / component, and / or power transformer(s).

[0017] In some implementations, one or more predictive outputs include a predicted fault. Incorporated features that enhance interactivity such as a corresponding signal graph and a graphic of a three-phase transformer highlight the specific areas affected by the fault.

[0018] In some implementations, the instructions further cause the at least one processor to: generate the plurality of models including the classification model and the second model; and train the plurality of models using historical energy system data.

[0019] In some implementations, the instructions further cause the at least one processor to: pre-process the historical energy system data prior to training the plurality of models.

[0020] In some implementations, the historical energy system data is retrieved from one or more publicly available databases.

[0021] In some implementations, the instructions further cause the at least one processor to: output, via a display or graphical user interface, the one or more predictive outputs and / or one or more fault classification outputs / characteristics. Attorney Docket No.11605-066WO1 FAU 202408

[0022] In some implementations, the output includes conversational text, graphs, and / or charts.

[0023] In some implementations, the second model includes a chat-based model or model architecture.

[0024] In some implementations, the classification model is seamlessly integrated into a backend Application Programming Interface (API).

[0025] In some implementations, the plurality of models includes one or more of a deep neural network model, a Transformer model, and a large language model.

[0026] In some implementations, the energy system data includes real-time data obtained from sensors operatively coupled with energy system components, and the instructions further cause the at least one processor to: continuously analyze streaming data from the sensors in order to promptly response to grid issues.

[0027] In some implementations, the instructions further cause the at least one processor to: trigger or cause a corrective action with respect to at least one identified fault classification output (e.g., generate an alert, deploy personnel to a target location, and / or automatically activate or deactivate a system component).

[0028] In some implementations, a method for prognostic power monitoring is provided. The method includes: analyzing, using a classification model, energy system data to detect one or more fault classification outputs; responsive to receiving a direct user input, determining, using a second model, one or more user input parameters; determining based at least in part on one or more fault characteristics associated with the one or more fault classification outputs and the one or more user input parameters, at least one of a plurality of models; and generating, using at least one of a plurality of models, one or more predictive outputs based at least in part on the one or more fault characteristics and the one or more user input parameters.

[0029] In some implementations, a non-transitory computer-readable medium, including a memory having instructions stored thereon to perform / implement any of the operations / system described herein, is provided.

[0030] Other systems, methods, features, and / or advantages will be or may become apparent to one with skill in the art upon examination of the following drawings and detailed description. It is intended that all such additional systems, methods, features, and / or advantages be included within this description and be protected by the accompanying claims. Attorney Docket No.11605-066WO1 FAU 202408 BRIEF DESCRIPTION OF DRAWINGS

[0031] The components in the drawings are not necessarily to scale relative to each other. Like reference, numerals designate corresponding parts throughout the several views.

[0032] FIG.1 is an example system in accordance with certain embodiments of the present disclosure.

[0033] FIG. 2A and FIG. 2B are flowcharts depicting example operations for the exemplary system in accordance with certain embodiments of the present disclosure.

[0034] FIG. 3A and FIG. 3B are flowcharts showing example user-privileged and administrative-privileged (admin-privileged) use cases, respectively, in accordance with certain embodiments of the present disclosure.

[0035] FIG.4 is an example structure of the exemplary system comprising a Transformer- based fault classification model and a natural language processor (e.g., NLP model), wherein the classification model is configured to classify faults in the smart grid system and the NLP model is configured to output responses and information using the classified faults.

[0036] FIGS.5A – 5C show the classification results for the models and evaluation results for the graphical user interface (GUI) of the exemplary system. FIG. 5A shows this dataset with a total of 45 fault classes requiring a multiclass classification solution. FIG. 5B shows classification matrices generated by two fault classification models, e.g., Indirect Symmetrical Phase Angle Regulator (ISPAR) series transformer and power transform. FIG. 5C shows the GUI configured to perform and visualize the fault diagnosis of the classification model.

[0037] FIG.6 shows an example computing device having a basic configuration. DETAILED DESCRIPTION

[0038] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, can also be provided in combination with a single embodiment. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, can also be provided separately or in any suitable sub-combination. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure. Attorney Docket No.11605-066WO1 FAU 202408 DEFINITIONS

[0039] In this specification and in the claims that follow, reference will be made to a number of terms, which shall be defined to have the following meanings:

[0040] Throughout the description and claims of this specification, the word “comprise” and other forms of the word, such as “comprising” and “comprises,” means including but not limited to, and are not intended to exclude, for example, other additives, segments, integers, or steps. Furthermore, it is to be understood that the terms comprise, comprising, and comprises as they relate to various embodiments, elements, and features of the disclosure also include the more limited embodiments of “consisting essentially of” and “consisting of.”

[0041] As used herein, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a “software system” includes embodiments having multiple components or modules unless the context clearly indicates otherwise.

[0042] Ranges can be expressed herein as from “about” one particular value and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another embodiment. It should be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint.

[0043] As used herein, the terms “optional” or “optionally” mean that the subsequently described event or circumstance may or may not occur, and that the description includes instances where said event or circumstance occurs and instances where it does not.

[0044] For the terms “for example” and “such as,” and grammatical equivalences thereof, the phrase “and without limitation” is understood to follow unless explicitly stated otherwise.

[0045] As used herein, the terms “data,” “content,” “information,” and similar terms may be used interchangeably to refer to data capable of being transmitted, received and / or stored in accordance with embodiments of the present invention.

[0046] Embodiments of the present disclosure provide a generative pre-trained Transformer (GPT)-based system for smart grid data analysis and management with many embedded functions, such as the grid virtualization and power transformer condition monitoring capabilities. Some benefits of the proposed systems and methods include: (a) Improved efficiency and reliability: By analyzing vast amounts of data from the smart grid, the Attorney Docket No.11605-066WO1 FAU 202408 system could identify patterns of energy usage, predict outages, and optimize energy distribution leading to a more efficient and reliable grid. (b) Enhanced power transformer maintenance: Continuous monitoring the condition of power transformers, a critical component of the grid, would allow for preventive maintenance and avoid costly failures. (c) Cost savings: Early detection of transformer issues can prevent catastrophic failures and the associated repair costs. Additionally, improved grid efficiency can lead to reduced energy waste and lower overall costs. (d) Integration with renewables: As the use of renewable energy sources like solar and wind power increases, the grid needs to become more flexible and adaptable. The proposed system could help manage the fluctuations in renewable energy production and ensure grid stability through accurate wind power generation forecasting.

[0047] Additional benefits of the proposed systems and methods include: (a) Proactive System Data Management: Enables predictive maintenance and optimization to enhance grid performance and prevent potential failures. (b) Streamlined Grid Control: Features a user- friendly interface with visual diagrams and an interactive chatbot for efficient system management. (c) All-in-One Platform: Integrates wind power forecasting, power plant and asset mapping, fault diagnosis, and an interactive data dashboard to provide a comprehensive grid management solution. (d) Easy to Scale: Seamlessly adapts to various grid sizes and infrastructures, with a modular design that allows for effortless integration of new functions.

[0048] Embodiments of the present disclosure address all or some of the following challenges: (a) Aging infrastructure: The electrical grid in many areas is aging and in need of modernization. The proposed system could be a key tool for improving the efficiency and reliability of existing infrastructure. (b) Growth of renewable energy: The increasing use of renewable energy sources creates challenges for grid management. This system could help address these challenges and facilitate the integration of renewables. (c) Focus on grid resilience: Extreme weather events and cyberattacks are becoming more common, highlighting the need for a more resilient grid. This system could play a role in improving grid resilience by providing early warnings of potential problems.

[0049] The success of the proposed system depends on its ability to handle the complex and real-time nature of smart grid data. Security would be a major concern, as the system would have access to sensitive data about the grid. Regulatory hurdles may need to be addressed before the system can be widely deployed. The commercial applications of the technology span across various sectors of the energy industry as outlined below Attorney Docket No.11605-066WO1 FAU 202408

[0050] For Utility Companies: (a) Predictive Maintenance: By analyzing sensor data from transformers and other grid components, the system can predict potential failures and schedule maintenance before critical breakdowns occur. This reduces downtime repair costs and improves grid reliability. (b) Demand Forecasting: The system can analyze historical data and real-time usage patterns to forecast energy demand with greater accuracy. This allows utilities to optimize generation and distribution to meet demand efficiently, reducing costs and minimizing the need for peak power plants. (c) Dynamic Pricing: The system can be used to implement dynamic pricing models where energy prices fluctuate based on real-time demand. This incentivizes consumers to shift their energy usage to off-peak hours, reducing peak load on the grid and lowering overall costs. (d) Integration of Renewables: The system can help manage the variable nature of renewable energy sources like solar and wind. By predicting generation and optimizing grid operations, the system can ensure smooth integration of renewables into the grid.

[0051] For Businesses and Industries: (a) Energy Cost Optimization: Businesses can leverage the system’s insights to optimize their energy consumption patterns by offering a subscription-based model for utility companies, allowing them to efficiently manage energy distribution and grid performance. This can be achieved by identifying and reducing peak usage periods or scheduling energy-intensive tasks for off-peak hours. (b) Improved Facility Management: The system can monitor energy usage in different parts of a facility, allowing businesses to identify areas for improvement and implement targeted energy-saving measures. (c) Demand Response Programs: Businesses can participate in demand response programs where they agree to reduce their energy usage during peak periods in exchange for financial incentives. The system can help businesses identify opportunities to participate in such programs and optimize their response.

[0052] For Consumers: (a) Smart Metering and Home Energy Management: The system can be integrated with smart meters to provide consumers with real-time information about their energy usage. This empowers consumers to make informed choices about their energy consumption and potentially lower their energy bills. (b) Time-of-Use Pricing: Consumers can benefit from time-of-use pricing models where electricity costs vary depending on the time of day. The system can help them adjust their usage to take advantage of lower off-peak rates.

[0053] The proposed technology would hold several competitive advantages over existing solutions in smart grid management: (a) Advanced Data Analytics: Large language models (LLMs) like GPT are adept at handling complex and multifaceted data sets. This allows for a Attorney Docket No.11605-066WO1 FAU 202408 more comprehensive analysis of smart grid data, leading to deeper insights and more accurate predictions compared to traditional grid management systems. (b) Real-Time Processing: The ability to process data in real-time is crucial for effective grid management. GPT-based systems can continuously analyze streaming data from sensors, enabling them to identify and respond to grid issues promptly. This is a significant edge over conventional systems that rely on periodic data collection and analysis. (c) Unstructured Data Integration: The grid generates a vast amount of unstructured data, including sensor readings, maintenance logs, and weather data. LLMs excel at handling unstructured data, allowing them to extract valuable insights that may be missed by conventional grid management systems focused on structured data formats. (d) Scalability and Flexibility: A GPT-based system can be readily scaled to accommodate the growing complexity of the grid as more renewable energy sources and distributed generation come online. Additionally, the inherent flexibility of LLMs allows them to adapt to evolving grid conditions and integrate with new technologies seamlessly.

[0054] The above-noted advantages translate into at least the following practical benefits: (a) Reduced Downtime and Improved Reliability: Real-time transformer condition monitoring and predictive maintenance capabilities can significantly reduce the risk of unexpected failures and associated downtime. This translates to a more reliable grid with fewer outages. (b) Enhanced Grid Resilience: The ability to analyze grid data in real time and predict potential issues allows for proactive measures to be taken, improving the grid’s resilience against extreme weather events or cyberattacks. (c) Optimized Energy Delivery and Lower Costs: Accurate demand forecasting and improved grid management can lead to more efficient energy distribution, reducing overall costs for utilities and consumers. (d) Faster Integration of Renewable Energy: The ability to handle the variable nature of renewable energy sources like solar and wind paves the way for smoother integration into the grid, facilitating the transition towards a more sustainable energy future.

[0055] In summary, some of the advantages of the proposed system lie in its superior data processing capabilities, real-time responsiveness, and adaptability. This translates into a more efficient, reliable, and resilient grid, with significant benefits for utilities, businesses, and consumers alike.

[0056] Embodiments of the present disclosure leverage a Transformer-based architecture (e.g., a state-of-the-art deep learning architecture) and Large-Scale Foundation (LSF)-Models for prognostic health management of smart grid power transformers. This architecture and system perform complex classification predictions and provide users with an appropriately Attorney Docket No.11605-066WO1 FAU 202408 predicted response. The advanced natural language processing techniques of the Transformer architecture interpret the nuances of power transformer-related data and efficiently extract valuable insights for fault detection and diagnosis. The multiple self-attention mechanisms of the architecture provide an enhanced form of feature extraction which expedites the creation of comprehensive models for large multiclass datasets. During testing, this fault classification method performed with 97.2% accuracy according to the Matthews Correlation Coefficient when evaluating 45 classes of simulated power transformer internal faults and external transient disturbances.

[0057] To better represent the diverse multimodal data that compose smart grid systems, this disclosure explores a unified system that integrates multiple machine-learning models into one simple and easy-to-use interface. The backend of this interface again takes advantage of the Transformer architecture to perform conversation-based classification and to provide a response or prediction from the appropriate system-integrated model. The success of the Transformer architecture in diverse applications within the same overall system showcases its potential to analyze the wide range of data typically found throughout a robust smart grid system. Example System

[0058] FIG.1 depicts an example system 100 (e.g., Power GPT system) in accordance with certain embodiments described herein. The system 100 includes a data engine 102, a cloud- based web service 107, and one or more user computing devices 115. Each system component can include or comprise at least one computing device 600 described below in connection with FIG.6.

[0059] In various implementations, DevOps 103 practices can be used to manage and deploy machine learning models 109 (e.g., Model A, Model B…Model Z) via a smart grid management API for smart / power grid data management, integrating data preprocessing, machine learning model training, and prediction services. In some implementations, as illustrated, users interact through one or more user computing devices 115 that each include a graphical user interface (GUI), sending data and chat queries, with outputs hosted via the cloud- based web service 107 (e.g., Amazon Elastic Compute Cloud (EC2) and managed via an identifier assigned to a URL route (e.g., Flask endpoint). The cloud-based web service 107 can include or be operatively coupled to a first database 110 (e.g., PostgreSQL database) storing, for example, chat history and other relevant data. As illustrated, the system 100 / GUI 115 can be accessible via a subscription service. Attorney Docket No.11605-066WO1 FAU 202408

[0060] As shown, the data engine 102 comprises a second database 104 storing smart grid data / power grid data. The data engine 102 includes a data-processing component 105 configured to pre-process power grid data obtained from the second database 104. The pre- processing step can comprise obtaining and conditioning (e.g., filtering, selecting) smart grid sensor data obtained from one or more smart grid systems / energy system components (e.g. transformers). The data engine 102 includes a model training and testing component 106 configured to use the pre-processed data to train and / or test one or more test machine learning models. Subsequent to training / testing, the data engine 102 can deploy one or more trained prediction models 108 for use. Example Method

[0061] FIG.2A and FIG.2B are flowcharts of example methods 200, 250 (e.g., processes, computer-implemented methods) in accordance with certain embodiments described herein. In some implementations, the methods 200, 250 are performed by a processing circuitry (for example, but not limited to, an application-specific integrated circuit (ASIC), or a central processing unit (CPU)). In some examples, the processing circuitry may be electrically coupled to and / or in electronic communication with other circuitries of an example computing device, such as, but not limited to, the example computing device 600 described below in connection with FIG.6. In some examples, embodiments may take the form of a computer program product on a non-transitory computer-readable storage medium storing computer-readable program instruction (e.g., computer software). Any suitable computer-readable storage medium may be utilized, including non-transitory hard disks, CD- ROMs, flash memory, optical storage devices, or magnetic storage devices. This disclosure contemplates that each method 200, 250 can be partially or fully implemented using the system 100 of FIG.1.

[0062] With reference to FIG.2A, at step / operation 210, the method 200 includes analyzing, using a classification or large language model, energy system data to identify one or more faults and / or data for visualization and interpretation for decision-making. Step / operation 210 can also include receiving (retrieving, obtaining) the energy system data, for example as a signal file or from one or more databases. In some embodiments, the classification or large language model is a generative Transformer-based model. The classification or large language model can be seamlessly integrated into a backend API. In some implementations, the energy system data comprises real-time data obtained from sensors operatively coupled with energy system components, and the system 100 is Attorney Docket No.11605-066WO1 FAU 202408 configured to continuously analyze streaming data from the sensors in order to promptly response to grid issues.

[0063] At step / operation 220, the method 200 includes, responsive to receiving a direct user input (e.g., in a language format), determining, using a second model, one or more user input parameters. In some implementations, the second model is an NLP model and can comprise a chat-based model or model architecture. The classification model or large language model can be configured to receive input data comprising the direct user input and a signal file including power grid data, organize the input data into data matrices, and classify / identify the one or more faults using a series of dense ReLU layers and dropouts. In some embodiments, the second model (e.g., NLP model) uses the data matrices containing the one or more classified faults as input data and employs one or more databases or libraries to process the received data matrices.

[0064] At step / operation 230, the method 200 includes determining, based, at least in part, on one or more word-based tokens generated from the user input and the one or more user input parameters, at least one of a plurality of models. In some embodiments, the plurality of models includes a deep neural network model and / or large language model. In some implementations, the method 200 includes generating the plurality of models including the second model and training the plurality of models using historical energy system data (e.g., at or prior to step / operation 210). In some examples, the method 200 includes pre- processing the historical energy system data prior to training the plurality of models. The historical energy system data can be retrieved from one or more publicly available databases.

[0065] At step / operation 240, the method 200 includes generating, using at least one of the plurality of models, one or more visualization outputs based, at least in part, on the one or more user input parameters and / or the one or more identified faults. In some embodiments, generating the one or more visualized outputs includes directing the energy system data to at least one of the plurality of models for analysis or processing.

[0066] Optionally, at step / operation 240, the method 200 can include outputting, via a display or graphical user interface, the one or more visualization outputs as conversational text, graphs, and / or charts.

[0067] Referring now to FIG.2B, beginning at step / operation 260 the exemplary method 250 includes analyzing, using a classification model, energy system data to detect one or more fault classification outputs. Step / operation 260 can also include receiving Attorney Docket No.11605-066WO1 FAU 202408 (retrieving, obtaining) the energy system data, for example as a signal file or from one or more databases.

[0068] At step / operation 270, the method 250 includes, responsive to receiving a direct user input, determining, using a second model, one or more user input parameters.

[0069] At step / operation 280, the method 250 includes determining, based at least in part, on one or more fault characteristics associated with the one or more fault classification outputs and the one or more user input parameters, at least one of a plurality of models. In some implementations, the method 250 includes generating the plurality of models including the second model and training the plurality of models using historical energy system data (e.g., at or prior to step / operation 260).

[0070] At step / operation 290, the method 250 includes generating, using at least one of a plurality of models, one or more predictive outputs based, at least in part, on the one or more fault characteristics and the one or more user input parameters. The one or more predictive outputs and / or the one or more user input parameters can include at least one of a location, a fault type, an energy system / component, and / or target power transformer(s) that are associated with a fault or likely to fail.

[0071] Optionally, at step / operation 295, the method 250 includes outputting, via a display or GUI, the one or more predictive outputs, and / or one or more fault classification outputs / characteristics.

[0072] Additionally and / or alternatively, at step / operation 297, the method 250 includes triggering or causing a corrective action with respect to at least one identified fault classification output (e.g., generate an alert, deploy personnel to a target location, and / or automatically activate or deactivate a system component).

[0073] With reference to FIG.3A and FIG.3B, the exemplary system 100 can provide two use cases: an user-privileged use case and admin-privileged use case. FIGS.3A – 3B each show an example operation flow / flowchart 300a, 300b for the user-privileged use case and the admin-privileged use case, respectively. A generalized operation flow for both use cases of the exemplary system is further detailed in FIG.4. This disclosure contemplates that each operation flow 300a, 300b can be partially or fully performed using the system 100 described in connection with FIG.1.

[0074] FIG.3A is the operation flow 300a (e.g., process, method) for the user- privileged use case. At step / operation 302, a user can log into the exemplary system 100, for example via the user computing device 115 described in connection with FIG.1. At Attorney Docket No.11605-066WO1 FAU 202408 step / operation 304, the user retrieves an application programming interface (API). At step / operation 306, the user proceeds to open a graphical user interface (GUI). At step / operation 308, the user can request the exemplary system 100 to find a power transformer fault. Additionally, and / or alternatively, at step / operation 310, the user can estimate the remaining useful life (RUL) of a power grid. At step / operation 312, the user inputs a user prompt. At step / operation 314, the user uploads a signal file via the GUI (e.g., GUI chatbot). At step / operation 316, the user prompt and the signal file is processed by a classification model and / or a natural language processor (e.g., 108, 109). At step / operation 318, the system 100 generates outputs such as diagnostic information, natural language responses, prognostic information, and / or graphs.

[0075] FIG.3B is the operation flow 300b (e.g., process, method) for the admin- privileged use case. As shown, at step / operation 320, the user (admin) logs into the exemplary system 100 using their administrator access. At step / operation 321, the system authenticates the user to confirm that the user has administrative access. At step / operation 322 and step / operation 324, the user can modify source code, train new classification or NLP models, or both. At step / operation 326, the user can then validate new code and models of the exemplary system locally (e.g., in a virtual / local / offline environment). After the validation, at step / operation 328, the user can push the changes to the code or the classification / NLP models to a database storage (shown as 407A in FIG.4) (e.g., local disk drive, GitHub, Docker Hub) operatively coupled to the exemplary system 100. Then, at step / operation 330, the user can deploy an instance (e.g., containerized instance) of the exemplary system to a remote server (e.g., cloud-based web service 107, Amazon EC2) for mass production use. Example Transformer-based Models for Exemplary System

[0076] In some embodiments, the exemplary system can employ state-of-the-art machine learning (ML) models (e.g., random forest, gradient boost, vector classification) to extract a discrete wavelet transform (DWT) of a power signal (e.g., three-phase power) in a power grid and detect abnormalities in the power signal by performing fault classification on the extracted DWT. The three state-of-the-art ML models can classify a subset of the faults with an accuracy of 90% on average, indicating that pairing the full dataset with a multiclass classification ML model may provide better results in practical applications.

[0077] The limited efficiency and precision of state-of-the-art ML models (e.g., random forest, gradient boost, vector classification) can restrict the exemplary system from classifying more than four fault classes at a time. To handle large multiclass classification Attorney Docket No.11605-066WO1 FAU 202408 problems (e.g., tens of classes of power transformer faults and transient disturbances), the exemplary system can employ a Transformer-based large language model (LLM) (e.g., PowerGPT, ChatGPT, Transformer).

[0078] FIG.4 is an example structure of the exemplary system comprising a Transformer-based fault classification model 406 and a natural language processor 408 (e.g., NLP model), wherein the classification model 406 is configured to classify faults in the smart grid system and the NLP model 408 is configured to output responses and information from the classified faults. As shown, the exemplary system can receive signals 404 (e.g., three- phase signals) from a smart grid system 402 as input data (e.g., file format) to a Transformer- based fault classification model 406. The classification model 406 can then organize the input data into data matrices and classify the faults using a series of dense rectified linear unit (ReLU) layers and dropouts. The data matrices containing the classified faults can then be transmitted to a natural language processor 408 (e.g., NLP model) as input data, wherein the NLP model employs one or more JSON libraries to process data matrices received from the classification model 406.

[0079] The NLP model 408 can vectorize the text data in the data matrices into a sequence of numbers, split the sequence into lists of word-based tokens, and reorganize the lists into a matrix for generating responses 410 or diagnostic information 412. In some embodiments, the NLP model is operatively coupled with a graphical user interface (GUI) configured to store historical data 407a (e.g., past responses, diagnostic information) and user inputs 407b.

[0080] In some embodiments, the exemplary system further comprises one or more additional prognostic models 409 configured to receive multimodal data 403, generate data matrices of classified faults, and then transmit the data matrices to the NLP model 408, along with the classification model 406, for further processing.

[0081] The library can be processed into organized lists that are then labeled for training. The text data (from data file received from power grid) can be vectorized, by converting the text into integers and removing all punctuation, to generate a sequence of numbers, which can be split into lists of word-based tokens to retain relational components. The lists of word-based tokens can then be reorganized into an m × n matrix (e.g., n=20) for m total pattern entries.

[0082] Transformer-Based Fault Classification Model. The exemplary system can employ a Transformer-based multiclass classification model (e.g., LLM) (shown as 406 in Attorney Docket No.11605-066WO1 FAU 202408 FIG.4), compatible with various data formats

[0018] , to monitor the conditions of power grids and transformers. The classification model can employ a comma-delimited text file data processing method to configure a data file (received from power grids) into data matrices (e.g., 726 × 3 Numpy arrays) for current measurements (e.g., three-phase differential current measurements). Since the number of fault conditions varies between model classes, the training data for the classification model can be equalized using class normalization methods. The classification model can struggle to improve when trained solely on the normalized data, so the first epoch of the classification model should be trained on standard data, and then the remaining epochs should be trained with the proper normalized data. When the classification model is trained with an appropriate number of total epochs, the inconsistency of data normalization in training data can have a negligible impact on overall model accuracy.

[0083] The classification model can be initialized by instantiating a tensor (e.g., Keras tensor) with the same dimensions as the restructured data files (e.g., data file restructured as 726 × 3 Numpy arrays). After additional normalization, the tensor can be reshaped to incorporate a batch size dimension and then flattened with respect to each time step. Some dense layers, for example, first dense layer (e.g., of dimension 1024) and second dense layer (e.g., of dimension 64), of the classification model can be applied and followed by a dropout rate (e.g., of 20%), wherein dropout is a regularization method to prevent data overfitting during a LLM training. The rectified linear unit (ReLU) activation function

[0019] can be used for all intermediate dense layers of the classification model, as shown in Equation 1. ^= ^^^^ = ^^, ^ > 00, ^ ≤ 0(Eq.1)

[0084] The process of positional embedding can be applied using the batch size as the embedding vocabulary size and the dense layer dimension as the dense embedding output dimension. The body of the classification model can comprise a plurality (e.g., stack of 4) of Transformer encoders, wherein each encoder can perform layer normalization on the input tensor and sequentially perform multi-headed attention that consists of a plurality (e.g., 4) of attention heads (e.g., of size 64) with a dropout probability (e.g., of 20%). With dk as the dimension of keys K, the scaled dot-product self-attention mechanism, at a layer of the classification model, can be defined as a function of the queries Q, keys K, and values V per Equation 2. ^^^ ^^^^^^^ ^^^ ^^ Attorney Docket No.11605-066WO1 FAU 202408 (Eq.2)

[0085] A residual connection at a layer can be established by adding the output of the self-attention mechanism to the original normalized tensor. After additional normalization, a feedforward neural network (FNN) can be applied, which comprises several iterations of the ReLU activation function dense layers and dropouts. A final residual connection can be established between the input and output of the FFN.

[0086] Following the data preprocessing stage, the classification model can start the fault classification process. The data from the stacked encoder of the classification model can be flattened and processed through the following series of dense ReLU layers and dropouts: dense (e.g., 256), dropout (e.g., 20%), dense (e.g., 128), dense (e.g., 32), and dropout (e.g., 20%). Including frequent dropout layers can keep the model from overfitting the data, ensuring future generalizability of the classification model. The final dense layer dimension can equal the total number of evaluated classes. The layer can also apply a softmax activation function, which normalizes the data into representative probabilities. The classification model can use the Adam optimizer

[0020] , an extended version of stochastic gradient descent optimization, with a learning rate α of 0.001. After some epochs (e.g., 150), training checkpoints, training history, and the trained classification model can be saved in corresponding file formats (e.g., training checkpoints as an H5 file, training history as a Pickle file, and trained model as a Keras file)

[0087] Natural Language Processor or model. In addition to using the Transformer- based classification model for fault classification (e.g., of three-phase time-series signals), the exemplary system also employs a natural language processor (e.g., NLP model) (shown as 408 in FIG.4). In some embodiments, the NLP is supported by a JSON library that stores topics of conversation as distinct classes, where each class contains a tag, a list of patterns, and a list of responses, and the pattern list can contain words or phrases associated with the corresponding class. The library can be processed into organized lists that are then labeled for training. The text data (from data file received from power grid) can be vectorized, by converting the text into integers and removing all punctuation, to generate a sequence of numbers, which can be split into lists of word-based tokens to retain relational components. The lists of word-based tokens can then be reorganized into an m × n matrix (e.g., n=20) for m total pattern entries.

[0088] The NLP model first utilizes an embedding layer (e.g., with a vocabulary size of 1000 and a dense embedding dimension of 32) to keep the length of input sequences Attorney Docket No.11605-066WO1 FAU 202408 constant (e.g., at 20). A one-dimensional global average pooling operation layer can be added to the NLP model to map the features of the embedding layer. A plurality (e.g., 2) of dense ReLU activation function layers (e.g., of dimension 16) can also be added sequentially to the NLP model. The final dense layer of the NLP model can contain the same number of nodes as the total evaluated classes and apply the softmax activation function to obtain the distribution of probabilities. The NLP model can use sequential grouping to create a complete model object, which can be further compiled using the Adam optimizer. The exemplary system can run the NLP model for many epochs (e.g., 550 epochs) and save the NLP model as an H5 file upon completion. Experimental Results and Additional Examples

[0089] A study was conducted to develop the exemplary systems, methods, and a software system for proactive data management of energy systems (e.g., smart grid power transformers, wind farms, and power plant asset mapping), featuring an interactive data dashboard enhanced with natural language processing (NLP).

[0090] Dataset Preparation and Validation. Before evaluating the exemplary system and its associated models, the study evaluated a subset of a three-phase time-series signal- based dataset

[0015] assembled by researchers at Syracuse University. This dataset recorded differential current as a function of time for a 5-bus system simulated using PSCAD / EMTDC software.100,908 transient cases were simulated by modifying various system parameters. Differential current measurements for each transient example were saved in comma-delimited plain text files. Measurements were taken every 100 microseconds for 0.0726 seconds to produce 726 data points for each of the three phases. In other words, the raw data was organized as a 726 × 4 matrix with the first column as time and columns two, three, and four as the differential current of phases A, B, and C, respectively. FIG.5A shows this dataset with a total of 45 fault classes requiring a multiclass classification solution.

[0091] Three types of transformers were observed in this simulation: power transformers, Indirect Symmetrical Phase Angle Regulator (ISPAR) series transformers, and ISPAR exciting transformers.13 internal faults were observed for each of the three transformer types. In FIG.5A, the first 11 faults were labeled “Class1” through “Class11” and represented the following transformer winding locations: phase A to ground, B to ground, C to ground, A to B to ground, A to C to ground, B to C to ground, A to B to C to ground, phase A to phase B, A to C, B to C, and A to B to C. The twelfth and thirteenth classes for each transformer type reflected turn-to-turn and winding-to-winding fault Attorney Docket No.11605-066WO1 FAU 202408 locations, respectively. Six additional external transient disturbances were observed as individual fault classes: capacitor switching, external faults with current transformer (CT) saturation, ferro resonance, magnetizing in-rush, non-linear load switching, and sympathetic inrush. The 13 internal faults for each of the three transformer types, along with the six additional external transient disturbances, composed the 45 total fault classes.

[0092] The study verified the validity of this dataset through its use as training data for multiclass classification ML techniques, the workflows of which were based on analyses of various ML classifiers for signal classification using Discrete Wavelet Transform (DWT) decomposition for feature extraction

[0016] , wherein DWT involved the convolution of a discrete signal and a preselected mother wavelet. The two signals were multiplied at increasing distances from the initial position of the discrete signal. Various families of mother wavelets were applied in various individual patterns to find the most optimal comparison for a particular application and signal type. The convolution output passed through a series of high- and low-pass filters, which produced waveforms of detail and approximation coefficients, respectively. The waveforms of these coefficients represented the high-and low- frequency components of the original signal as a function of time. The abnormal frequencies present in signal anomalies were exaggerated in these decomposed forms and can be detected and categorized during signal analysis processes.

[0093] The study compared the features of each decomposed detail coefficient level and the last approximate coefficient level with the corresponding levels of other signals. The study calculated twelve statistical features, including entropy, 5th percentile, 25th percentile, 75th percentile, 95th percentile, median, mean, standard deviation, variance, and root mean square, for each decomposition level and phase. For example, a level-five decomposition analysis evaluated twelve statistical features for six sets of DWT decomposition coefficients for three phases of data.216 features, in this example, were extracted per data file in the randomly selected training set. For all training instances, the number of decomposition levels was logistically related to model accuracy while maintaining a linear relationship to the number of features and overall training time. Therefore, due to the diminishing returns of the asymptotic logistic relationship, only five decomposition levels were used for the ML-based dataset validation.

[0094] Fault Classification Model Performance. The study developed a predictive Transformer-based fault classification model to analyze three-phase signals recorded over time for feature extraction and classification. The classification model was configured to Attorney Docket No.11605-066WO1 FAU 202408 detect and classify faults among 45 different power transformer fault types using a dataset of over 100,000 simulated examples

[0015] . The classification model achieved a Matthew’s Correlation Coefficient accuracy score of 97.2% in classifying the 45 classes of fault signals.

[0095] FIG.5B shows classification matrices generated by two fault classification models, e.g., Indirect Symmetrical Phase Angle Regulator (ISPAR) series transformer (subpanel (a)) and power transform (subpanel (b)). In FIG.5B, a specific handful of classes in the matrices were more prone to error than others, including the exciting turn-to-turn (exciting-tt), exciting winding-to-winding (exciting-ww), transformer-Class1, transformer-tt, magnetic inrush, and series-tt classes. Among the exciting categories, testing samples from exciting-tt and exciting-ww were classified as exciting-Class1 and exciting-Class2, and vice versa. Additional outliers included the erroneous classification of magnetic inrush as sympathetic inrush. The most prominent outlier from the testing data was the misclassification of transformer-tt as transformer-Class1. The most inaccurate predictions stemmed from the exciting-tt and transformer-tt classes, which exhibited a total misclassification rate of 65%. From FIG.5B, the study demonstrated the feasibility of classifying large-scale phase signals while recognizing the presence of certain outlier classes. Viewing the concept of predictability and fault classification through the Transformer-based classification model demonstrated its potential for contributing to a largely accurate mechanism for reliable multiclass classification of three-phase time-series signals found in smart grid systems.

[0096] Features of Graphical User Interface. To demonstrate the diagnostic capabilities of the Transformer-based classification model, the study developed a Linux- based graphical user interface (GUI). FIG.5C shows a graphical user interface (GUI) configured to perform and visualize the fault diagnosis of the classification model. The GUI's home page enabled dynamic interaction with a chatbot. Table 1 shows the key components of the GUI shown in FIG.5C. Table 1 GUI component Description he he in es Attorney Docket No.11605-066WO1 FAU 202408 History Log (not shown) History log records conversations and corresponding graphs for future review. h the chatbot through the GUI for fault diagnosis. The study uploaded the signal files to the exemplary system via the GUI, and the classification model of the exemplary system analyzed and classified the data. Table 2 shows the responses of the GUI chatbot to the user queries and instructions. Table 2 User queries GUI chatbot response User greeting and request The GUI chatbot greeted users, initiated interactions, and responded to e, k. he se l. nd nd ly he es he as

[0098] Discussion #1. In 2022, the United States utilized 4.05 trillion kilowatt-hours of electricity to support residential, commercial, and industrial sectors [1]. With technologies connecting to the grid daily, various energy consumption has strained the electrical grid system. The transmission of electrical energy from current state-of-the-art distribution systems is only possible due to the proper maintenance of electrical power transformers. These devices step up or down voltage according to physical properties that use electromagnetic induction. This process, however, dissipates energy to the environment in the form of heat, thereby exposing the power transformer to thermal malfunctions. The general wear and tear built up throughout the lifespan of a power transformer contributes to transient disturbances and mechanical failures. Not only does this compromise the mechanical design of the power transformer, but it also affects the stability of its electrical output. Anomalous voltage drops, voltage rises, and Attorney Docket No.11605-066WO1 FAU 202408 phase shifts all impact the overall reliability of the electrical grid and can be damaging to connected devices. Therefore, a power transformer operating with untreated mechanical faults becomes a weak link that, if left untreated, disrupts the chain of power distribution to industrial, commercial, and residential sectors.

[0099] The rise in electrical power dependence requires an improvement in reliability. Because internal faults and external transient disturbances also impact the electrical signal, modern sensors incorporated into smart grid systems can be used to collect signal data and alert maintenance crews to defective power transformers. Although contemporary systems have shown results in detecting faults [2], classifying those detections poses a computational challenge. The amount of data needed to construct a fault classification model for an entire smart grid system requires machine learning (ML) models with the capability of handling extensive datasets, such as large-scale foundation models (LSF-Models) (e.g., GPT-3.5 (ChatGPT) [3] and the Segment Anything Model (SAM) [4]). Applying an LSF model for smart grid system condition monitoring can be a solution that provides corroborated information by utilizing multiple models and simulations.

[0100] The success of many LSF-Models can be attributed to the innovative Transformer deep learning model [5]. The Transformer model has become a popular foundation for many natural language processing (NLP) and computer vision (CV) models in the few years since its infancy in 2017. The Transformer model’s recent success in the form of ChatGPT proves its capability of handling large multiclass classification problems. The instant study developed an exemplary system that employs a Transformer-based model to classify signals in smart grids and provide a conversational prediction similar to ChatGPT. Instead of processing global web data, this fault classification model (of the exemplary system) was trained on anomalous time- series signals to provide a fault classification prediction. The classification model was integrated with a basic natural language processor, which also used a simplified Transformer- based structure to constitute the back end of the exemplary system’s GUI.

[0101] Discussion #2. The exemplary system employs a Transformer-based model configured to analyze signals (e.g., three-phase) for power transformer fault classification. The classification model’s accuracy of 97.2% in identifying fault types highlighted its efficacy. Using multiclass classification (MCC), the results probed further discussion of finer metrics that delved deeper into the classification model’s performance, which aligned with an improvement from the initial machine-learning models. The results were promising, with most confusion matrix aligning accurately with expectations. Attorney Docket No.11605-066WO1 FAU 202408

[0102] The high overall accuracy, although impressive, was tempered by the challenges posed by outlier classes. The transformer and exciting classes exhibited inconsistencies in classification that required further investigation into their outliers and erroneous patterns. Although misclassifications existed, they served as insights into areas where the classification model’s sensitivity can be refined. The study emphasized that while automated fault classification was obtainable, continuous refinement was crucial to address variations in fault behavior.

[0103] The development of a Linux-based GUI further enhanced the classification model's accessibility. The GUI provided a user-friendly platform for dynamic interaction and fault diagnosis, amplifying the model’s applicability to real-world scenarios.

[0104] In future studies, other enhancements to this interface can prove crucial to improving the predictive Transformer-based model. These potential improvements can include: 1) Expanding the framework to cross-reference historical prognostic databases, enhancing fault classification accuracy and condition-based maintenance for improved power grid reliability; 2) Integrating real-time online sensor condition monitoring to provide continuous diagnostics and proactive maintenance insights; 3) Additionally, cybersecurity measures such as secure password protection, secure API integrations, and role-based access control (RBAC) may strengthen data protection; and 4) Collaborations with utility companies for industry-aligned development and pilot testing may further validate the platform’s effectiveness through real-world grid simulations.

[0105] In conclusion, the study contributed to advancing smart grid data analysis and management, focusing on integrating various features that support the overall functionality of smart grid systems into a single robust platform. The study of prognostics and health management and fault diagnosis for power transformers highlighted one of the key applications, while the exemplary system also included capabilities like GridWatch for visualizing power assets, wind power prediction models for managing renewable energy, and interactive dashboards for comprehensive grid insights. The exemplary system’s artificial intelligence (AI) features, such as visual fault diagrams and chatbot-driven explanations, provided operators with actionable insights. The GUI of the exemplary system served as a proof-of-concept, demonstrating how a unified user interface can enhance the reliability and performance of multiple grid assets, providing a holistic approach to smart grid management.

[0106] Discussion #3. The popularity, efficiency, and universality of LSF-Models have all expanded. Foundation models attempt to encompass multiple facets of data within a particular Attorney Docket No.11605-066WO1 FAU 202408 network while remaining applicable to multiple tasks within the same sector. Models such as Bidi-rectional Encoder Representation from Transformers (BERT)[6], Enhanced Representation through kNowledge IntEgration (ERNIE) [7], and Large Language Model Meta AI (LLaMA)[8] have showcased approaches to large-scale NLP applications. Additionally, the GPT series has pushed the boundaries of the Transformer architecture by developing increasingly complex optimization structures [9]

[0010] .

[0107] Despite the historic success of NLP and CV LSF models, complex smart grid systems require flexible deep learning models that can interpret several data modes, including signals, images, videos, and text. Even the textual mode requires the analysis of various data structures such as maintenance records, ongoing maintenance work orders, and culminating project reports. Assistive systems, therefore, require a deep integration of multiple ML models. Previous studies into such a capable multimodal model are still behind comparable designs from the NLP and CV fields

[0011] . The majority of current state-of-the-art solutions (to the smart grid systems) include wavelet-based convolutional neural networks (CNNs)

[0012]

[0013] and recurrent neural networks (RNNs)

[0014] , which both contain disadvantages within smart grid system applications.

[0108] The success of the Transformer-based architecture / model for large text-based multiclass classification (e.g., ChatGPT) prompted the search for a custom architecture with accurate signal-based classification. At the beginning of 2023, researchers constructed two models based on Transformer architecture. The research compared a basic “vanilla” Transformer to an innovative Differential Architecture Search (DARTS) algorithm, which produced accurate results for detecting 23 fault types and 15 fault locations for power transformers

[0018] . The baseline Transformer-based model preprocessed the input data by equalizing the samples per class and converting the arrays to TensorFlow tensors. The preprocessed data can be transformed into attention embeddings through several levels of multi-head attention mechanisms and feedforward networks. These methods analyze a series of ‘queries’, ‘keys’, and ‘values’ to calculate the relative connections between individual data tokens. This process mirrors the structure of the foundational Transformer architecture encoder. Although being a more efficient method for deep data analysis than RNNs, the mathematical processes of the attention encoder are not time-dependent. To retain time-dependent information, the Transformer-based model also applies positional embeddings to the data, reconstructing the context of the data directly after the completion of the main attention encoder process. Attorney Docket No.11605-066WO1 FAU 202408 Machine Learning

[0109] In some embodiments, the exemplary system can be implemented using one or more artificial intelligence and machine learning operations. The term “artificial intelligence” can include any technique that enables one or more computing devices or computing systems (i.e., a machine) to mimic human intelligence. Artificial intelligence (AI) includes but is not limited to knowledge bases, machine learning, representation learning, and deep learning. The term “machine learning” is defined herein to be a subset of AI that enables a machine to acquire knowledge by extracting patterns from raw data. Machine learning techniques include, but are not limited to, Transformer-based models (e.g., Bidirectional Encoder Representations from Transformers (BERT), Naïve Bayes classifiers, and artificial neural networks. The term “representation learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, or classification from raw data. Representation learning techniques include, but are not limited to, autoencoders and embeddings. The term “deep learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, classification, etc., using layers of processing. Deep learning techniques include but are not limited to artificial neural networks or multilayer perceptron.

[0110] Machine learning models include supervised, semi-supervised, and unsupervised learning models. In a supervised learning model, the model learns a function that maps an input (also known as feature or features) to an output (also known as target) during training with a labeled data set (or dataset). In an unsupervised learning model, the algorithm discovers patterns among data. In a semi-supervised model, the model learns a function that maps an input (also known as feature or features) to an output (also known as a target) during training with both labeled and unlabeled data.

[0111] Neural Networks. An artificial neural network (ANN) is a computing system including a plurality of interconnected neurons (e.g., also referred to as “nodes”). This disclosure contemplates that the nodes can be implemented using a computing device (e.g., a processing unit and memory as described herein). The nodes can be arranged in a plurality of layers such as input layer, an output layer, and optionally one or more hidden layers with different activation functions. An ANN having hidden layers can be referred to as a deep neural network or multilayer perceptron. Each node is connected to one or more other nodes in the ANN. For example, each layer is made of a plurality of nodes, where each node is connected to all nodes in the previous layer. The nodes in a given layer are not interconnected with one Attorney Docket No.11605-066WO1 FAU 202408 another, i.e., the nodes in a given layer function independently of one another. As used herein, nodes in the input layer receive data from outside of the ANN, nodes in the hidden layer(s) modify the data between the input and output layers, and nodes in the output layer provide the results. Each node is configured to receive an input, implement an activation function (e.g., binary step, linear, sigmoid, tanh, or rectified linear unit function), and provide an output in accordance with the activation function.

[0112] Additionally, each node is associated with a respective weight. ANNs are trained with a dataset to maximize or minimize an objective function. In some implementations, the objective function is a cost function, which is a measure of the ANN’s performance (e.g., error such as L1 or L2 loss) during training, and the training algorithm tunes the node weights and / or bias to minimize the cost function. This disclosure contemplates that any algorithm that finds the maximum or minimum of the objective function can be used for training the ANN. Training algorithms for ANNs include but are not limited to backpropagation. It should be understood that an artificial neural network is provided only as an example machine learning model. This disclosure contemplates that the machine learning model can be any supervised learning model, semi-supervised learning model, or unsupervised learning model. Optionally, the machine learning model is a deep learning model. Machine learning models are known in the art and are therefore not described in further detail herein.

[0113] A Transformer model is a type of deep neural network that is configured to process sequential data such as natural language inputs. Transformer models generally include a self- attention mechanism that facilitates weighing the importance of words within a sentence in relation to one another which in turn facilitates determining dependencies and context in a superior fashion. A Transformer model comprises an encoder-decoder structure where an encoder processes an input sequence and transforms it into an abstract representation, and the decoder is configured to generate an output based on the abstract representation. The ability to parallelize data also speeds up model training and makes it possible to train large models on big data sets.

[0114] In the context of Large Language Models (LLMs), a token is a fundamental unit of text that the model processes, representing a word, part of a word, or a symbol, and is used to efficiently handle and process language data.

[0115] Other Supervised Learning Models. A logistic regression (LR) classifier is a supervised classification model that uses the logistic function to predict the probability of a target, which can be used for classification. LR classifiers are trained with a data set (also Attorney Docket No.11605-066WO1 FAU 202408 referred to herein as a “dataset”) to maximize or minimize an objective function, for example, a measure of the LR classifier’s performance (e.g., error such as L1 or L2 loss), during training. This disclosure contemplates that any algorithm that finds the minimum of the cost function can be used. LR classifiers are known in the art and are therefore not described in further detail herein.

[0116] A Naïve Bayes’ (NB) classifier is a supervised classification model that is based on Bayes’ Theorem, which assumes independence among features (i.e., the presence of one feature in a class is unrelated to the presence of any other features). NB classifiers are trained with a data set by computing the conditional probability distribution of each feature given a label and applying Bayes’ Theorem to compute the conditional probability distribution of a label given an observation. NB classifiers are known in the art and are therefore not described in further detail herein.

[0117] A k-NN classifier is an unsupervised classification model that classifies new data points based on similarity measures (e.g., distance functions). The k-NN classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize a measure of the k-NN classifier’s performance during training. This disclosure contemplates any algorithm that finds the maximum or minimum. The k-NN classifiers are known in the art and are therefore not described in further detail herein. Computing Devices and Methods of Use

[0118] It should be appreciated that the logical operations described herein with respect to the various figures may be implemented (1) as a sequence of computer-implemented acts or program modules (i.e., software) running on a computing device, (2) as interconnected machine logic circuits or circuit modules (i.e., hardware) within the computing device and / or (3) a combination of software and hardware of the computing device. Thus, the logical operations discussed herein are not limited to any specific combination of hardware and software. The implementation is a matter of choice dependent on the performance and other requirements of the computing device. Accordingly, the logical operations described herein are referred to variously as operations, structural devices, acts, or modules. These operations, structural devices, acts and modules may be implemented in software, in firmware, in special-purpose digital logic, and any combination thereof. It should also be appreciated that more or fewer operations may be performed than shown in the figures and described herein. These operations may also be performed in a different order than those described herein. Attorney Docket No.11605-066WO1 FAU 202408

[0119] It should be understood that the computing device is only one example of a suitable computing environment upon which embodiments of the present disclosure may be implemented. Optionally, the computing device can be a well-known computing system including, but not limited to, personal computers, servers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, personal network computers (PCs), minicomputers, mainframe computers, embedded systems, and / or distributed computing environments including a plurality of any of the above systems or devices. Distributed computing environments enable remote computing devices, which are connected to a communication network or other data transmission medium, to perform various tasks. In the distributed computing environment, the program modules, applications, and other data may be stored on local and / or remote computer storage media.

[0120] FIG. 6 shows an example computing device 600 having a basic configuration. As shown, in its basic configuration, the computing device 600 includes at least one processing unit 606 and system memory 604. Depending on the exact configuration and type of computing device, system memory 604 may be volatile (such as random-access memory (RAM)), non- volatile (such as read-only memory (ROM), flash memory, etc.), or some combination of the two. The processing unit 606 may be a standard programmable processor that performs arithmetic and logic operations necessary for the operation of the computing device. The computing device 600 may also include a bus or other communication mechanism for communicating information among various components of the computing device.

[0121] Computing device 600 may have additional features / functionality. For example, the computing device may include additional storage such as removable storage 608 and non- removable 610 storage, including, but not limited to, magnetic or optical disks or tapes. Computing device 600 may also contain network connection(s) 616 that allow the device to communicate with other devices. Computing device 600 may also have input device(s) 614 such as a keyboard, mouse, touch screen, etc. Output device(s) 612, such as a display, speakers, printer, etc., may also be included. The additional devices may be connected to the bus to facilitate the communication of data among the components of the computing device. All these devices are well-known in the art and need not be discussed at length here.

[0122] The processing unit 606 may be configured to execute program code encoded in tangible, computer-readable media. Tangible, computer-readable media refers to any media that is capable of providing data that causes the computing device (i.e., a machine) to operate in a particular fashion. Various computer-readable media may be utilized to provide Attorney Docket No.11605-066WO1 FAU 202408 instructions to the processing unit for execution. Example of tangible, computer-readable media may include but is not limited to, volatile media, non-volatile media, removable media, and non-removable media implemented in any method or technology for the storage of information such as computer-readable instructions, data structures, program modules, or other data. System memory 604, removable storage 608, and non-removable storage 610 are all examples of tangible computer storage media. Examples of tangible, computer-readable recording media include but are not limited to, an integrated circuit (e.g., field-programmable gate array or application-specific IC), a hard disk, an optical disk, a magneto-optical disk, a floppy disk, a magnetic tape, a holographic storage medium, a solid-state device, RAM, ROM, electrically erasable program read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices.

[0123] In an example implementation, the processing unit 606 may execute program code stored in the system memory. For example, the bus may carry data to the system memory, from which the processing unit receives and executes instructions. The data received by the system memory 604 may optionally be stored on the removable storage 608 or the non-removable storage 610 before or after execution by the processing unit 606.

[0124] It should be understood that the various techniques described herein may be implemented in connection with hardware or software or, where appropriate, with a combination thereof. Thus, the methods and apparatuses of the presently disclosed subject matter, or certain embodiments or portions thereof, may take the form of program code (i.e., instructions) embodied in tangible media or removable storage media, such as floppy diskettes, CD-ROMs, hard drives, or any other machine-readable storage medium wherein, when the program code is loaded into and executed by a machine, such as a computing device, the machine becomes an apparatus for practicing the presently disclosed subject matter. In the case of program code execution on programmable computers, the computing device generally includes a processor, a storage medium readable by the processor (including volatile and non- volatile memory and / or storage elements), at least one input device, and at least one output device. One or more programs may implement or utilize the processes described in connection with the presently disclosed subject matter, for example, through the use of an application programming interface (API), reusable controls, or the like. Such programs may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, the program(s) can be implemented in Attorney Docket No.11605-066WO1 FAU 202408 assembly or machine language if desired. In any case, the language may be a compiled or interpreted language, and it may be combined with hardware implementations.

[0125] In one embodiment, disclosed herein is a non-transitory computer-readable storage medium comprising instructions that, when executed, cause at least one processor to perform the method of any preceding embodiments.

[0126] Although certain implementations may refer to utilizing aspects of the presently disclosed subject matter in the context of one or more stand-alone computer systems, the subject matter is not so limited but rather may be implemented in connection with any computing environment. For example, the components described herein can be hardware and / or software components in a single or distributed system, or in a virtual equivalent, such as, a cloud computing environment. Still further, aspects of the presently disclosed subject matter may be implemented in or across a plurality of processing chips or devices, and storage may similarly be affected across a plurality of devices. Such devices might include personal computers, network servers, and handheld devices, for example.

[0127] The following patents, applications, and publications, as listed below and throughout this document, describe various applications and systems that could be used in combination with the exemplary system and are hereby incorporated by reference in their entirety herein. References [1] U.S. Energy Information Administration - EIA – independent statistics and analysis. Use of electricity - U.S. Energy Information Administration (EIA). (2023, April). https: / / www.eia.gov / energyexplained / electricity / use-of-electricity.php [2] H. Ma, T. K. Saha, C. Ekanayake and D. Martin, “Smart Transformer for Smart Grid— Intelligent Framework and Techniques for Power Transformer Asset Management,” in IEEE Transactions on Smart Grid, vol.6, no.2, pp.1026-1034, March 2015, doi: 10.1109 / TSG.2014.2384501. [3] L. Ouyang, J. Wu, X. Jiang, D. Almeida and C. Wainwright, et al., “Training language models to follow instructions with human feedback,” in Proc. Advances in Neural Information Processing Systems, 2022, pp.27730-27744. [4] A. Kirillov, E. Mintun, N. Ravi, H. Mao and C. Rolland, et al., “Segment Anything,” [Online]. Available: https: / / arxiv.org / abs / 2304.02643. [5] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit and L. Jones, et al., “Attention is All You Need,” Advances in neural information processing systems, vol.30, 2017. Attorney Docket No.11605-066WO1 FAU 202408 [6] J. Devlin, M. Chang, K. Lee and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” [Online]. Available: https: / / arxiv.org / abs / 1810.04805. [7] Y. Sun, S. Wang, Y. Li, S. Feng and X. Chen, et al., “ERNIE: Enhanced Representation through Knowledge Integration,” [Online]. Available: https: / / arxiv.org / abs / 1904.09223. [8] H. Touvron, T. Lavril, G. Izacard, X. Martinet and M. Lachaux, et al., “Llama: Open and efficient foundation language models,” [Online]. Available: https: / / arxiv.org / abs / 2302.13971. [9] A. Radford, K. Narasimhan, T. Salimans and I. Sutskever, “Improving language understanding by generative pre-training,” [Online]. Available: https: / / openai.com / blog / chatgpt.

[0010] OpenAI, “GPT-4 Technical Report,” [Online]. Available: https: / / arxiv.org / abs / 2303.08774.

[0011] Li, Y., Wang, H., and Sun, M. (2023). ChatGPT-Like Large-Scale Foundation Models for Prognostics and Health Management: A Survey and Roadmaps. [Online]. Available: https: / / arxiv.org / abs / 2305.06472

[0012] H. Wang, Z. Liu, D. Peng and M.J. Zuo, “Interpretable convolutional neural network with multilayer wavelet for Noise-Robust Machinery fault diagnosis,” Mech. Syst. Signal Pr., vol.195, pp.110314, 2023.

[0013] Y. Wang and H. Wang, “Wavelet attention-powered neural network framework with hierarchical dynamic frequency learning for lithiumion battery state of health prediction,” Journal of Energy Storage, vol.61, pp.106697, 2023.

[0014] H. Liu, J. Zhou, Y. Zheng, W. Jiang and Y. Zhang, “Fault diagnosis of rolling bearings with recurrent neural network-based autoencoders,” ISA T., vol.77, pp.167-178, 2018.

[0015] Pallav K Bera, Can Isik, Vajendra Kumar, February 13, 2020, “Transients and Faults in Power Transformers and Phase Angle Regulators (DATASET)”, IEEE Dataport, doi: https: / / dx.doi.org / 10.21227 / 1d1wq940.

[0016] “A guide for using The wavelet transform in machine learning,” ML Fundamentals, https: / / ataspinar.com / 2018 / 12 / 21 / a-guide-for-usingthe-wavelet-transform-in-machine- learning / [Accessed Aug.29, 2023].

[0017] F. Wasilewski, “Discrete wavelet transform (DWT),” Discrete Wavelet Transform (DWT) - PyWavelets Documentation, https: / / pywavelets.readthedocs.io / en / latest / ref / dwt- discrete-wavelettransform. html [Accessed Aug.29, 2023]. Attorney Docket No.11605-066WO1 FAU 202408

[0018] J. B. Thomas and S. K.V., “Neural architecture search algorithm to optimize deep transformer model for fault detection in Electrical Power Distribution Systems,” Engineering Applications of Artificial Intelligence, vol.120, p.105890, 2023. doi:10.1016 / j.engappai.2023.105890

[0019] V. Nair and G. Hinton, “Rectified Linear Units Improve Restricted Boltzmann Machines.” Available: https: / / www.cs.toronto.edu / ˜hinton / absps / reluICML.pdf

[0020] D. Kingma and J. Lei Ba, “ADAM: A METHOD FOR STOCHASTIC OPTIMIZATION,” Jan.2017. Available: https: / / arxiv.org / pdf / 1412.6980.pdf

Claims

Attorney Docket No.11605-066WO1 FAU 202408 What is claimed is:

1. A smart grid data management system comprising: at least one processor; and a memory having instructions thereon, wherein the instructions when executed by the at least one processor, cause the at least one processor to: analyze, using a classification or large language model, energy system data to identify one or more faults; responsive to receiving a direct user input, determine, using a second model, one or more user input parameters; determine, based, at least in part, on one or more word-based tokens generated from the user input and the one or more user input parameters, at least one of a plurality of models; and generate, using the at least one of the plurality of models, one or more visualization outputs based, at least in part, on the one or more user input parameters and / or the one or more identified faults.

2. The system of claim 1, wherein the classification or large language model is a generative Transformer-based model.

3. The system of claim 1 or 2, wherein the second model comprises a Natural language Processor (NLP) model.

4. The system of any one of claims 1-3, wherein the classification model or large language model is configured to: receive input data comprising the direct user input and a signal file including power grid data; organize the input data into data matrices; classify / identify the one or more faults using a series of dense rectified linear unit (ReLU) layers and dropouts.

5. The system of claim 4, wherein the second model uses the data matrices containing the one or more classified faults as input data, and wherein the second model employs one or more databases or libraries to process the received data matrices.Attorney Docket No.11605-066WO1 FAU 202408 6. The system of any one of claims 1-5, wherein generating the one or more visualized outputs includes directing the energy system data to at least one of the plurality of models for analysis or processing.

7. The system of any one of claims 1-6, wherein the instructions further cause the at least one processor to: generate the plurality of models including the second model; and train the plurality of models using historical energy system data.

8. The system of claim 7, wherein the instructions further cause the at least one processor to: pre-process the historical energy system data prior to training the plurality of models.

9. The system of claim 7 or 8, wherein the historical energy system data is retrieved from one or more publicly available databases.

10. The system of any one of claims 1-9, wherein the instructions further cause the at least one processor to: output, via a display or graphical user interface, the one or more visualization outputs.

11. The system of claim 10, wherein the one or more visualization output comprises conversational text, graphs, and / or charts.

12. The system of any one of claims 1-11, wherein the second model comprises a chat- based model or model architecture.

13. The system of any one of claims 1-12, wherein the classification or large language model is seamlessly integrated into a backend Application Programming Interface (API).

14. The system of any one of claims 1-13, wherein the plurality of models comprise one or more of a deep neural network model and a large language model.

15. The system of any one of claims 1-14, wherein the energy system data comprises real- time data obtained from sensors operatively coupled with energy system components, and theAttorney Docket No.11605-066WO1 FAU 202408 instructions further cause the at least one processor to: continuously analyze streaming data from the sensors in order to promptly response to grid issues.

16. A method for prognostic power monitoring comprising: analyzing, using a classification model, energy system data to detect one or more fault classification outputs; responsive to receiving a direct user input, determining, using a second model, one or more user input parameters; determining based at least in part on one or more fault characteristics associated with the one or more fault classification outputs and the one or more user input parameters, at least one of a plurality of models; and generating, using at least one of a plurality of models, one or more predictive outputs based at least in part on the one or more fault characteristics and the one or more user input parameters.

17. The method of claim 16, further comprising: generating the plurality of models including the classification model and the second model; and training the plurality of models using historical energy system data.

18. The method of claim 16 or 17, wherein the one or more predictive outputs and / or the one or more user input parameters include at least one of a location, a fault type, an energy system / component, and / or power transformer(s).

19. The method of any one of claims 16-18, further comprising: outputting, via a display or graphical user interface, the one or more predictive outputs and / or one or more fault classification outputs / characteristics.

20. The method of any one of claims 16-19, further comprising: triggering or causing a corrective action with respect to at least one identified fault classification output (e.g., generate an alert, deploy personnel to a target location, and / or automatically activate or deactivate a system component).Attorney Docket No.11605-066WO1 FAU 202408 21. A non-transitory computer readable medium comprising a memory having instructions stored thereon to perform / implement any one of the system of claims 1-15 or the method of claims 16-20.

Citation Information

Patent Citations

  • Systems and methods for language model-based text insertion

    US11886826B1

  • Systems and methods for ai-assisted electrical power grid fault analysis

    US20230035930A1

Cited By

  • Power dispatching optimization method and system based on deep reinforcement learning

    CN121618635A