System and method for root cause analysis in building automation systems
The large language model-enhanced knowledge graph system addresses the challenges of misdiagnosis in building automation by automating data extraction and standardizing root cause analysis, improving accuracy and efficiency through dynamic adaptability and collaborative insights.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SIEMENS SCHWEIZ AG
- Filing Date
- 2025-10-08
- Publication Date
- 2026-05-15
AI Technical Summary
Conventional building automation systems face challenges in accurately diagnosing complex, multi-source faults due to reliance on static data-driven techniques, manual data interpretation leading to misdiagnosis, and lack of standardization, which hinders collaborative analysis and results in energy wastage and operational inefficiencies.
A large language model-enhanced knowledge graph system that automates the extraction and organization of unstructured maintenance data, applying proximity algorithms to generate weighted relationships for accurate root cause analysis, with mechanisms for schema validation and adaptability to dynamic data changes.
Enhances the precision and efficiency of fault diagnosis by standardizing root cause analysis, reducing misdiagnosis, and enabling seamless collaboration across professionals, while maintaining data accuracy and adaptability to evolving conditions.
Smart Images

Figure US2025050010_15052026_PF_FP_ABST
Abstract
Description
202420026System and Method for Root Cause Analysis in Building Automation SystemsTechnical Field
[0001] The present disclosure relates generally to building automation systems and fault diagnosis, and particularly to root cause analysis using large language models.Background
[0002] In the building automation domain, diagnosing building anomalies often requires facility managers to manually review large amounts of unstructured maintenance-related data, such as symptom reports and history data, before providing responsive actions. Beyond the time-consuming nature of this task, a more critical issue is the potential for misdiagnosis or incorrectly addressing the root cause of equipment failures. The reliance on facility managers to manually interpret unstructured natural language reports introduces significant opportunities for error, which can lead to incomplete or inconsistent diagnoses. Consequently, building anomalies may persist unaddressed, resulting in ongoing energy wastage and inconvenience for building occupants.
[0003] The open approach of conventional systems to fault diagnostics presents a significant challenge. This lack of standardization not only complicates the generation of consistent recommendations for fault remediation but also hinders the ability to leverage insights across professionals, thereby reducing the potential for collaborative analysis. Consequently, the diagnostic process becomes fragmented and challenging to optimize, leading to a higher risk of misdiagnoses and incorrect resolutions.
[0004] Conventional approaches in building automation fault detection often rely on autoencoder-based and inverse-model-based fault detection and diagnosis methods. These methods are typically restricted by limited data and are generally tailored to address isolated faults rather than interconnected ones. Recent advancements focus on real-time monitoring but struggle to identify complex, multi-source faults over time due to a reliance on static data-driven techniques. Furthermore, conventional systems predominantly address equipment-level issues and may overlook system- wide interactions and dependencies.
[0005] While knowledge graphs can help organize and connect data for efficient analysis, extracting entity types and relationships to build knowledge graphs from202420026 vast amounts of unstructured data presents significant technical challenges. Managing large graph databases can lead to long query times, increased service times, and reduced system responsiveness. Additionally, ensuring data consistency and accuracy when integrating information from diverse sources and formats remains problematic.Summary
[0006] In accordance with one embodiment of the disclosure, there is provided a unified and systematic approach to fault diagnostics approach for building management systems. Aspects of the present disclosure address and overcome at least some of the above-described limitations in building automation fault diagnosis by providing a large language model-enhanced knowledge graph system that can efficiently extract and organize unstructured maintenance data while enabling accurate root cause analysis through weighted relationship modeling.
[0007] A first aspect is a method for root cause analysis in a building automation system. Unstructured raw data related to building maintenance is received. A knowledge graph schema is generated from the unstructured raw data using a large language model agent based on a predetermined knowledge graph schema quality criteria. Concept pairs and semantic relationships are extract from the unstructured raw data using the large language model agent based on the knowledge graph schema and a predetermined knowledge graph quality criteria. A proximity algorithm is applied to add contextual proximity relationships between the concept pairs. The proximity algorithm is applied, at least in part, by grouping node pairs based on proximity and calculating weights for groups of the node pairs. A final knowledge graph is generated with weighted relationships for root cause analysis of a fault of the building automation system.
[0008] A second aspect is a system for root cause analysis in a building automation system. The system comprises one or more processors and memory stored instructions. When executed by the one or more processors, the memory stored instructions cause the system to perform a method for root cause analysis. Unstructured raw data related to building maintenance are received. A knowledge graph schema is generated from the unstructured raw data using a large language model agent based on a predetermined knowledge graph schema quality criteria. Concept pairs and semantic relationships are extracted from the unstructured raw data using the large language model agent based on the knowledge graph schema202420026 and a predetermined knowledge graph quality criteria. A proximity algorithm is applied to add contextual proximity relationships between the concept pairs by grouping node pairs based on proximity and calculating weights for the groups of the node pairs. A final knowledge graph is generated with weighted relationships for root cause analysis of a fault of the building automation system.
[0009] Embodiments described herein address various problems through various innovative technical features, including, for example and without limitation:
[0010] 1. Application of Knowledge Graphs and LLMs in the Root Cause Analyzer: Embodiments can leverage knowledge graphs to organize and store building maintenance documents, significantly improving data structure and accessibility. By utilizing large language models (LLMs), the system not only identifies potential root causes, symptoms, and actions from vast amounts of unstructured documents but also enables iterative interactions with users to refine and enhance the consistency and completeness of the knowledge graph. This iterative feedback loop improves the precision of the analysis, thereby enhancing both the accuracy and efficiency of root cause identification.
[0011] 2. Standardization of Fault Diagnosis via Knowledge Graph Schema: A predetermined knowledge graph schema standardizes the root causes diagnosis process, ensuring consistent representation of root causes, symptoms, and actions across various data inputs. This standardization facilitates interoperability between systems and professionals, enabling seamless sharing and application of insights. Moreover, the structured nature of the schema allows for more reliable querying, thus reducing variability in the diagnostic process. Iterative interactions with LLMs further contribute to the schema’s refinement, ensuring that the diagnosis becomes more accurate over time.
[0012] 3. Validation, Completeness, and Optimization of the Knowledge Graph Schema: To maintain the accuracy and relevance of the knowledge graph, embodiments can define mechanisms for validating the schema’s quality and completeness. Completeness is introduced as an additional evaluation metric, ensuring that all relevant data points are captured while minimizing redundancy. To avoid unnecessary complexity, the system applies optimization techniques such as graph simplification and deduplication, thereby improving the accuracy of query results and reducing end-to-end processing time. These optimizations enable202420026 quicker, more reliable access to maintenance information, significantly enhancing the overall user experience.
[0013] 4. Adaptability to Dynamic Data Changes: Embodiments an employ a flexible and dynamic knowledge graph structure that can seamlessly adapt to changes in data, including additions, deletions, and modifications. This adaptability ensures that the knowledge graph remains up-to-date, relevant, and accurate, reflecting the most current information available. In addition to adaptability, optimization approaches such as redundancy elimination and iterative validation further maintain the graph’s efficiency while minimizing unnecessary complexity. The combination of these strategies ensures that users have access to the most current and optimized data, improving decision-making and operational efficiency.
[0014] The above described features and advantages, as well as others, will become more readily apparent to those of ordinary skill in the art by reference to the following detailed description and accompanying drawings. While it would be desirable to provide one or more of these or other advantageous features, the teachings disclosed herein extend to those embodiments which fall within the scope of the appended claims, regardless of whether they accomplish one or more of the above- mentioned advantages.Brief Description of the Drawings
[0015] For a more complete understanding of the present disclosure, and the advantages thereof, reference is now made to the following descriptions taken in conjunction with the accompanying drawings, wherein like numbers designate like objects.
[0016] FIG. 1 is a block diagram illustrating a building management system including various subsystems and management devices according to some embodiments.
[0017] FIG. 2 is a block diagram illustrating the internal architecture of a building management device including processors, memory, and communication components according to some embodiments.
[0018] FIG. 3 is a flowchart diagram illustrating an example pipeline for generating a knowledge graph from unstructured raw data using large language models and clustering algorithms according to some embodiments.202420026
[0019] FIG. 4 is a flowchart diagram illustrating a process for adding unstructured raw data with new entities and relationships to an existing knowledge graph according to some embodiments.
[0020] FIG. 5 is a flowchart diagram illustrating a process for adding unstructured raw data without new entities and relationships to an existing knowledge graph according to some embodiments.Detailed Description
[0021] Various technologies that pertain to systems and methods that facilitate root cause analysis in building automation systems will now be described with reference to the drawings, where like reference numerals represent like elements throughout. The drawings discussed below, and the various embodiments used to describe the principles of the present disclosure in this patent document are by way of illustration only and should not be construed in any way to limit the scope of the disclosure. Those skilled in the art will understand that the principles of the present disclosure may be implemented in any suitably arranged apparatus. It is to be understood that functionality that is described as being carried out by certain system elements may be performed by multiple elements. Similarly, for instance, an element may be configured to perform functionality that is described as being carried out by multiple elements. The numerous innovative teachings of the present application will be described with reference to exemplary non-limiting embodiments.
[0022] A large language model (LLM)-enhanced approach is provided for root cause analysis in building automation systems. The capabilities of generative LLMs are leveraged to systematically organize and connect unstructured maintenance data through knowledge graph construction, enabling efficient analysis and accurate fault diagnosis. The system addresses the critical challenges of manual fault diagnosis by automating the extraction of entities and relationships from vast amounts of unstructured data while maintaining data consistency and accuracy.
[0023] The systems and methods overcome the limitations of conventional building automation fault detection systems that rely on static data-driven techniques and are restricted to isolated fault analysis. By integrating LLMs with knowledge graph technology, the systems and methods provide a unified and systematic approach to fault diagnostics that reduces the risk of misdiagnosis and enables collaborative analysis across professionals.202420026
[0024] Advantages of the systems and methods include minimizing manual labor through automated data processing, leveraging advanced LLM capabilities for accurate entity recognition and relationship extraction, handling linguistic and technical diversity in maintenance documentation, and providing efficient and scalable knowledge graph construction. The systems and methods also feature dynamic adaptability to accommodate new data while maintaining consistency and accuracy through continuous quality validation and optimization mechanisms.
[0025] Referring to FIG. 1, there is shown a building management system 100 of a facility 10 in an example implementation that is operable to employ techniques described herein. The building management system 100 provides the infrastructure for implementing the LLM-enhanced knowledge graph approach for root cause analysis in building automation systems. The system 100 includes one or more management devices 102, 106, 108 that facilitate data collection, processing, and analysis. Where multiple management devices 102, 106, 108 are utilized, the devices may operate cooperatively from remote locations. Each management device 102, 106, 108 may operate as a workstation and / or server that provides local data processing and storage for facility managers and technicians. For some embodiments, one or more management devices 108 may be utilized as a remote server for extended computational capabilities and cloud-based processing.
[0026] The management devices 102, 106, 108 may be interconnected through a management network 104, which enables communication and data exchange among the various devices. The system 100 further includes connectivity to a cloud network 110, allowing for scalable processing capabilities and access to external services and computational resources, such as the large language model (LLM).
[0027] The building management system 100 manages multiple subsystems 120, 122, 124 within the building infrastructure. A building environmental comfort subsystem 120 includes comfort devices 126 that monitor and control heating, ventilation, air conditioning (HVAC), lighting, and other comfort-related systems. A building security subsystem 122 includes security devices 128 that handle access control, surveillance, and security monitoring functions. A building fire safety subsystem 124 includes fire safety devices 130 that manage fire detection, suppression, and emergency response systems. Each of the devices 126, 128, and 130 generates unstructured raw data related to building maintenance, including fault reports, sensor readings, maintenance logs, and operational status information. This202420026 unstructured raw data serves as input to the LLM-enhanced knowledge graph system for root cause analysis and fault diagnosis.
[0028] Referring to FIG. 2, there is shown an internal architecture of a building management device 102, 106, 108 according to some embodiments of the present disclosure. The device components 200 shown in FIG. 2 represent internal components of any of the management devices 102, 106, or 108 shown in FIG. 1 and provide the computational infrastructure for implementing the LLM-enhanced knowledge graph system.
[0029] The device components 200 are interconnected through a communication bus 202 that enables efficient data transfer and coordination between the various system elements. The device components 200 include a communication component 204 that handles data exchange with other devices and systems through the management network 104 and cloud network 110.
[0030] The device components 200 include a processor 206 that serves as the central processing unit for executing the various algorithms and processes described herein. The processor 206 includes a first module 210 and a second module 212 that can be configured to handle different aspects of the knowledge graph generation and analysis processes. For example, the first module 210 may be dedicated to LLM operations and natural language processing tasks, while the second module 212 may handle graph computational tasks, such as GON processing and clustering algorithms.
[0031] A memory component 208 provides data storage capabilities for the system 100. The memory component 208 includes first data 214 and second data 216 that can store different types of information required for the knowledge graph operations. The first data 214 may include unstructured raw data, knowledge graph schemas, and quality criteria, while the second data 216 may store processed data such as entity embeddings, knowledge data frames, and final knowledge graphs.
[0032] Input / output (I / O) interfaces 222 provide connectivity for external devices and data sources. A user interface 224 within the I / O interfaces 222 enables facility managers and technicians to interact with the system, provide feedback for iterative refinement of the knowledge graph, and access root cause analysis results.202420026
[0033] It is to be understood that FIG. 2 is provided for illustrative purposes only to represent examples of the device components 200 of one or more management devices 102, 106, 108 and is not intended to be a complete diagram of the various components that may be utilized by the system 100. Therefore, the device or devices 102, 106, 108 may include various other components not shown in FIG. 2, may include a combination of two or more components, or a division of a particular component into two or more separate components, and still be within the scope of the present invention.
[0034] Referring to FIG. 3, there is shown an example first pipeline 300 for generating a knowledge graph from unstructured raw data according to some embodiments of the present disclosure. The first pipeline 300 represents a first scenario where the complete knowledge graph generation process is performed from initial unstructured data inputs. The systems and methods employ a multi-phase pipeline that transforms unstructured raw data into a comprehensive knowledge graph with weighted relationships. In the first phase, a large language model (LLM) agent generates a knowledge graph schema based on predetermined quality criteria, followed by entity node embedding computation using graph convolutional networks and recursive clustering to optimize the schema structure. In the second phase, the system extracts concept pairs and semantic relationships, applies proximity algorithms to establish contextual connections, and generates a final knowledge graph optimized for root cause analysis.
[0035] The first pipeline 300 is organized into multiple phases that systematically transform unstructured raw data into a comprehensive knowledge graph. The process begins with unstructured raw data 310 and knowledge graph schema (KGS) quality criteria 312 as inputs to the first phase (320, 324, 328) of the pipeline 300. The unstructured data 310 is raw, unstructured data such as text documents, articles, or any data source that hasn’t been structured.
[0036] In Phase 1.1 , a first LLM agent receives (320) the unstructured raw data 310 and applies the KGS quality criteria 312 to generate an initial knowledge graph schema 322. The first LLM agent 320 leverages natural language processing capabilities to identify potential entities, relationships, and structural patterns within the unstructured data, creating a preliminary schema, namely the initial knowledge graph schema 322, that serves as the foundation for subsequent processing. The initial knowledge graph schema 322 represents the structured entities and202420026 relationships identified in the raw data. The KGS quality criteria 312 is a predetermined criteria to ensure the quality and accuracy of the knowledge graph schema.
[0037] An example template for KGS criteria, designed to ensure the quality and accuracy of the KGS may be represented by the table below.The above template covers key metrics, including consistency, completeness, redundancy, and scalability, to support the effectiveness and sustainability of the schema.
[0038] Phase 1.2 involves computing entity node embeddings 326 using a graph convolutional network (GON). The GON processes (324) the initial knowledge graph schema 322 to generate entity node embeddings 326, which provide numerical representations of the entities that capture their semantic relationships and contextual information within the graph structure. The entity node embeddings represent entities in a high-dimensional space.
[0039] In Phase 1.3, a KMeans clustering algorithm is applied to the entity node embeddings 326. The clustering algorithm recursively clusters (328) the entities and relationships until the predetermined KGS quality criteria is met, thereby producing a final knowledge graph schema 330. This iterative clustering process optimizes the graph structure by grouping related entities and refining the overall schema organization. The final knowledge graph schema 330 is a refined and structure knowledge graph schema that accurately represents the entities and their relationships.
[0040] The second phase (350, 356) of the pipeline utilizes unstructured raw data 340 along with a final knowledge graph schema 342 and a knowledge graph (KG) quality criteria 344. For some embodiments, the unstructured raw data 340 is the same as the unstructured raw data 310, and the knowledge graph schema 342 is the202420026 same as the final knowledge graph schema 330. In Phase 2.1, an LLM agent is employed to dive deeper into the raw data and extract concept pairs and semantic relationships. This phase emphasizes identifying key concepts that are contextually linked in the data. A second LLM agent extracts (350) concept pairs and semantic relationships from the unstructured raw data 340 based on the optimized schema 342 and quality criteria 344, generating a knowledge data frame 352. For some embodiments, the first and second LLM agents are the same LLM agent. The unstructured raw data 340 are the unstructured data that continues to be analyzed for more detailed insights into concept pairs and relationships. The knowledge graph schema 342 is the refined knowledge graph schema from Phase 1.3 used as the foundation. The knowledge graph quality criteria are predetermined criteria that ensure the quality and accuracy of the extracted relationships. The knowledge data frame 352 is a data frame including the extracted concept pairs and their semantic relationships, providing a structured dataset for further refinement.
[0041] An example template for KG quality criteria, aimed at ensuring data quality and structure, may be represented by the table below.The above template covers key metrics including semantic accuracy, where relationships reflect real scenarios; information density, providing sufficient operational detail; completeness, capturing all essential relationships; and consistency, ensuring unique IDs and consistent relationships. These criteria help evaluate the knowledge graph's effectiveness in supporting operational needs. Unlike knowledge graph schema quality criteria, which focuses on the design and structure of the schema itself, the knowledge graph quality criteria can evaluate the actual data and relationships within the constructed knowledge graph, ensuring that it accurately represents and supports real-world applications.202420026
[0042] Phase 2.2 applies a proximity algorithm to the knowledge data frame 352. The proximity algorithm adds (356) contextual proximity relationships between the extracted concept pairs, groups node pairs based on proximity, and calculates weights for the groups, ultimately producing the final knowledge graph 360 with weighted relationships optimized for root cause analysis. For example, a Cosine Similarity based algorithm may be applied to the knowledge data frame to add contextual proximity relationships between the extracted concepts, group node pairs based on their proximity and contextual relevance, and calculate the sum weights for these groups, capturing the strength of the relationships between the concepts. The final knowledge graph may include enriched relationships that reflect the contextual and weighted connections between entities. This final knowledge graph is ready for use in various applications, such as root cause analysis or decision support systems.
[0043] Referring to FIG. 4, there is shown a second pipeline 400 for adding unstructured raw data with new entities and relationships to an existing knowledge graph according to some embodiments of the present disclosure. The second pipeline 400 represents a second scenario where the system 100 dynamically updates the knowledge graph schema to incorporate new entities and relationships from additional data sources.
[0044] The second pipeline 400 demonstrates the capability of the system 100 to handle evolving data requirements while maintaining consistency and accuracy. The process begins with multiple data inputs, specifically existing unstructured raw data 410, new unstructured raw data with new entity and relationship 414, and existing unstructured raw data 440 that serves as the baseline for the updated pipeline.
[0045] The first phase (420, 424, 428) of the second pipeline 400 follows a similar structure to the first pipeline 300, but with enhanced input handling. KGS quality criteria 412 guide the processing of both existing and new data streams. In Phase 1.1 , the LLM agent analyzes the new and existing unstructured raw data to update the knowledge graph schema, ensuring alignment with KGS quality criteria. A first LLM agent processes (420) the combined data inputs 410, 414 to generate an initial knowledge graph schema 422 based on the knowledge graph schema quality criteria 412 that accommodates both the existing entities and the newly identified entities and relationships. The new unstructured raw data 414 includes new entities and relationships, providing previously unseen information to be integrated. The existing202420026 unstructured raw data 410 are provided by general raw data sources, such as text documents or articles. The KGS quality criteria 412 are predetermined standards for evaluating the quality of the knowledge graph schema. The initial knowledge graph schema 422 is a preliminary schema updated with new entities and relationships.
[0046] Phase 1.2 employs a graph convolutional network (GCN) to compute (424) entity node embeddings 426 for all entities in the updated schema, including both existing and new entities, based on the initial knowledge graph schema 422. This ensures that the numerical representations capture the expanded semantic relationships within the enhanced graph structure. The GCN transforms embeddings into a high-dimensional space for easier clustering and analysis. The entity node embeddings 426 are high-dimensional representations of entities, prepared for the next clustering phase.
[0047] In Phase 1.3, a KMeans clustering algorithm recursively clusters (428) the entities and relationships based on the entity node embeddings 426, taking into account the new additions while maintaining the integrity of existing relationships. The clustering process continues iteratively until the predetermined KGS quality criteria are satisfied, producing a final knowledge graph schema 430 that incorporates the new entities and relationships. The final knowledge graph schema 430 is a refined knowledge graph schema that accurately incorporates and clusters the new entities and relationships, ready for integration into the knowledge graph.
[0048] The second phase (450, 456) utilizes knowledge graph (KG) quality criteria 444 to guide the extraction and integration process. In Phase 2.1 , a second LLM agent extracts (450) concept pairs and semantic relationships from the combined data sources (440, 442, 444), generating a knowledge data frame 452 that includes both existing and new relationship information. For some embodiments, the first and second LLM agents are the same LLM agent.
[0049] Phase 2.2 applies (456) a proximity algorithm to the enhanced knowledge data frame 452. The proximity algorithm 456 adds (456) contextual proximity relationships between all concept pairs, including new relationships involving the newly added entities, groups node pairs based on proximity calculations, and calculates weights for all groups. This process produces a final knowledge graph 460 that seamlessly integrates the new entities and relationships while preserving the202420026 existing graph structure and optimizing the overall system for comprehensive root cause analysis.
[0050] Referring to FIG. 5, there is shown a third pipeline 500 for adding unstructured raw data without new entities and relationships to an existing knowledge graph according to some embodiments of the present disclosure. The third pipeline 500 represents a third scenario where the system 100 integrates additional data into the final knowledge graph without modifying the existing knowledge graph schema. New data points are integrated into the existing knowledge graph without introducing new entities or relationships.
[0051] The pipeline 500 demonstrates the efficiency of the system 100 in handling incremental data updates that do not require structural changes to the knowledge graph. The process begins with multiple inputs. A final knowledge graph schema 542 from a previous processing cycle, unstructured raw data 540, new unstructured raw data without new entities and relationships 546, and knowledge graph (KG) quality criteria 544. Since this scenario does not involve new entities or relationships, the third pipeline 500 bypasses the schema generation and optimization phases, similar to the first phases of the first and second pipelines 300, 400, and proceeds directly to the data integration phases. The existing final knowledge graph schema 542 serves as the structural foundation for processing the additional data.
[0052] In Phase 2.1 , an LLM agent 550 extracts concept pairs and semantic relationships from both the unstructured raw data 540 and the new unstructured raw data 546. The LLM agent generates (550) an initial knowledge data frame 552 and a new knowledge data frame 554, both based on the established knowledge graph schema 542 and knowledge graph (KG) criteria 544. The unstructured raw data 540 are fresh, unstructured data inputs. The new unstructured raw data 546 are data points the enhance or add context to existing entities and relationships. The knowledge graph criteria 544 are guidelines to maintain the integrity and quality of the knowledge graph during integration. The new knowledge data frame 554 is a structured representation of the new data, derived from the initial knowledge data frame 552. This new knowledge data frame 554 becomes the input for the next phase.
[0053] In Phase 2.2, an algorithm incorporates contextual relationships by grouping node pairs and summing their weights, enabling the integration of new data into the202420026 existing knowledge graph. This phase applies (556) a proximity algorithm to both knowledge data frames 552 and 554. The proximity algorithm adds (556) contextual proximity relationships, groups node pairs based on proximity calculations, and calculates weights for the groups, producing a final knowledge graph 560 and new knowledge graph entity and relationship information 562. The new knowledge data frame 554 is a structured dataset of extracted concept pairs and relationships from raw data, ready for integration into the knowledge graph without adding new entities. The final knowledge graph 560 is an updated graph that reflects the added context from the new data points without introducing any new entities or relationships.
[0054] Phase 2.3 introduces an additional processing step specific to the third scenario of the third pipeline 500. An updating process integrates (564) the new knowledge graph entity and relationship information 562 with the final knowledge graph 560, ensuring seamless integration without schema modifications. This process produces a new final knowledge graph 570 that incorporates all the additional data while maintaining the integrity and structure of the original knowledge graph.
[0055] The third pipeline 500 provides significant computational efficiency advantages by avoiding the resource-intensive schema generation and clustering processes when the additional data does not introduce new entity types or relationship categories. This approach of the third pipeline 500 enables rapid integration of incremental maintenance data while preserving the optimized structure of the existing knowledge graph.
[0056] The disclosed system provides significant advantages over conventional building automation fault detection approaches. By leveraging the capabilities of large language models for natural language processing and knowledge graph construction, the system enables automated extraction and organization of complex maintenance data while maintaining high accuracy and consistency. The multi-phase pipeline approach ensures optimal schema generation and relationship modeling, while the dynamic adaptation capabilities allow the system to evolve with changing data requirements.
[0057] The proximity algorithms employed in the system enable sophisticated contextual relationship modeling that captures subtle dependencies between building components and fault conditions. This enhanced relationship modeling facilitates202420026 more accurate root cause identification and enables facility managers to address underlying issues rather than merely treating symptoms.
[0058] The system's ability to handle different data integration scenarios provides operational flexibility while maintaining computational efficiency. The complete pipeline (the first scenario) enables comprehensive knowledge graph generation from scratch, while the adaptive scenarios (the second and third scenarios) allow for efficient incremental updates that preserve existing knowledge while incorporating new information.
[0059] Alternative embodiments may employ different clustering algorithms such as hierarchical clustering or DBSCAN instead of KMeans clustering, depending on the specific characteristics of the building data and performance requirements. Similarly, various proximity algorithms including Euclidean distance, Manhattan distance, or custom similarity measures may be utilized based on the semantic characteristics of the extracted entities and relationships.
[0060] The system may also incorporate feedback mechanisms that enable continuous learning and improvement of the knowledge graph quality. User feedback on root cause analysis results can be integrated back into the system to refine the quality criteria and improve future knowledge graph generation and analysis processes.
[0061] Embodiments distinguish from known solutions, for example and without limitation, by integrating advanced Al capabilities, specifically large language models (LLMs), with knowledge graph technology for superior data storage and analysis. Unlike traditional databases, embodiments can leverage LLMs for efficient and accurate identification of root causes, symptoms, and actions from unstructured documents, enhancing data organization and accessibility. Embodiments also feature quality validation and optimization of the knowledge graph schema, reducing complexity and redundancy, leading to faster data retrieval and processing times. Furthermore, it incorporates a flexible structure that automatically adapts to dynamic data changes, ensuring the graph remains current and accurate. This adaptability, combined with advanced Al-driven data processing, significantly improves accuracy and efficiency while minimizing manual intervention. These innovations, among others described herein, collectively offer a more user-friendly, consistent and responsive system, resulting in faster query times, enhanced data retrieval accuracy,202420026 less manual effort and an overall improved user experience, ultimately leading to more effective and timely maintenance management.
[0062] The embodiments of the present disclosure may be implemented with any combination of hardware and software. In addition, the embodiments of the present disclosure may be included in an article of manufacture (e.g., one or more computer program products) having, for example, a non-transitory computer-readable storage medium. The computer readable storage medium has embodied therein, for instance, computer readable program instructions for providing and facilitating the mechanisms of the embodiments of the present disclosure. The article of manufacture can be included as part of a computer system or sold separately.
[0063] The computer readable storage medium can include a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network.
[0064] The system and processes of the figures are not exclusive. Other systems, processes and menus may be derived in accordance with the principles of the disclosure to accomplish the same objectives. Although this disclosure has been described with reference to particular embodiments, it is to be understood that the embodiments and variations shown and described herein are for illustration purposes only. Modifications to the current design may be implemented by those skilled in the art, without departing from the scope of the appended claims.
[0065] Those skilled in the art will recognize that, for simplicity and clarity, the full structure and operation of all data processing systems suitable for use with the present disclosure are not being depicted or described herein. Also, none of the various features or processes described herein should be considered essential to any or all embodiments, except as described herein. Various features may be omitted or duplicated in various embodiments. Various processes described may be202420026 omitted, repeated, performed sequentially, concurrently, or in a different order. Various features and processes described herein can be combined in still other embodiments as may be described in the claims.
[0066] It is important to note that while the disclosure includes a description in the context of a fully functional system, those skilled in the art will appreciate that at least portions of the mechanism of the present disclosure are capable of being distributed in the form of instructions contained within a machine-usable, computer-usable, or computer-readable medium in any of a variety of forms, and that the present disclosure applies equally regardless of the particular type of instruction or signal bearing medium or storage medium utilized to actually carry out the distribution. Examples of machine usable / readable or computer usable / readable mediums include: nonvolatile, hard-coded type mediums such as read only memories (ROMs) or erasable, electrically programmable read only memories (EEPROMs), and user- recordable type mediums such as floppy disks, hard disk drives and compact disk read only memories (CD-ROMs) or digital versatile disks (DVDs).
[0067] Although an example embodiment of the present disclosure has been described in detail, those skilled in the art will understand that various changes, substitutions, variations, and improvements disclosed herein may be made without departing from the spirit and scope of the disclosure in its broadest form.
Claims
202420026Claims1. A method for root cause analysis in a building automation system, the method comprising: receiving unstructured raw data related to building maintenance; generating, using an LLM agent, a knowledge graph schema from the unstructured raw data based on a predetermined knowledge graph schema quality criteria; extracting, using the LLM agent, concept pairs and semantic relationships from the unstructured raw data based on the knowledge graph schema and a predetermined knowledge graph quality criteria; applying a proximity algorithm to add contextual proximity relationships between the concept pairs, wherein applying the proximity algorithm including grouping node pairs based on proximity and calculating weights for groups of the node pairs; and generating a final knowledge graph with weighted relationships for root cause analysis of a fault of the building automation system.
2. The method as described in claim 1, further comprising: determining entity nodes and their relationships based on an initial knowledge graph schema generated by the LLM agent using a graph convolutional network in response to generating the knowledge graph schema.
3. The method as described in claim 2, further comprising: producing the knowledge graph schema by recursively clustering the entity nodes and their relationships using a clustering algorithm until the predetermined knowledge graph schema quality criteria is met.
4. The method as described in claim 1, wherein the predetermined knowledge graph schema quality criteria comprise at least one of consistency, completeness, redundancy, scalability, semantic accuracy, or information density metrics.
5. The method as described in claim 1, further comprising: receiving user feedback on the concept pairs and the semantic relationships; and iteratively refining the knowledge graph schema based on the user feedback.2024200266. The method as described in claim 1, further comprising: receiving additional unstructured raw data including new entities and relationships; and dynamically updating the knowledge graph schema to incorporate the new entities and relationships.
7. The method as described in claim 1, further comprising: receiving additional unstructured raw data without new entities and relationships; and integrating the additional unstructured raw data into the final knowledge graph without modifying the knowledge graph schema.
8. The method as described in claim 1, wherein: the proximity algorithm comprises a cosine similarity algorithm; the clustering algorithm comprises a k-means clustering algorithm; and the unstructured raw data comprises at least one of maintenance reports, symptom descriptions, historical fault data, or technical documentation.
9. The method as described in claim 1, further comprising: identifying a root cause of the fault of the building automation system by querying the final knowledge graph with fault symptoms.
10. The method as described in claim 9, further comprising: generating a recommended action for addressing the root cause based on the weighted relationships in the final knowledge graph.20242002611. A system for root cause analysis in a building automation system, the system comprising: one or more processors; and memory stored instructions that, when executed by the one or more processors, cause the system to: receive unstructured raw data related to building maintenance; generate, using an LLM agent, a knowledge graph schema from the unstructured raw data based on a predetermined knowledge graph schema quality criteria; extract, using the LLM agent, concept pairs and semantic relationships from the unstructured raw data based on the knowledge graph schema and a predetermined knowledge graph quality criteria; apply a proximity algorithm to add contextual proximity relationships between the concept pairs by grouping node pairs based on proximity and calculating weights for the groups of the node pairs; and generate a final knowledge graph with weighted relationships for root cause analysis of a fault of the building automation system.
12. The system as described in claim 11 , wherein the instructions further cause the system to: determine entity node and their relationships based on an initial knowledge graph schema generated by the LLM agent using a graph convolutional network in response to generating the knowledge graph schema.
13. The system as described in claim 12, wherein the instructions further cause the system to: produce the knowledge graph schema by recursively cluster the entity nodes and their relationships using a clustering algorithm until the predetermined knowledge graph schema quality criteria is met.
14. The system as described in claim 11 , wherein the predetermined knowledge graph schema quality criteria comprise at least one of consistency, completeness, redundancy, and scalability, semantic accuracy, or information density metrics.
15. The system as described in claim 11 , wherein the instructions further cause the system to: receive user feedback on the concept pairs and the semantic relationships;202420026 and iteratively refine the knowledge graph schema based on the user feedback.
16. The system as described in claim 11 , wherein the instructions further cause the system to: receive additional unstructured raw data including new entities and relationships; and dynamically update the knowledge graph schema to incorporate the new entities and relationships.
17. The system as described in claim 11 , wherein the instructions further cause the system to: receive additional unstructured raw data without new entities and relationships; and integrate the additional unstructured raw data into the final knowledge graph without modifying the knowledge graph schema.
18. The system as described in claim 11 , wherein: the proximity algorithm comprises a cosine similarity algorithm; the clustering algorithm comprises a k-means clustering algorithm; and the unstructured raw data comprises at least one of maintenance reports, symptom descriptions, historical fault data, or technical documentation.
19. The system as described in claim 11 , wherein the instructions further cause the system to: identify a root cause of the fault of the building automation system by querying the final knowledge graph with fault symptoms.
20. The system as described in claim 19, wherein the instructions further cause the system to: generate a recommended action for addressing the root cause based on the weighted relationships in the final knowledge graph.