Systems and Methods for Predictive Maintenance of HVAC Equipment Using Machine Learning

By applying machine learning and survival analysis in building automation systems, the problem of inaccurate maintenance methods is solved, more accurate equipment failure prediction and maintenance cost estimation are achieved, and maintenance efficiency and equipment reliability are improved.

CN115943352BActive Publication Date: 2025-07-29SIEMENS INDUSTRY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180051333.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-21
Filing Date
2021-07-21
Publication Date
2025-07-29
Estimated Expiration
2041-07-21

AI Technical Summary

Technical Problem

Maintenance of building automation systems is both expensive and time-consuming, and equipment failures can affect production and comfort. The existing maintenance methods are heuristic and inaccurate, and cannot accurately predict equipment failures and maintenance requirements.

Method used

Using machine learning technology, it uses inference engines and predictive maintenance engines to perform survival analysis, generate updated failure data and conduct root cause analysis, combining Bayesian networks and survival curves to predict equipment failure and maintenance costs.

Benefits of technology

It realizes more accurate equipment failure prediction and maintenance cost estimation, improves maintenance efficiency, and reduces unnecessary maintenance overhead and equipment failure impact.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115943352B_ABST
    Figure CN115943352B_ABST
Patent Text Reader

Abstract

A method of predictive maintenance using machine learning in a building automation system (100) and corresponding systems and computer-readable media. A method includes receiving (1902) device event data (522) corresponding to a device (112) and executing an inference engine (522) to determine root cause failure data (524) corresponding to the device event data (522). The method includes executing (1906) a predictive maintenance engine (508) to generate a survival analysis (402, 406, 410, 1604) of a physical device (112) based on the root cause failure data (524). The method includes generating (1910) updated failure data (526) by the predictive maintenance engine (508) based on the survival analysis (402, 406, 410, 1604) and providing the updated failure data (526) to the inference engine (522). Thereafter, the inference engine (522) uses the updated failure data (526) in subsequent root cause analysis. The method includes outputting (1912) the survival analysis (402, 406, 410, 1604).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to systems and methods for predictive maintenance in building control systems and other systems. Background Art

[0002] Building automation systems include a variety of systems that assist in monitoring and controlling various aspects of building operation. Building automation systems include security systems, fire safety systems, lighting systems, and heating, ventilation, and air conditioning (HVAC) systems. The components of building automation systems are widely distributed throughout the facility. For example, an HVAC system can include temperature sensors and damper controls, as well as other components located in nearly every area of the facility. These building automation systems typically have one or more central control stations from which system data can be monitored and various aspects of system operation can be controlled and / or monitored.

[0003] To allow monitoring and control of the decentralized control system components, building automation systems typically employ a multi-level communication network to transfer operation and / or alarm information between the operating components (such as sensors and actuators) and the central control station. An example of a building automation system is the DXR controller, available from the Building Technologies Division of Siemens Industry, Inc. (“Siemens”) in Buffalo Grove, Illinois. In this system, several control stations connected by Ethernet or another type of network can be distributed at one or more building locations, each control station having the ability to monitor and control system operation.

[0004] Maintenance of building automation systems is both expensive and time-consuming. Equipment failures can affect production, comfort levels, and facility operation, and can occur without warning. Improved systems are desired. Summary of the Invention

[0005] The present disclosure describes systems and methods for using machine learning for predictive maintenance in building automation systems and corresponding systems and computer-readable media. A method performed by a data processing system of a building automation system includes receiving device event data corresponding to a physical device of the building automation system. The method includes executing an inference engine to determine root cause failure data corresponding to the device event data. The method includes executing a predictive maintenance engine to generate a survival analysis of the device based on the root cause failure data. The method includes generating updated failure data by the predictive maintenance engine based on the survival analysis and providing the updated failure data to the inference engine, wherein the inference engine thereafter uses the updated failure data in subsequent root cause analysis. The method includes outputting the survival analysis, such as displaying or sending the survival analysis.

[0006] In some embodiments, device event data includes sensor data of a device for a particular time instance and the failure rate of the device at the same time instance. Some embodiments further include generating an aggregated survival curve for multiple devices using AND SA or OR SA operators that specify survival relationships between the multiple devices. In some embodiments, device event data is received directly or indirectly from one or more event detection applications that identify device or system events based on sensor data. In some embodiments, root cause failure data includes the device failure probability for a particular time instance. In some embodiments, performing a predictive maintenance engine to generate a survival analysis includes receiving root cause failure data, generating an augmented time-event table based on the root cause failure data, and using the augmented time-event table and the similarity between the device event data corresponding to a device and the device event data of other devices to generate a similarity-based survival curve. In some embodiments, the inference engine includes a Bayesian network and makes decisions based on the Bayesian network, which associates device events with device failures. In some embodiments, the inference engine combines the device event data with the output of the Bayesian network to generate root cause failure data. In some embodiments, the survival analysis includes one or more survival curves generated by performing a parametric survival analysis process, a non-parametric survival analysis process, or a probability-similarity-based survival analysis process. Some embodiments further include performing a predictive maintenance engine to generate a cost analysis corresponding to a device based on the survival analysis. Some embodiments further include performing a predictive maintenance engine to generate a cost analysis corresponding to a device based on the survival analysis, and generating an aggregated cost analysis for multiple devices using AND CA or OR CA operators based on the survival relationships between the multiple devices. In some embodiments, the survival analysis includes performing a probability-similarity-based survival analysis (SSA) process and includes performing principal component analysis and regression analysis on selected device event data to establish a health index representing the survival probability of the device.

[0007] The foregoing has outlined rather broadly some features and technical advantages of the present disclosure so that the detailed description that follows may be better understood by those skilled in the art. Additional features and advantages of the present disclosure will be described hereinafter, which form the subject matter of the claims. Those skilled in the art will understand that they can readily use the disclosed concepts and specific embodiments as a basis for modifying or designing other structures for achieving the same purposes of the present disclosure. Those skilled in the art will also recognize that such equivalent structures do not depart from the spirit and scope of the broadest form of the present disclosure.

[0008] Before proceeding with the following detailed description, it may be advantageous to set forth certain definitions of words and phrases used throughout this patent document: The terms "comprising" and "including" and their derivatives mean including without limitation; the term "or" is inclusive and means and / or; the phrases "associated with" and "associated therewith" and their derivatives may mean including, being included therein, interconnected with, containing, being contained therein, connected to or coupled with, communicable with, cooperating with, interlaced, juxtaposed, adjacent to, bound to or bound with, having, having the attribute of, etc.; the term "controller" refers to any device, system or portion thereof that controls at least one operation, whether such device is implemented in hardware, firmware, software or some combination of at least two of hardware, firmware, software. It should be noted that the functions associated with any particular controller may be centralized or distributed, whether local or remote. Certain definitions of words and phrases are provided in this patent document, and those of ordinary skill in the art will understand that these definitions apply in many, if not most, cases to the prior and future use of such defined words and phrases. Although some terms may encompass a variety of embodiments, the appended claims may specifically limit these terms to particular embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more fully understand the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which like numerals represent like objects and in which:

[0010] Figure 1 A block diagram of a building automation system according to the present disclosure is shown, in which the data quality of a heating, ventilation, and air conditioning (HVAC) system or other systems can be improved;

[0011] Figure 2 Shows a detail of one of the field panels according to the present disclosure Figure 1 of;

[0012] Figure 3 Shows a detail of one of the field controllers according to the present disclosure Figure 1 of;

[0013] Figure 4A , Figure 4B and Figure 4C Show exemplary survival curves according to embodiments of the present disclosure;

[0014] Figure 5 Shows an example of elements of a software architecture capable of performing the processes disclosed herein;

[0015] Figure 6 Shows a non-limiting example of historical data according to the disclosed embodiments;

[0016] Figure 7 shows a non - limiting example of a maintenance cost curve according to the disclosed embodiments;

[0017] Figure 8 shows an example of a time - event table that can be used according to the disclosed embodiments;

[0018] Figure 9 shows an example output of an inference engine according to the disclosed embodiments;

[0019] Figure 10 shows an example of a standardized time - event table for a single device;

[0020] Figure 11 shows an example of a fitted sample according to the disclosed embodiments;

[0021] Figure 12 shows an example of sample generation according to the disclosed embodiments;

[0022] Figure 13 shows an example of an extended time - event table according to the disclosed embodiments;

[0023] Figure 14A and Figure 14B shows a similarity - based survival analysis process according to the disclosed embodiments;

[0024] Figure 15A and Figure 15B shows the use of PCA on sensor values according to the disclosed embodiments;

[0025] Figure 16 shows an example of a cost analysis process according to the disclosed embodiments;

[0026] Figure 17A and Figure 17B shows an example of calculating cost analysis from survival analysis according to the disclosed embodiments;

[0027] Figure 18 shows the AND CA aggregation of a maintenance cost curve according to the disclosed embodiments;

[0028] Figure 19 and Figure 20 shows an example of a process according to the disclosed embodiments;

[0029] Figure 21 shows an example of using the logical operators disclosed herein; and

[0030] Figure 22 shows a block diagram of a data processing system in which various embodiments can be implemented. Detailed implementation manners

[0031] The following discussion Figures 1 to 22 and the various embodiments used to describe the principles of the present disclosure in this patent document are for illustration only and should not be construed in any way as limiting the scope of the present disclosure. Those skilled in the art will understand that the principles of the present disclosure can be implemented in any appropriately arranged device. Many innovative teachings of this application will be described with reference to exemplary non-limiting embodiments.

[0032] A building automation system (BAS) as disclosed herein can operate in an automatic operation mode, which helps to effectively operate the systems in a space to save energy. The BAS continuously evaluates the environmental conditions and energy usage in the space, and can determine and indicate to the user when the space is operating most effectively. Similarly, the BAS can determine and indicate when the system is operating inefficiently, such as when the occupants are not in control of the room due to personal preferences or due to a sharp change in weather conditions. The BAS can automatically, or upon user input, adjust the control settings to make the system operate effectively again.

[0033] For the normal operation of the BAS, the physical devices and other devices in the BAS must be maintained and occasionally replaced. Delaying maintenance or replacement until a failure actually occurs will result in unnecessary costs and inconvenience and discomfort to the occupants.

[0034] The disclosed embodiments include systems and methods for using machine learning for predictive maintenance of HVAC devices and other physical devices to ensure the correct operation of the BAS or other systems.

[0035] Figure 1 A block diagram of a building automation system 100 in which the disclosed embodiments can be implemented is shown. The building automation system 100 is an environmental control system configured to control at least one parameter among a plurality of environmental parameters in a building, such as temperature, humidity, lighting, and / or the like. For example, for a specific embodiment, the building automation system 100 can include a DXR controller building automation system that allows setting and / or changing various controls of the system. Although a brief description of the building automation system 100 is provided below, it should be understood that the building automation system 100 described herein is only an example of a specific form or configuration of a building automation system, and the system 100 can be implemented in any other suitable manner without departing from the scope of the present disclosure.

[0036] For the illustrated embodiment, building automation system 100 includes a site controller 102, a reporting server 104, a plurality of client stations 106a-c, a plurality of field panels 108a-b, a plurality of field controllers 110a-e, and a plurality of field devices 112a-d. Although three client stations 106, two field panels 108, five field controllers 110, and four field devices 112 are shown, it should be understood that system 100 may include any suitable number of any of these components 106, 108, 110, and 112 based on the particular configuration of a particular building.

[0037] The site controller 102, which may include a computer or a general-purpose processor, is configured to provide overall control and monitoring of building automation system 100. The site controller 102 may operate as a data server capable of exchanging data with various elements of system 100. Thus, the site controller 102 may allow access to system data by various applications that may be executed on the site controller 102 or other monitoring computers ( Figure 1 not shown).

[0038] For example, the site controller 102 may be capable of communicating via a management-level network (MLN) 120 with other supervisory computers, an Internet gateway, or other gateways to other external devices, and to additional network managers (which, in turn, may be connected to more subsystems via additional lower-level data networks). The site controller 102 may use the MLN 120 to exchange system data with other elements on the MLN 120, such as the reporting server 104 and one or more client stations 106. The reporting server 104 may be configured to generate reports regarding various aspects of system 100. Each client station 106 may be configured to communicate with system 100 to receive information from and / or provide modifications to system 100 in any suitable manner. The MLN 120 may include an Ethernet or similar wired network and may employ TCP / IP, BACnet, and / or other protocols that support high-speed data communication.

[0039] The site controller 102 may also be configured to accept modifications and / or other input from a user. This may be accomplished via the user interface of the site controller 102 or any other user interface that may be configured to communicate with the site controller 102 via any suitable network or connection. The user interface may include a keyboard, a touch screen, a mouse, or other interface components. The site controller 102 is configured to, among other things, affect or change the operating data of the field panels 108 and other components of system 100. The site controller 102 may use a building-level network (BLN) 122 to exchange system data with other elements on the BLN 122, such as the field panels 108.

[0040] Each field panel 108 may include a general-purpose processor and is configured to provide control of one or more corresponding field controllers 110 using data and / or instructions from the site controller 102. While the site controller 102 is generally used to modify one or more of the various components of the building automation system 100, the field panel 108 is also capable of providing certain modifications to one or more parameters of the system 100. Each field panel 108 may use a field-level network (FLN) 124 to exchange system data with other elements on the FLN 124, such as a subset of the field controllers 110 coupled to the field panel 108.

[0041] Each field controller 110 may include a general-purpose processor and may correspond to one of a plurality of localized standard building automation subsystems, such as a building space temperature control subsystem, a lighting control subsystem, etc. For a particular embodiment, the field controller 110 may include a DXR-type controller available from Siemens. However, it should be understood that the field controller 110 may include any other suitable type of controller without departing from the scope of the present invention.

[0042] To perform control of its corresponding subsystem, each field controller 110 may be coupled to one or more field devices 112. Each field controller 110 is configured to provide control of one or more corresponding field devices 112 using data and / or instructions from its corresponding field panel 108. For some embodiments, some field controllers 110 may control their subsystems based on sensed conditions and desired setpoint conditions. For these embodiments, these field controllers 110 may be configured to control the operation of one or more field devices 112 to attempt to bring the sensed conditions to the desired setpoint conditions. Note that in the system 100, information from the field devices 112 may be shared among the field controllers 110, the field panels 108, the site controller 102, and / or any other element on or connected to the system 100.

[0043] To facilitate sharing of information between subsystems, groups of subsystems may be organized into the FLN 124. For example, the subsystems corresponding to field controller 110a and field controller 110b may be coupled to field panel 108a to form FLN 124a. Each FLN 124 may include a low-level data network, which may employ any suitable proprietary or open protocol.

[0044] Each field device 112 can be configured to measure, monitor, and / or control various parameters of the building automation system 100. Examples of field devices 112 include lights, thermostats, temperature sensors, lighting sensors, fans, damper actuators, heaters, coolers, alarms, HVAC equipment, blind controllers and sensors, and many other types of field devices. The field devices 112 are capable of receiving control signals from and / or sending signals to the field controllers 110, field panels 108, and / or site controllers 102 of the building automation system 100. Thus, the building automation system 100 is capable of controlling various aspects of building operation by controlling and monitoring the field devices 112. Specifically, each field device 112 or any field device 112 can generate data processed as described herein. The physical devices of the building automation system 100 can include any one of these exemplary field devices 112.

[0045] As Figure 1 shown, any field panel 108 (such as field panel 108a) can be directly coupled to one or more field devices 112 (such as field devices 112c and 112d). For this type of embodiment, the field panel 108a can be configured to provide direct control of the field devices 112c and 112d, rather than through one of the field controllers 110a or 110b. Thus, for this embodiment, the functions of the field controller 110 for one or more specific subsystems can be provided by the field panel 108 without the need for a field controller 110.

[0046] Figure 2 Details of one of the field panels 108 according to the present disclosure are shown. For this particular embodiment, the field panel 108 includes a processor 202, a memory 204, an input / output (I / O) module 206, a communication module 208, a user interface 210, and a power module 212. The memory 204 includes any suitable data storage capable of storing data, such as instructions 220 and a database 222. It should be understood that the field panel 108 can be implemented in any other suitable manner without departing from the scope of the present disclosure.

[0047] The processor 202 is configured to operate the field panel 108. Thus, the processor 202 can be coupled to other components 204, 206, 208, 210, and 212 of the field panel 108. The processor 202 can be configured to execute program instructions or programming software or firmware stored in the instructions 220 of the memory 204, such as the BAS application 230. In addition to storing the instructions 220, the memory 204 can also store other data for use by the system 100 in the database 222, such as various records and configuration files, graphical views, and / or other information.

[0048] The execution of the BAS application 230 by the processor 202 can cause control signals to be sent to any field device 112 that can be coupled to the field panel 108 via the I / O module 206 of the field panel 108. The execution of the BAS application 230 can also cause the processor 202 to receive status signals and / or other data signals from the field device 112 coupled to the field panel 108, and store the relevant data in the memory 204, and the data can be processed as described herein. In one embodiment, the BAS application 230 can be provided or implemented by a DXR controller commercially available from Siemens Industry, Inc. However, it should be understood that the BAS application 230 can include any other suitable BAS control software.

[0049] The I / O module 206 can include one or more input / output circuits, and the I / O module is configured to communicate directly with the field device 112. Thus, for some embodiments, the I / O module 206 includes an analog input circuit for receiving analog signals and an analog output circuit for providing analog signals.

[0050] The communication module 208 is configured to provide communication with the site controller 102, other field panels 108, and other components on the BLN 122. The communication module 208 is also configured to provide communication to the field controller 110 and other components on the FLN 124 associated with the field panel 108. Thus, the communication module 208 can include a first port that can be coupled to the BLN 122 and a second port that can be coupled to the FLN 124. Each port can include an RS-485 standard port circuit or other suitable port circuit.

[0051] The on-site panel 108 can be locally accessible via the interactive user interface 210. A user can control the data collection from the on-site devices 112 through the user interface 210. The user interface 210 of the on-site panel 108 can include devices for displaying data and receiving input data. These devices can be permanently fixed to the on-site panel 108 or be portable and removable. For some embodiments, the user interface 210 can include an LCD-type screen, etc. and a keyboard. The user interface 210 can be configured to both change and show information about the on-site panel 108, such as status information and / or other data regarding the operation, function, and / or modification of the on-site panel 108.

[0052] The power module 212 can be configured to supply power to the components of the on-site panel 108. The voltage module 212 can operate on standard 120-volt AC power, other AC voltages, or DC power provided by one or more batteries.

[0053] Figure 3 Details of one of the on-site controllers in the on-site controller 110 according to the present disclosure are shown. For this particular embodiment, the on-site controller 110 includes a processor 302, a memory 304, an input / output (I / O) module 306, a communication module 308, and a power module 312. For some embodiments, the on-site controller 110 may also include a user interface ( Figure 3 not shown in the figure), which is configured to change and / or display information about the on-site controller 110. The memory 304 includes any suitable data memory capable of storing data, such as instructions 320 and a database 322. It should be understood that the on-site controller 110 can be implemented in any other suitable manner without departing from the scope of the present disclosure. For some embodiments, the on-site controller 110 can be located inside or adjacent to a room in a building, in which the temperature or another environmental parameter associated with the subsystem can be controlled by the on-site controller 110.

[0054] The processor 302 is configured to operate the on-site controller 110. Thus, the processor 302 can be coupled to the other components 304, 306, 308, and 312 of the on-site controller 110. The processor 302 can be configured to execute program instructions or programming software or firmware stored in the instructions 320 in the memory 304, such as the subsystem application 330. For a specific example, the subsystem application 330 can include a temperature control application, which is configured to control and process data from all components of the temperature control subsystem, such as temperature sensors, damper actuators, fans, and various other on-site devices. In addition to storing the instructions 320, the memory 304 can also store other data for the subsystem in the database 322, such as various configuration files and / or other information.

[0055] Execution of the subsystem application 330 by the processor 302 can cause control signals to be sent to any field device 112 that may be coupled to the field controller 110 via the I / O module 306 of the field controller 110. Execution of the subsystem application 330 can also cause the processor 302 to receive status signals and / or other data signals from the field device 112 coupled to the field controller 110 and store the relevant data in the memory 304.

[0056] The I / O module 306 can include one or more input / output circuits and is configured to communicate directly with the field device 112. Thus, for some embodiments, the I / O module 306 includes an analog input circuit for receiving analog signals and an analog output circuit for providing analog signals.

[0057] The communication module 308 is configured to provide communication with the field panel 108 corresponding to the field controller 110 and other components on the FLN 124 (such as other field controllers 110). Thus, the communication module 308 can include a port that can be coupled to the FLN 124. The port can include an RS-485 standard port circuit or other suitable port circuit.

[0058] The power module 312 can be configured to power the components of the field controller 110. The power module 312 can operate on standard 120-volt AC power, other AC voltages, or DC power provided by one or more batteries.

[0059] As described above, in commercial buildings and other facilities, HVAC equipment should be maintained regularly. Early replacement and additional maintenance will incur additional hardware / equipment and labor costs. Too little maintenance may damage the health of the machine and its remaining useful life (RUL).

[0060] Standard corrective maintenance (CM) and scheduled maintenance (SM) methods are heuristic and imprecise. For example, application engineers can estimate the mean time to failure (MTTF), mean time between failures (MTBF), or mean time before repair (MTBR) of new hardware from vendor manuals or historical log data. Then, the engineer prescribes the CM or SM method.

[0061] Under the CM framework, the engineer does not replace the equipment until it has failed, whether detected manually or automatically by a software system. Due to equipment failure, the CM method can result in significant commercial costs and loss of comfort.

[0062] Under the SM framework, engineers regularly replace equipment even if it is still operating. While this approach can reduce potential downtime, it is only applicable to low-cost equipment such as air filters. First, replacing functional hardware is a waste. Second, due to their probabilistic nature, it is inaccurate to estimate the maintenance budget assuming that equipment always fails at the MTTF time. For example, if 100 valves are deployed together in a system and the RUL of each valve is 20 years, it is unlikely that all 100 valves will fail on the same day after 5 years. Instead, the valves will fail gradually following a certain distribution, such as the Weibull distribution. Counterintuitively, the probability of a valve failing in the first year is small. Therefore, an appropriate maintenance plan should also reserve sufficient budget for the first year. To perform reliable maintenance, facility managers (FM) and engineers need to determine the RUL with higher accuracy than the SM framework.

[0063] In some maintenance analysis processes, engineers diagnose the root cause of equipment failures based on sensor data and historical maintenance data. Due to the limited sensors in the HVAC industry, faults usually cannot be directly measured by sensors. Instead, engineers need to infer the root cause based on domain knowledge and experience.

[0064] A typical maintenance prediction process may include manual root cause analysis combined with manual RUL and budget estimation, tracking data such as static failure rates in spreadsheets. Current methods are limited and ineffective because they do not incorporate survival analysis into the remaining useful life estimation (RUL). Therefore, conventional maintenance budget estimation methods are inaccurate, and the maintenance process cannot prevent high-impact failures. In addition, current technologies do not make any comparison of maintenance log data between similar devices. Therefore, the maintenance efficiency does not improve over time.

[0065] The disclosed embodiments include an automated process for predicting maintenance needs and costs using machine learning techniques.

[0066] The disclosed embodiments can use a "survival curve" for planned maintenance scheduling. A survival curve predicts the probability of a device surviving or failing over time.

[0067] Figures 4A to 4C An example of a survival curve is shown, where the x-axis is future time, measured in any time interval suitable for a particular device, and the y-axis is the probability that the device will survive to that time. This curve provides more information than an independent RUL number.

[0068] Figure 4AIt is an example of a survival curve 402 of non-parametric survival analysis using known Kaplan-Meyer estimation, generated by using, for example, maintenance log data. In this figure, band 404 reflects the 95% confidence interval of the survival curve 402 at the teaching point in the time line.

[0069] Figure 4B It is an example of a survival curve 406 of parametric survival analysis using Weibull estimation as described herein, where band 408 reflects the 95% confidence interval of the survival curve 406 at the teaching point in the time line. The survival curve 406 can be generated by using, for example, maintenance log data.

[0070] Figure 4C It is an example of a survival curve 410 of similarity-based survival analysis. As described herein, the survival curve 410 can be determined based on the similarity of operating parameters and operating conditions using multiple actual-history-based survival curves 412, as can be determined from sensor data (generally, device event data of other devices). In this example, each curve 412 reflects the health index of a device (such as a VAV) under similar conditions and with similar operating parameters, which is used to generate the survival curve 410. For example, the survival curve 410 of a specific VAV with specific operating parameters and conditions can be developed using similarity analysis of other VAVs operating under similar operating parameters and conditions, as reflected by the curve 412. Using these techniques, the survival curve 410 can be more accurate than a general survival curve because the system as described herein can use machine learning to analyze the actual survival curves 412 of similar devices and apply that analysis to generate the survival curve 410.

[0071] The disclosed embodiments perform predictive maintenance analysis for different types of HVAC equipment requirements. For example, various embodiments can estimate the survival curve of a hardware type using device survival curve estimation based on historical maintenance data without requiring sensor data. Various embodiments can use sensor data from a specific piece of hardware and use sensor-based survival curve estimation to accurately estimate the survival curve of that piece of hardware. Given data such as the cost of parts within a device, various embodiments can perform budget estimation to estimate the total maintenance cost over a future period of time.

[0072] Figure 5An example of elements of a software architecture 502 that can be implemented in a BAS or other data processing system 500 to perform the processes disclosed herein is shown. The data processing system 500 can be an example of an implementation of, for example, a site controller data processing system 102, a client station 106, a reporting server 104, or other client or server data processing system or controller configured to operate as disclosed herein. The software architecture 502 described herein is exemplary and non - limiting; particular implementations may use alternative architecture components to perform similar functions, may call the various components by different names, may combine or divide the various operations differently with respect to different components, or otherwise use different logical structures to perform the processes described herein, and the scope of the present disclosure is intended to encompass such variations.

[0073] The inference engine 504 determines a device failure, i.e., a root cause, based on events occurring in the system. Even the detection application (app) 506 is used to collect system events based on sensor data. The inference engine 504 can include a Bayesian network (BN) 520 and make its decisions based on the Bayesian network (BN) 520, which will be described in more detail below.

[0074] The input from the event detection app 506 to the inference engine 504 can include the k - th time instance, s[k], and the simultaneous expected failure rate, r[k], and other data, and can generally be referred to as device event data 522. For the i - th device, for the hardware device h i the failure rate is denoted as r i [k]. For the initial failure rate r i of a given hardware device, in some cases, it can be stored or manually input into the inference engine 504 and can be updated later by the predictive maintenance engine 508, as described below.

[0075] The output of the inference engine 504 can include the failure probability p i of the hardware device h i at the k - th time instance (current or future time instance), denoted as p i [k]. For simplicity, when the time k is obvious, p

[0076] is used here. The root cause, failure probability, and other outputs of the inference engine 504 can generally be referred to as root - cause failure data 524. i The predictive maintenance (PM) engine 508 uses survival analysis to calculate the failure probability p i [k] of the hardware device h i at the k - th time instance. For simplicity, k is ignored when there is no misunderstanding in time. The PM engine 508 calculates the failure rate r i of the hardware device hDevice h i The failure rate r i and any other updated data generated by the PM engine 508, collectively referred to as updated failure data 526, can then be fed back to the inference engine 504 to improve root cause analysis.

[0077] The PM engine 508 can include a survival analysis component 510. The survival analysis component 510 can generate a survival curve from historical data, as Figures 4A to 4C shown. Note that the output (e.g., survival curves 402, 406, 410) or the processes performed by the survival analysis component 510 can generally be referred to as "survival analysis". Also note that, as described herein, generating a curve can include actually generating a curve that can be graphically displayed to a user and, as reflected in the example figures below, but can also include simply generating the data and / or formulas required to reflect such a curve, regardless of whether the graph of the curve has ever been generated or displayed.

[0078] Figure 6 FIG. shows a non-limiting example of historical data 600 that can be used as input data for survival analysis in various embodiments. In this example, data from multiple machines / devices is shown, line data indicating loss of tracking (a flat line with no end point), line data indicating when a failure has occurred and when (the line terminates at a point indicating when the failure occurred), and line data indicating no failure yet (a line with an arrow indicating the device is still operating normally). Each line in the line data can be associated with a specific machine or device identifier. Of course, this exemplary illustration does not limit how such data is recorded, stored, or displayed. The historical data can include any number of samples, including data from hundreds or thousands of devices.

[0079] Return Figure 5 , the PM engine 508 can include a budget forecasting component 512. The budget forecasting component 512 generates a maintenance cost curve and a cost analysis based on the survival curve and / or other data.

[0080] Figure 7 FIG. shows a non-limiting example of a maintenance cost curve that can be generated as the output of the budget forecasting component 512 in various embodiments. Figure 7 FIG. shows an example of an estimated maintenance cost projection 700 that can be generated by a system as disclosed herein, a graph plotting the maintenance cost (y-axis) of an example VAV reheat valve over time (x-axis). In this figure, the maintenance cost curve 702 shows the expected cost of the VAV reheat valve at each point in its life, measured relative to its total replacement cost of $300. Curve 704 reflects the lower 95% confidence survival curve, and curve 706 reflects the upper 95% confidence survival curve.

[0081] ReturnFigure 5 For survival analysis (SA) and cost analysis (CA), the PM engine 508 can operate using multiple unique operators 514. The PM engine 508 can use "AND Survival Analysis (AND SA)", "OR SA", "AND Cost Analysis (AND CA)", and "OR CA" operators, where the AND SA or OR SA operators specify the survival relationships between multiple devices for a more advanced system that aggregates these devices. For the SA operators, the PM engine 508 combines the survival curves of individual components to represent a larger device composed of the individual components. The CA operators predict the life cycle maintenance costs of these devices based on the device costs and survival curves.

[0082] The "AND SA" operator applies to devices that require all components to work properly; that is, if any of the "ANDed" devices fails, the system fails. The "OR SA" operator is designed for the case where if one component works, the entire device works. Of course, in this discussion, AND and OR are used as exemplary operators, and various implementations may use different terms to accomplish the same aggregation function. AND CA and OR CA reflect cost analysis based on the AND SA and OR SA survival relationships.

[0083] The PM engine 508 can use a digital twin 516, which in this example is an HVAC digital twin. As used herein, a digital twin refers to a computerized (or digital) model of a physical asset / device and / or process. The digital twin can receive data representing real-time or archived information about the physical asset, such as sensor data, and can be used to model the behavior and response of the "twin" physical asset.

[0084] In a particular implementation, the HVAC digital twin 516 can include digital twins for components such as air handling units (AHUs), rooftop units (RTUs), variable air volume (VAV) units, and other elements, as well as other elements of subsystems or supersystems that include these units, any of which can be considered a physical device (such as field device 112). The digital twin 516 can include software containers that aggregate the SA and CA functions of smaller components into larger devices.

[0085] The PM engine 508 can read a time-event table from a data lake 518 or other source as described below, and then calculate survival curves (SCs) such as Figures 4A to 4C as shown in

[0086] Figure 8An example of a time-event table 800 that can be used according to the disclosed embodiments is shown. Examples of time-event tables include an event column to indicate whether an abnormal event has been detected, and other information such as the time and date of the identified event, the device and site where the event was detected, and a probability value. Such a table can include any other relevant information, such as a duration column for storing how long a given event lasted.

[0087] Using this example of data recorded from October 1 to 2, 2019, the survival curve generated shortly thereafter will indicate that the replaced valve 1 has not yet failed, and it is a "right-censored" record because its actual lifespan is not yet known. The old valve 1 failed on October 1, 2019, and its lifespan has been measured, which is referred to here as "left-censored". If the relevant value in the "manual" column is "no", the value in the "probability" column can be, for example, a probability calculated from the Bayesian network 520. If the entry in the "manual" column is "yes", the probability number is always 100, which means the result has been verified by an expert (i.e., an application engineer). The use of the BN in various embodiments is described in more detail below.

[0088] The data lake 518 represents a data repository layer that aggregates data from multiple databases of different types. Such databases can include, for example, time series databases, SQL and non-SQL databases, graph databases, and any other databases or repositories that may be useful for performing the processes described herein. Specifically, data such as access runtime and maintenance log data, as well as the above-mentioned time-event table, can be stored in the data lake 518.

[0089] Figure 9 An example output 900 of the inference engine 504 is shown, where the y-axis is p i (In this example, it represents the probability that the cooling valve fails due to improper opening), and the x-axis is time.

[0090] The disclosed embodiments can employ probabilistic parametric and non-parametric methods to estimate survival curves based on the time-event table. The system can apply the time-event table, sensor data, and the output of the BN to a probability similarity-based process as described herein to estimate the survival curve of an individual device, as described above with respect to Figure 7 shown, based on the similarity between the device event data corresponding to the specific device being analyzed and the device event data of other devices. Based on the survival curve and cost of the device, the system can also estimate the future maintenance cost of one or a group of devices.

[0091] As described above, the inference engine 504 may include a Bayesian network 520. The disclosed embodiments may combine the BN of the inference engine 504 with the survival analysis 510 of the predictive maintenance engine 508 for continuous machine learning. To achieve this, in some embodiments, the system 500 uses a novel time-event table augmentation process to connect the output of the BN from the inference engine 504 to the SA 510.

[0092] In conventional SA, the input to the survival analysis process is limited to deterministic data. However, in the disclosed embodiments, the output of the BN 520 may indicate the failure probability as shown in the time-event table such as Figure 8 and Figure 11 . In Figure 8 , data from multiple devices is mixed in a single time-event table.

[0093] Figure 10 FIG. shows an example of a standardized time-event table for a single device. After we extract the information of a device, the system can generate a time-event table 1000, where the BN may detect a failure multiple times with different probabilities. For stable regression, the system can standardize the failure time to a fixed range, which is 1.0 and 1.1 in this example. The exact range can be selected as needed and is not intended to limit the present disclosure.

[0094] Then, the system can fit the data with a selected distribution, such as a normal distribution, as Figure 11 shown

[0095] Figure 11 FIG. shows an example of fitting a sample with the cumulative density function of the normal distribution in the graph 1100, which shows the failure risk over time corresponding to the standardized time-event table 1000. The probability density function (PDF) of the normal distribution in this example is:

[0096]

[0097] This equation indicates that the variable y follows a normal distribution, where x is the input, y is the output; and σ and μ are the standard deviation and the mathematical expectation respectively; e is the Euler number.

[0098] The associated cumulative density function (CDF) y2 or f c (x) is:

[0099]

[0100] where σ and μ can be calculated using non-linear regression methods. Note that the normal distribution example used here is just one possible distribution among many alternatives that can be used in a particular implementation. In this equation, σ and μ are the standard deviation and the mathematical expectation defined as above, and erf represents the "error function" known in statistical mathematics, described at the time of submission as en.wikipedia.org / wiki / Error_function.

[0101] Then, the system can apply Figure 11 the standardized time-event table 1100, and use a common curve fitting function to fit the unknown parameters. In this example, after regression, σ = 0.0594 and μ = 1.0588. Based on these fitted parameters, the system can use the statistical "bootstrap" technique to generate many samples that follow this distribution.

[0102] Figure 12 Shows an example of this sample generation based on the fitted distribution according to the disclosed embodiments.

[0103] Then, the system can replace the data in the standardized time-event table with the generated sample data. In this example, Figure 10 the data in the standardized time-event table 1000 of Figure 12 is replaced with the generated data shown in

[0104] Figure 13 Shows an example of the augmented time-event table 1300 according to the disclosed embodiments. In this example, the augmented time-event table 1300 only has one column of the standardized failure times of the device, and the probability column is removed. Since all the enhanced probabilities are 100%, the system can process the data using standard deterministic survival analysis tools.

[0105] The system can also perform probability survival analysis. Using the augmented time-event table, the system can convert the probability output of BN 520 into deterministic data. On this basis, the system can also convert all standard deterministic SA data into probability SA data.

[0106] The system disclosed herein can perform a probability non-parametric survival analysis process. When the survival curve is unknown, the system can fit the maintenance log data using non-parametric SA methods, such as the Kaplan-Meier estimator, to generate the survival rate function S(t) as

[0107]

[0108] where k iis the time at which at least one virtual event (such as a device failure) occurs in a time-event table or a standardized time-event table, d i is a single device known to survive to time t i and n i is the population size.

[0109] The systems disclosed herein can perform a probabilistic parametric survival analysis process. If parameters such as the MTTF and / or the failure rate F for a device are known, the system can use them to fit a structured survival curve, such as a Weibull distribution, f[k], defined as:

[0110]

[0111] where f[k] is equal to W[k] and the MTTF, or F, is The Weibull distribution is a method of describing the failure rate of a device, where a represents the shape parameter, b represents the scale parameter, and k represents the input variable. The Weibull distribution is understood by those skilled in the art and is described at the time of filing as en.wikipedia.org / wiki / Weibull_distribution. For calculating F, these a and b parameters are the same as the parameters in the Weibull distribution definition. The gamma function Γ(x) is a standard mathematical function, described at the time of filing as en.wikipedia.org / wiki / Gamma_function, for example.

[0112] If the failure rate F is defined as a constant, then b = 1, and the system uses parametric SA, W[k], calculated as

[0113]

[0114] If F[k] is not constant, then the system uses

[0115]

[0116] Using a curve fitting library, such as the Python Scipy curvefit() function, the system can find a, b with nonlinear optimization using the following equation:

[0117]

[0118] where R[k] and the MTTF are, for example, from the manufacture of the device.

[0119] The systems disclosed herein can perform a survival analysis based on probabilistic similarity (SSA) process. The SSA process may be more accurate than other methods because it takes into account sensor data and thus has more information as input.

[0120] Figure 14A and Figure 14B illustrates a similarity-based survival analysis process according to the disclosed embodiments.

[0121] Figure 14A illustrates an SSA process 1400 that depends on sensor data according to the disclosed embodiments. In this figure, hardware 1402 provides sensor data for establishing a health index 1406, which reflects the survival probability of the device at a specific point in the device life cycle. Based on the sensor data, the system can use techniques such as principal component analysis (PCA) and regression analysis to construct the health index 1406. According to the health index 1406, the system can generate a similarity-based survival curve 1408 by using data from similar devices, processes, and environments. That is, for example, using the techniques described herein, a similarity-based survival curve can be generated based on the similarity between the device event data corresponding to the device and the device event data of other devices. The health index 1406 indicates the "instantaneous" survival probability of the device at a specific time point, while the survival curve 1408 represents the overall trend of the health index data when combined with data from similar devices, processes, and environments.

[0122] Figure 14B illustrates an SSA process 1450 that combines sensor data or other device event data with the output of a BN. In this figure, BN 1452 provides data p i [k] such as features, events, and failure probabilities to feature selection 1454. Feature selection 1454 selects from these data according to the survival analysis being performed and provides the selected data to establish a health index 1456. Based on the selected data, the system can use techniques such as principal component analysis (PCA) and regression analysis to construct the health index 1456. According to the health index 1456, the system can generate a similarity-based survival curve 1458 by using data from similar devices, processes, and environments. Note that BN 1452 includes and integrates sensor data.

[0123] For example, the system can perform similarity-based SA on a VAV box water valve. After the BN detects abnormal behavior of the VAV, the system can predict the remaining useful life of the VAV. In this example, in feature selection 1454, the system selects VAV data, including sensor data, from the controversial VAV. The system uses PCA and regression functions, such as functions from a machine learning library.

[0124] Figure 15A and Figure 15BShows that, according to the disclosed embodiments, for a portion of key features (sensor values or data), PCA is used on the sensor values. As Figure 15A shown, four sensor values 1502 are used as input features (in this example, discharge air temperature, supply air temperature, reheat valve command, and supply fan status). Then, in this case, the system can use the PCA feature engineering process 1504 to fit a health index from these readings. Assuming that the health index h[k] is a linear combination of the feature signals s[k], then h[k]=As[k].

[0125] In this example, the system uses four features for the VAV valve. This process can generally be applied to mechanical hardware devices. The sensors and feature inputs are different for different devices. Using singular value decomposition (SVD),

[0126] A = S∑D

[0127] where S and D are unitary matrices and ∑ is a diagonal matrix with singular values on the diagonal.

[0128] In this example, Figure 15B shows the relative feature importance PCA of the sensor data (pc1 - pc4) of the remaining useful life of the VAV reheat water valve. This process can eliminate or ignore features with small singular values. For this case, none of the features can be removed because Figure 15B none of the feature scores are close to 0.

[0129] As described above Figure 4C shows the resulting probability - similarity - based survival curve 410 for the VAV reheat valve example according to the disclosed embodiments. As Figure 4C shown, the probability - similarity - based survival analysis is applied to the VAV reheat valve. The x - axis of the survival curve 410 is time. The y - axis is the health index (survival probability) of the survival curve 410. Due to different usage patterns, some valves degrade faster than others, as shown by the other curves 412.

[0130] Different regression methods can be used in different embodiments. Products such as the XGBOOST open - source software library can be used because it has strong performance, can achieve high accuracy, can reduce false positives and false negatives, can achieve data scalability by performing parallel computing on random forest estimates, and has the ability to prevent overfitting of training data by using random forest structures, early stopping, and bagging techniques.

[0131] The disclosed embodiments may perform a cost analysis (CA) based on survival analysis. Given the equipment cost and the survival curve of the equipment, the system may estimate the maintenance cost, C[k]. For facility managers and other individuals, it is important to reserve sufficient maintenance budget at the beginning of the fiscal year. Simply retaining the same amount of funds each year is ineffective because as the equipment ages, higher maintenance costs are expected before the old equipment is replaced by new equipment.

[0132] Figure 16 An example of a cost analysis process 1600 according to the disclosed embodiments is shown. In such a process, the component cost 1602 and the survival analysis 1604 of the component may be used to generate a maintenance cost estimate 1606 (also referred to as a cost analysis 1606). Then the maintenance cost estimate 1606 may be used to calculate a cost budget plan 1608.

[0133] For the survival analysis 1604, in one example, if the equipment cost is c, the non-parametric SA can be calculated as

[0134] C[k] = c·S[k],

[0135] And the parametric SA can be calculated as

[0136] C[k] = c·W[k].

[0137] In this process, define as the general SA, that is, for non-parametric SA and for parametric SA. Then,

[0138] Figure 17A and Figure 17B show an example of calculating a cost analysis from a survival analysis according to the disclosed embodiments. Figure 17A Shows an example of the survival analysis of a leaking VAV reheat valve; Figure 17B Shows a corresponding example of the maintenance cost curve. Figure 17B Shows the same example as above Figure 7 Same example.

[0139] For budget planning, the total cost at a specific time may be important. For example, it may be valuable to know C

[52] , where 52 is the number of weeks, as an estimated cost for the next year.

[0140] The disclosed embodiments may use AND / OR operators to perform an aggregation process to combine the survival analysis and cost analysis from each component to the entire system. The SA and CA of large equipment or a set of equipment (such as a building cluster) are respectively represented as and C[k]. The SA and CA of the i-th component in a large device or a set of devices are respectively and C i [k].

[0141] If multiple components of a device need to work properly, then the SA or CA of the entire system can be calculated using the AND operator on the SA or CA. The AND SA operator is defined herein as

[0142]

[0143] where:

[0144]

[0145] Similarly, the AND CA operator is defined herein as

[0146]

[0147] where:

[0148]

[0149] The OR SA operator is defined herein as

[0150]

[0151] where:

[0152]

[0153] The OR CA operator is defined herein as

[0154]

[0155] where:

[0156]

[0157] Figure 18 Shows the AND CA aggregation of the maintenance cost curve according to the disclosed embodiments. In this example, the CA curve 1802 (for VAV 324E) is combined with the CA curve 1804 (for VAV 320A) using the AND CA operation to produce the CA curve 1806.

[0158] As Figure 18 shown, the AND CA operation means that if VAV 324E and VAV 320A fail, then a maintenance team is required to repair these components. Thus, the maintenance cost curves of the AND operation are combined and injected with the probabilities of the survival probabilities of the two VAVs as the maintenance cost.

[0159] In other examples, if a single component fails and a second component must make up the "shortfall," the single survival curve of the second device may change because it is carrying an additional load. The OR SA and OR CA aggregation techniques discussed herein can explain the interaction between component failures and changes in survival curves.

[0160] The system can also determine confidence intervals associated with all maintenance cost curves to indicate the highest possible amount to be spent and the lowest possible amount spent on maintenance costs. In this way, the aggregated maintenance cost curve, such as curve 1806, helps facility managers and other individuals estimate the maintenance cost budget for an entire HVAC system with multiple VAVs.

[0161] Figure 19 A process according to the disclosed embodiments is shown, which can be performed, for example, by a data processing system, a controller, or other processor in a BAS system or other system, or any combination of multiple such systems. Devices that perform such a process are collectively referred to herein as "systems." Any or all of the features discussed here can be used in the methods described below.

[0162] The system receives device event data (1902). As used in this process, "receiving" can include loading from memory, receiving from another device or process, receiving via interaction with a user, or other means. In a particular embodiment, the device event data is received directly or indirectly from one or more event detection applications that identify device or system events based on sensor data. In some embodiments, the device event data can include sensor data of a device at a particular time instance and the failure rate of the device at the same time instance. The sensor data and the failure rate can include historical data, such as runtime data and maintenance log data stored in a data lake repository. In a BAS implementation, the device can be any HVAC or other building device, and the sensor data can be any data received from field devices 112, field controllers 110, or other devices.

[0163] The system executes an inference engine to determine root cause failure data corresponding to the device event data (1904). The root cause failure data can include the probability of device failure at a particular time instance or a future time instance. The inference engine can include a Bayesian network that associates device events with device failures and makes decisions based on that network. The inference engine can combine the device event data with the output of the Bayesian network to produce the root cause failure data. Known techniques for determining the root cause of a failure based on events can be used, and techniques such as those described can be used.

[0164] The system executes a predictive maintenance engine to generate a survival analysis (1906) of a device based on root cause failure data. The survival analysis can include one or more survival curves as disclosed herein. The predictive maintenance engine can use digital twins of one or more devices to generate the survival analysis. The predictive maintenance engine can use operators such as AND and OR operators to combine device survival analysis data, thereby aggregating device data for the survival analysis. The predictive maintenance engine can be, for example, the predictive maintenance engine 508 described above.

[0165] The survival analysis can include performing a parametric survival analysis process, a non-parametric survival analysis process, and / or a survival similarity analysis (SSA) process based on probability. The survival analysis can include using a time-to-event table, a normalized time-to-event table, and / or an augmented time-to-event table to process the root cause failure data. In the SSA process, the system can also perform a principal component analysis and / or a regression analysis on selected device event data to establish a health index representing the survival probability of the device. Singular value decomposition calculations can be performed as part of the principal component analysis. The survival analysis can include generating a normalized time-to-event table from the time-to-event table using a cumulative density function and curve fitting, and thereafter generating an augmented time-to-event table that replaces the event time values and / or probability metrics with normalized time values.

[0166] The system can execute a predictive maintenance engine to generate a cost analysis (1908) corresponding to the device. The predictive maintenance engine can use AND and OR operators to aggregate device data for the cost analysis to combine device cost analysis data, and can use AND CA or OR CA operators to generate an aggregated cost analysis for multiple devices based on the survival relationships between the multiple devices.

[0167] The system generates updated failure data based on the survival analysis through the predictive maintenance engine and provides the updated failure data to an inference engine (1910). The inference engine can then use the updated failure data in subsequent root cause analysis to provide more accurate root cause failure data, thereby creating an ongoing machine learning process.

[0168] The system can output or display the updated failure data, survival curves, survival analysis, cost analysis, or any other such output described above to the user, and / or can store or transmit such outputs that may be useful in a given implementation (1912). Additionally, as part of this step, the system can predict, schedule, or command the replacement of physical devices based on the survival analysis (including the survival relationships between multiple devices in some cases).

[0169] Figure 20illustrates a process according to the disclosed embodiments, which can be performed, for example, by a data processing system, a controller, or other processor in a BAS system or other system, or any combination of multiple such systems, as the above-described predictive maintenance engine 508 or Figure 19 the predictive maintenance engine described in

[0170] The system receives failure probability data (2002). For a device or component, the failure probability data can be or correspond to, for example, the root cause failure data discussed above. The failure probability data can include, for example, the device h i of interest, and the failure probability p i received from the inference engine.

[0171] The system generates an augmented time-event table (2004) based on the failure probability data. For example, as described above with respect to Figures 10 to 14A and Figure 14B , this can be performed to produce an augmented time-event table that replaces the probability metric with a normalized time value for the device or component.

[0172] The system uses the augmented time-event table to generate one or more survival curves for the device or component (2006). This can include generating a survival curve using the Kaplan-Meyer estimates as shown in Figure 4A , generating a Weibull estimate as shown in Figure 4B , or generating a similarity-based survival curve as shown in Figure 4C using sensor data and operating parameters from similar devices. This can specifically include using the similarity between the augmented time-event table and the device event data corresponding to the device or component and the device event data of other devices (such as other devices subject to the same operating conditions and parameters) to generate the similarity-based survival curve described herein.

[0173] In some cases, processes 2002 to 2006 can be used to implement the process of executing the predictive maintenance engine to produce a survival analysis of the device based on the root cause failure data in 1906 above.

[0174] The system generates a cost curve (2008) using one or more survival curves. The system can use the survival curve of the device or component in combination with the cost data of the device or component to generate a cost curve for the device or component. The cost curve can be, for example, the cost curve shown in Figure 7 .

[0175] In some cases, process 2008 can be used to implement a process that executes a predictive maintenance engine to generate a cost analysis corresponding to the devices in 1908 above.

[0176] The system can generate a higher-level system curve (2010) using AND SA, OR SA, AND CA, and / or OR CA operators. The higher-level system can be any aggregation of lower-level devices or components, such as an air handling system for a building floor, an HVAC system for an entire building, physical devices for a campus, or others. This can include generating a survival curve for an aggregation of multiple devices using AND SA or OR SA operators that specify the survival relationship between multiple devices, and can include generating a cost analysis for an aggregation of multiple devices using AND CA or OR CA operators based on the survival relationship between multiple devices. The generated higher-level system curve can include a survival curve and / or a cost curve of the higher-level system. In this way, the system can predict maintenance problems, equipment failures, and associated costs for any combination of devices analyzed as described herein.

[0177] Process 2010 can be executed as part of executing a predictive maintenance engine to generate a survival analysis of a device based on the root cause failure data in 1906 above, and / or can be executed as part of executing a predictive maintenance engine to generate a cost analysis corresponding to the devices in 1908 above.

[0178] The system can estimate the future failure rate (2012) of a device or component. Then, the system can send the future failure rate to an inference engine. For device h i the future (predicted) failure rate p i [t] can be estimated using the survival curve generated for the device as described above.

[0179] Figure 21 An example of using the logical operators described herein is illustrated. In this example, assume that the system has generated certain survival curves and cost curves for various devices as described above - survival curve A 2102 and cost curve A 2112 for device A, survival curve B 2104 and cost curve B 2114 for device B, and survival curve C 2106 and cost curve C 2116 for device C. In this example, consider that the operation of devices A and C in a given subsystem AB 2120 is necessary and complementary. If any one of devices A or C fails, then subsystem 2120 fails. AND SA 2122 reflects the aggregated survival curve of subsystem AB, and combining survival curve A 2102 and survival curve B 2104 using AND SA 2122 shows the survival prediction of subsystem AB 2120 where both devices A and B survive.

[0180] Since the cost curve is calculated from the survival curve, this also shows that the cost curve prediction for subsystem AB 2120 is the combination AND CA 2126 of cost curve A 2112 and cost curve B 2114.

[0181] Also consider that the survival of system AB-C 2130 requires the survival of subsystem AB or the survival of device C. In this example, the system can use OR SA 2124 that combines the survival curve C 2106 and the result of AND SA2122 to generate a system survival curve 2108 that reflects the survival of system AB-C 2130.

[0182] Since the cost curve is calculated from the survival curve, this also shows that the cost curve 2118 prediction for system AB-C 2130 is the OR CA 2128 that combines the AND CA 2126 cost curve generated from the combination of cost curve A 2112 and cost curve B 2114 with cost curve C 2116.

[0183] Figure 22 A block diagram of a data processing system 2200 that can implement various embodiments is shown. The data processing system 2200 is Figure 1 an example of an implementation of the site controller data processing system 102 in Figure 5 and an example of an implementation of the data processing system 500 in , and can be used as an implementation of other data processing systems configured to operate as described herein.

[0184] The data processing system 2200 includes a processor 2202 connected to a secondary cache / bridge 2204, and the secondary cache / bridge 2204 is in turn connected to a local system bus 2206. The local system bus 2206 can be, for example, a Peripheral Component Interconnect (PCI) architecture bus. In the depicted example, the main memory 2208 and the graphics adapter 2210 are also connected to the local system bus 2206. The graphics adapter 2210 can be connected to a display 2211.

[0185] Other peripheral devices such as a local area network (LAN) / wide area network (WAN) / wireless (e.g., WiFi) adapter 2212 can also be connected to the local system bus 2206. The expansion bus interface 2214 connects the local system bus 2206 to the input / output (I / O) bus 2216. The I / O bus 2216 is connected to a keyboard / mouse adapter 2218, a disk controller 2220, and an I / O adapter 2222. The disk controller 2220 can be connected to a memory 2226, which can be any suitable machine-usable or machine-readable storage medium, including but not limited to non-volatile, hard-coded types of media such as read-only memory (ROM) or erasable, electrically programmable read-only memory (EEPROM), magnetic tape memory, and user-recordable types of media such as floppy disks, hard disk drives, and compact disc read-only memory (CD-ROM) or digital versatile disc (DVD), as well as other known optical, electrical, or magnetic storage devices.

[0186] The memory 2226 can store any program code or data useful in performing processes as disclosed herein or for performing building automation tasks. In a particular embodiment, the memory 2226 can include elements such as device event data 2252, root cause failure data 2254, analysis including survival analysis, cost analysis, and correlation curves and data 2256, and other data 2258, as well as a stored copy of the BAS application 2228. The other data 2258 can include software architecture, any of its elements, or any other data, programs, code, tables, data lake repositories, or other information or data discussed above.

[0187] In the example shown, an audio adapter 2224 is also connected to the I / O bus 2216, and speakers (not shown) can be connected to the audio adapter 2224 to play sound. The keyboard / mouse adapter 2218 provides connections for pointing devices such as a mouse, trackball, track pointer, etc. (not shown). In some embodiments, the data processing system 2200 can be implemented as a touchscreen device such as a tablet or touchscreen panel. In these embodiments, the elements of the keyboard / mouse adapter 2218 can be implemented in combination with the display 2211.

[0188] In various embodiments of the present disclosure, the data processing system 2200 can be used to implement a workstation or a site controller 102, where all or part of the BAS application 2228 is installed in the memory 2208, configured to execute the processes described herein, and is generally available as the BAS described herein. For example, the processor 2202 executes the program code of the BAS application 2228 to generate a graphical interface 2230 displayed on the display 2211. In various embodiments of the present disclosure, the graphical user interface 2230 provides an interface for a user to view information about and control one or more devices, objects, and / or points associated with the building automation system 100. The graphical user interface 2230 also provides a customizable interface that presents information and controls in an intuitive and user-modifiable manner.

[0189] Those of ordinary skill in the art will understand that Figure 22 the hardware depicted can vary for a particular implementation. For example, other peripheral devices such as optical disk drives can be used in addition to or in place of the hardware depicted. The examples depicted are provided for illustrative purposes only and are not meant to imply limitations regarding the architecture of the present disclosure.

[0190] If appropriately modified, one of various commercial operating systems can be used, such as a version of Microsoft Windows, a product of Microsoft Corporation located in Redmond, Washington. TM As described, the operating system can be modified or created in accordance with the present disclosure, for example, to implement object discovery and generation of a hierarchy for the discovered objects.

[0191] The LAN / WAN / WiFi adapter 2212 can be connected to a network 2232, such as Figure 1 the MLN 120 in. As further explained below, the network 2232 can be any public or private data processing system network or a combination of networks known to those skilled in the art, including the Internet. The data processing system 2200 can communicate with one or more computers via the network 2232, which are not part of the data processing system 2200 but can be implemented as, for example, a separate data processing system 2200.

[0192] Of course, those skilled in the art will recognize that some of the steps in the above processes can be omitted, performed simultaneously or sequentially, or in a different order, unless the operation sequence is specifically indicated or required.

[0193] Those skilled in the art will recognize that, for simplicity and clarity, the full structure and operation of all data processing systems suitable for use with the present disclosure are not depicted or described herein. Instead, only those data processing systems that are unique to the present disclosure or necessary for understanding the present disclosure are so depicted and described. The remainder of the construction and operation of the systems used herein may conform to any of a variety of current implementations and practices known in the art.

[0194] The disclosed embodiments provide significant advantages over other systems. For example, the disclosed method can reduce the operating cost of an HVAC system by determining the remaining useful life of VAV components, forecast maintenance budgets for facility management, and prevent downtime by forecasting fault occurrences using sensor and meter data.

[0195] By connecting Bayesian network-based fault detection and diagnosis with a survival analysis process and providing continuous feedback between these processes, the disclosed embodiments produce a "lifelong" machine learning system that continuously updates and improves the survival analysis of one or more devices, particularly in a BAS. The disclosed embodiments ensure that the performance of the machine learning process can be improved as more data is collected and processed. Lifelong machine learning is a desirable feature for "big data" applications.

[0196] When an abnormal or faulty hardware device is detected, the disclosed embodiments can apply SA methods to estimate the remaining useful life of that device and the associated group of devices. Such a process can combine BN and SA techniques to improve accuracy in a closed-loop manner. Thus, the prediction accuracy of both BN and SA can be improved as more data is collected. The disclosed SA process can use probabilistic or deterministic time-event data, while other SA processes only accept deterministic data.

[0197] The disclosed embodiments can accept the output from a BN and uncertain maintenance log data from a real-world HVAC job log. The disclosed embodiments can convert probabilistic time-event data into deterministic time-event data and apply the SA process, thus carefully converting the deterministic SA method process into a corresponding probabilistic process.

[0198] The disclosed SA process is not only applicable to conventional maintenance log data, but can also utilize sensor data and BN output to improve performance. In cases where conventional SA methods only estimate the average RUL of a class of devices, the disclosed embodiments can estimate the RUL of a device based on its usage pattern and thus be more accurate than other methods. Additionally, for devices with limited sensor measurements, the disclosed embodiments can collect data from associated devices and use BN to infer the state of the device of interest.

[0199] The disclosed technology can be applied to groups or aggregations of devices for joint survival analysis and cost analysis. Using novel operators such as AND SA, AND CA, OR SA, and OR CA, the disclosed embodiments can aggregate devices together for joint SA or CA.

[0200] It is important to note that although this disclosure includes a description in the context of a full-featured system, those skilled in the art will understand that at least portions of the mechanisms of this disclosure are capable of being distributed in any of a variety of forms as instructions contained in a machine-usable, computer-usable, or computer-readable medium, and this disclosure applies equally regardless of the particular type of medium used to actually execute the distributed instructions or signal-carrying medium or storage medium. Examples of machine-usable / readable or computer-usable / readable media include: non-volatile, hard-coded type media such as read-only memory (ROM) or erasable, electrically programmable read-only memory (EEPROM), and user-recordable type media such as floppy disks, hard disk drives, and compact disc read-only memory (CD-ROM) or digital versatile disc (DVD).

[0201] Although the exemplary embodiments of this disclosure have been described in detail, those skilled in the art will understand that various changes, substitutions, variations, and improvements disclosed herein can be made without departing from the spirit and scope of the broadest form of this disclosure. Specifically, any feature disclosed herein can be combined with any feature described in this application and other documents.

[0202] No description in this application should be construed as implying that any particular element, step, or function is an essential element required to be included in the scope of the claims: the scope of the patent subject matter is defined only by the allowed claims. Additionally, none of these claims are intended to invoke 35 U.S.C. § 112(f) unless the exact word "means" is followed by a participle.

Claims

1. A method for a building automation system, the method being executed by a data processing system of the building automation system and comprising: Receiving device event data corresponding to a physical device of the building automation system; Executing an inference engine to determine root cause fault data corresponding to the device event data; Executing a predictive maintenance engine to generate a survival analysis of the physical device based on the root cause fault data; Generating, by the predictive maintenance engine, updated fault data based on the survival analysis and providing the updated fault data to the inference engine, wherein the inference engine thereafter uses the updated fault data in subsequent root cause analysis; and Outputting the survival analysis, wherein executing the predictive maintenance engine to generate the survival analysis comprises: Receiving the root cause fault data; Generating an augmented time-event table based on the root cause fault data; and Using the augmented time-event table and the similarity between the device event data corresponding to the physical device and the device event data of other devices to generate a similarity-based survival curve.

2. The method according to claim 1, further comprising generating an aggregated survival curve for the plurality of devices using an AND SA or OR SA operator that specifies a survival relationship between the plurality of devices.

3. The method according to claim 1, wherein The device event data is received directly or indirectly from one or more event detection applications that identify device or system events based on sensor data.

4. The method according to claim 1, wherein, The root cause fault data includes a failure probability of the physical device for a particular time instance.

5. The method according to claim 1, wherein The inference engine combines the device event data with the output of a Bayesian network to generate the root cause fault data.

6. The method according to claim 1, wherein The survival analysis includes one or more survival curves generated by performing a probabilistic parametric survival analysis process, performing a probabilistic non-parametric survival analysis process, or performing a probability similarity-based survival analysis process.

7. The method according to claim 1, further comprising, based on the survival analysis, executing the predictive maintenance engine to generate a cost analysis corresponding to the physical device.

8. The method according to claim 1, further comprising, based on the survival analysis, executing the predictive maintenance engine to generate a cost analysis corresponding to the physical device, and generating an aggregated cost analysis for the plurality of devices using an AND CA or OR CA operator based on the survival relationship between the plurality of devices.

9. The method according to claim 1, wherein, The survival analysis includes performing a probability similarity-based survival analysis process and includes performing a principal component analysis and a regression analysis on selected device event data to establish a health index representing the survival probability of the physical device.

10. A building automation system includes a plurality of physical devices and at least one data processing system, the at least one data processing system being configured to process device event data corresponding to the plurality of physical devices of the building automation system, wherein, The building automation system is configured to: Receive device event data corresponding to a physical device among the plurality of physical devices of the building automation system; Execute an inference engine to determine root cause fault data corresponding to the device event data; Execute a predictive maintenance engine to generate a survival analysis of the physical device based on the root cause fault data; The predictive maintenance engine generates updated failure data based on the survival analysis and provides the updated failure data to the inference engine, where the inference engine thereafter uses the updated failure data in subsequent root cause analysis; and Output the survival analysis, where performing the predictive maintenance engine to generate the survival analysis includes: Receiving the root cause failure data; Generating an augmented time-event table based on the root cause failure data; and Using the augmented time-event table and the similarity between the device event data corresponding to the physical device and the device event data of other devices to generate a similarity-based survival curve.

11. The building automation system according to claim 10, wherein, The building automation system is further configured to generate an aggregated survival curve for the plurality of devices using an AND SA or OR SA operator that specifies the survival relationship between the plurality of devices.

12. The building automation system according to claim 10, wherein, The device event data is received directly or indirectly from one or more event detection applications that identify device or system events based on sensor data.

13. The building automation system according to claim 10, wherein, The root cause failure data includes the failure probability of the physical device for a specific time instance.

14. The building automation system according to claim 10, wherein, The survival analysis includes one or more survival curves generated by performing a parametric probability survival analysis process, performing a non-parametric probability survival analysis process, or performing a probability similarity-based survival analysis process.

15. The building automation system according to claim 10, wherein, The building automation system is further configured to, based on the survival analysis, perform the predictive maintenance engine to generate a cost analysis corresponding to the physical device.

16. The building automation system according to claim 10, wherein, The building automation system is further configured to, based on the survival analysis, perform the predictive maintenance engine to generate a cost analysis corresponding to the physical device, and generate an aggregated cost analysis for the plurality of devices using an AND CA or OR CA operator based on the survival relationship between the plurality of devices.

17. The building automation system according to claim 10, wherein, The survival analysis includes performing a probability similarity-based survival analysis process and includes performing principal component analysis and regression analysis on selected device event data to establish a health index representing the survival probability of the device.

Citation Information

Patent Citations

  • Systems and methods for using rule-based fault detection in a building management system

    US20110047418A1

  • System for predicting equipment failure events and optimizing manufacturing operations

    US20200265331A1